Originally posted on LinkedIn as a four-part series: Part 1, Part 2, Part 3, Part 4.
How do you build a great AI interaction for an enterprise customer who is not technical and does not want to be?
That’s what I’m working on at Levelpath right now, and I’d guess most product teams could say the same. It’s new, nobody has the patterns yet, and the ones that exist change every few months. So here’s where I’ve landed.
Underneath it there are three separate questions we’re all reasoning through:
Do you talk to one agent or many?
How does the customer interact with agents that run on their own?
Is it one never-ending conversation, or separate sessions?
1. One agent or many
The hot new pattern is a roster of named agents. Grok Bot nudges you to create a chief of staff, an inbox bot, an expenses bot, and pick who to talk to. It’s familiar. I talk to my procurement manager and my AP specialist the same way. It’s nice to know who you’re talking to.
But it pushes routing onto the customer. They have to know which agent is the right one.
We’ve been here once. When Levelpath started in 2022, LLMs could do far less and context windows were small, so we scoped the assistant to the page you were on. It frustrated users. The contracts page and the suppliers page each had an assistant, but they were different ones, and nothing you said to one carried over to the other. That’s the same problem the roster creates: nobody wants to track which agent knows what. We were already headed toward one universal assistant, the union of every role, pulling in the right skills and data as the task needs. The roster pattern only reinforced it.
Under the hood the universal assistant is still routing: to different skills, different data, effectively different specialists. There may be value in showing that. Grafana’s assistant tells you which skills it’s using as it works. You never ask for the AP specialist; it switches and you can see that it did. I think that’s useful mostly as information, and maybe more for me debugging than for the customer. Not sure yet.
In any big organization, getting one thing done means hunting down ten people. It would be simpler if one person knew it all. In human terms that’s impossible: roles exist because a person can only execute so much, so at scale you end up with a strategic sourcing manager, an AP specialist, a lawyer, a category manager. A sourcing manager agent feels familiar and safe to a company adopting AI. But it’s anthropomorphizing in a way that gives away the one thing AI is uniquely good for: a single entity that can actually do all the roles.
2. How does the customer interact with agents that run on their own?
A growing share of what an AI product does happens with nobody watching. An invoice arrives and an agent processes it. A supplier’s news gets scanned every day for adverse media. These agents are narrow, goal-driven, and mostly silent.
One hard problem with them is showing the customer what they did on their behalf. That’s a real design challenge, and it’s out of scope here.
What’s in scope is the moment a background agent needs a human. The adverse media agent finds something on a critical supplier. It could kick off a review process on its own, and for most suppliers that’s exactly what it does. But the customer configured it to check in before acting on anything tied to a critical supplier, so it stops and asks. Now the user has to steer something that mostly runs unattended and only occasionally surfaces.
The bot-first products handle this by making the background agent its own thing to talk to. That’s the roster problem again. The customer now has to know that the supplier risk bot exists, find it, and remember what it knows.
The answer I keep landing on is that the escalation should arrive in the same universal assistant from part 1. The agent pauses, the question shows up where the customer already is, the customer answers, the agent resumes. If the customer wants the agent to behave differently next time, they say so in that same conversation.
Under the hood I still want those agents to be discrete, with defined, readable prompts, so a non-technical admin can open one and see exactly what it will and won’t do. But that’s structure. The customer interacts with one thing.
What I haven’t settled is what that one thing is when it pauses. Is my assistant relaying a message from the supplier risk agent and carrying my answer back? Or does it take on the risk agent’s context, so I’m talking to the combination? Relay is the familiar model, a chief of staff telling you what someone on the team said. Merge is the AI-native one, since the model is the only thing in the system that can hold both contexts at once. With the right UX either can feel like a single conversation.
If you’ve built one of these, which did you pick, and would you pick it again?
3. Is it one never-ending conversation, or separate sessions?
The new paradigm is the conversation that never ends. Grok Bot and Muse both do this: each agent is one unbounded thread. You never start a new chat. You just keep talking to the same thing, for months.
What that buys you is the absence of a seam: no new-chat button, no deciding what carries forward. For a non-technical customer that’s appealing. They shouldn’t have to think about context at all.
The trouble starts when you ask what the customer is expecting from it. The natural expectation is that the thing remembers everything that was ever said. It doesn’t. Users report Grok Bot reloading 200 to 250k tokens on every reply, and it still summarizes the transcript as it nears the limit. It just doesn’t tell you. Muse is less forthcoming about how it works or what it costs, but the latency is similarly sluggish, and one can only guess it’s doing something close to the same thing. So the customer pays 5 to 15 times the cost of a compacted context, plus seconds on every message, for a promise of total recall that compaction quietly breaks.
There’s a plainer problem too. I want threads. (Remember Slack before threads? The horror!) On a normal day I’m working ten different things, and I want to click into each one, not scroll through a whole day of interleaved chat to find where a conversation left off. A never-ending chat might be fine for the quick back-and-forth with an assistant. For anything that’s a job, it’s the wrong shape.
What I want instead is closer to how a good human assistant works. They haven’t memorized everything you’ve said to them over the years. What they have is a sense of how you like to work and what you do often. Hand them a new job and they know whether it’s a repeat, and if it is, they can do most of it or at least tell you what comes next.
That’s a fine starting model. New thing to do, hit new. That’s the whole burden on the user. Everything about that job stays in its thread until the job is done, whether that takes a minute or a month. The assistant brings how you work and what you’ve done before. Things you ask for repeatedly become skills. If this is something you did last quarter, it says so and suggests where to start, instead of pretending it never forgot. And where it beats the human assistant: it has every past transcript and all the data, and it can look any of it up in seconds. The exact record comes from a lookup, not from hoping it’s still in the context.
If you’ve lived with the never-ending version for a while, how is it working for you?
4. What it costs to build
Across the first three sections the answer kept coming out the same way. Named agents, a separate bot for each background job, one thread that never ends: each is a real piece of the system’s structure pushed out onto the customer. Every time I checked a design against how I’d treat a human in that role, the human analogy helped a little and then started costing something only the AI could give. So the spec is short. One assistant, every role at once, new thread per job, memory for the shape of how you work, lookup for anything exact.
That is a single entity that is simultaneously the sourcing manager, the AP specialist, the category manager, and the lawyer, that knows how you work without being told, that has read every conversation you’ve had with it and can find any one of them in seconds, and that runs a dozen jobs for you in parallel without you keeping track of any of them. No human has ever been that. That’s what we’re asking for, and the only reason it’s askable is that the intelligence behind it is not a person.
None of that structure goes away. It moves to engineering. Here is what we have to solve to make the customer’s version that simple.
Routing. The customer no longer picks the specialist, so the assistant has to, on every turn: which skills, which data, and only the permissions this job calls for. Too loose and it answers an AP question like a sourcing manager. Too tight and it can’t do the job. Too broad and it has access it shouldn’t.
Context assembly. Every new thread starts empty. Filling it with the right working patterns, skills, data, and prior runs, inside a budget that keeps it fast, is the whole game.
Memory. How you work and what you do often should come forward on its own. What was actually said or done should never sit in context by default. It comes from fast, permission-aware retrieval that the assistant trusts more than its own guess.
That’s my current thesis on the ideal assistant inside an enterprise SaaS product. I haven’t seen anyone fully build it yet, and some of what it needs are hard engineering problems I haven’t seen anyone else solve either. Those are the problems we’re working on now.
