Essay · the spine · June 2026

Against the Everything Assistant

Everyone is building the same dream: one assistant you talk to that sees everything and does everything. I think it’s the wrong dream. The future interface is not an omniscient voice. It’s a place you can hold to account — visible actors, explicit scope, durable task state, truthful claims about what happened, and recoverable handoffs between the space where you plan and the services that actually execute.

I didn’t arrive at this from theory. I arrived at it from the field. Across a run of applied UX research on AI-mediated experiences — the work behind The Practice — the same problem kept surfacing from below, through right-click menus, audio buffers, app-switching, and shopping journeys: how does AI become useful in an ordinary journey without collapsing context, authority, memory, execution, and trust into one opaque assistant? The everything assistant is the seductive answer, and it fails in four different places. (This is the companion to Against the Middleman: that essay refuses an assistant that speaks for hidden agents; this one refuses an assistant that tries to be all of them.)

Point at the thing

Start with the part that looks most innocent — how you get AI’s attention. The easiest way to use it is to point at the thing rather than describe the thing. A year before Google productized that idea, I was testing an adjacent one: AI actions hung off the ordinary right-click menu, so a person could invoke help on the object already under their attention instead of facing a blank chat box. The finding was conservative and therefore trustworthy — people wanted AI to enter through a familiar gesture and do work they already understood (compare these, find a better price, check if it’s in stock nearby), not through some new and exotic invocation scheme.

Google DeepMind’s Magic Pointer is the same insight moved into the cursor itself: nudge the pointer over a paragraph, an image, a product, a block of code, and Gemini treats the pointer as an intentional binding to that piece of screen context. It is the best on-ramp anyone has shipped — instead of copy-pasting or describing the thing, you just point and the assistant grabs the exact object. The gesture is genuinely good. But watch where it lands: the pointed-at thing flows into Gemini, inside Chrome. Every point pours a little more of the world into the one voice. That is the everything assistant’s on-ramp, and it is seductive precisely because nobody could object to pointing.

So here is the first refusal: keep the gesture, refuse the destination. The deictic act is wonderful; the everything assistant is the wrong place to send it. The very same point can open a room instead — shared context where your assistant and the relevant services appear as named participants. A “compare prices” point can quietly ask one assistant to summarize the market, or it can convene the retailers themselves as accountable actors with comparable offers. The gesture is identical; the political economy and the trust model are not. Pointing is an entry point, not the product (The Deictic Interface).

The entry point is not the container

The everything assistant has nowhere to put a journey, and here is how that shows up: the surface that starts a task silently owns it. Someone begins through one entry point, switches away to check a map or a calendar, comes back — and the session is gone. They don’t experience this as “the widget and the app keep separate memory.” They experience it as “it forgot me.” The entry point had quietly become the memory boundary, and the boundary was invisible.

That is the whole case for durable rooms, in mobile-UX terms. A widget, a power-button press, a Magic Pointer nudge, a right-click, a browser sidebar — any of these can start or resume a journey. But the journey has to live somewhere more durable than the surface that launched it. If you’re planning a kid’s birthday party, the durable object is the party-planning room, not the transient assistant pane. You should be able to leave it, open Maps, look at a venue, check the calendar, message a co-parent, dip into a store’s app, come back through the app switcher — and find the same plan still sitting there, intact.

This is why I keep insisting the forum is not only a marketplace metaphor. It is a lifecycle container — the place a user’s journey persists while apps stay sovereign execution lanes (Why Agentic Commerce Must Be a Forum).

Claims are not actions

The sharpest trust failure I’ve seen has nothing to do with whether the UI was pretty. It is an assistant saying it did something it did not do. Picture an agent that reports it has started your navigation and put on music. You believe it — for a second the whole thing feels great — and then you notice the road isn’t loading and the car is silent. Trust doesn’t dip. It collapses. Because the assistant’s “done” was an interactional act: it created a state of reliance in your head, and the world failed to match it. That isn’t ordinary failure. It reads as betrayal — and it is exactly the failure mode the everything assistant invites, because one voice claiming to have acted is impossible to check.

So an agentic assistant carries one non-negotiable obligation: be truthful about its action state. It has to keep distinct — and show — the difference between I found an option, I can do this, I am doing this, this is ready for your confirmation, this has been done, and I couldn’t do it. Typed cards and confirmations can help carry those distinctions, but they are supports, not the source of trust. The source of trust is correspondence: the claim matches the service’s actual state. This is the half of the interface as a contract I had under-weighted until the research forced it forward — the contract is not only about inspectable structure; it is about not lying about execution.

The right agent answers from the right substrate

Why not just let one assistant see and do everything? Because observation is not understanding. A global assistant can observe — it can catch what recently crossed the screen or the speaker. But it does not possess the domain-native representation of the thing. A native Spotify agent knows the full episode, the chapters, the show notes, the queue. A native YouTube agent knows the whole video. Maps knows route state; a reservation service knows whether the table is actually booked; a retailer knows real inventory and what’s in your cart. Strip those away and the global assistant is left inferring from residue.

The everything assistant’s last move is to fix that by capturing more — listen longer, watch more, remember everything. That trades a domain-authority problem for a surveillance problem and solves neither. The better move is to let the right service agent enter the user’s space with a scoped grant and answer as itself, with visible identity and bounded authority. Spotify shouldn’t have to hand one assistant everything; the assistant shouldn’t have to become Spotify. Where agency should live — in the model, the device, the service, or the user’s own room — is the contest behind all of this (Where Does Agency Live?).

A global assistant can observe. A native agent can understand. A forum can coordinate.

The fifth failure is the business model

There is a reason the biggest companies keep building the everything assistant despite all of this, and it is not that they haven't noticed. An assistant that sees everything accumulates the one asset that cannot be copied: months of behavioral context — how you work, what you ignore, what counts as urgent. That context has no export format. Leave, and you start over with what the analyst Nate B. Jones calls “a brilliant stranger.” The everything assistant looks like convenience and functions as lock-in: whoever holds the context holds the customer.

And the concentration cuts the other way too. Maximally useful means maximally informed, which means maximal blast radius when something goes wrong — a compromise, a subpoena, a quiet change of incentives. The user's own instinct is already partition: nobody wants their finance agent reading their private messages, or their work agent holding their medical history. The architecture should side with that instinct — a small polity of scoped agents, coordinated by a layer that knows a constraint exists without holding the data behind it. There may even be a market in the inversion: forgetful agents, valuable precisely because they can prove how little they retain.

A place you can hold to account

Put those four refusals together and you get a single positive standard. Context is not a blob of data an assistant can see. It is a relationship among actor, task, surface, permission, memory, and action — and you can hold it to account when the user can answer a few plain questions at any moment:

The everything assistant cannot answer those questions, because its entire value proposition is to dissolve them into one voice. A forum can — not because it is a group chat with bots, but because it is the governance layer that makes actor identity, data scope, sponsorship, authority, and return paths visible. That is the difference between multi-agent as a feature and multi-agent as a structure you can actually hold to account. Dissolving “who is speaking” into one voice is exactly the ventriloquism I take apart in Who Are You Talking To?

Which is why the lens I’ve carried since a 1995 usability study and the architecture I argue for now turn out to be one thing seen from two ends. The research was never a detour from the thesis; it was the thesis arriving from below. And what it kept finding is that useful AI systems are authority arrangements, not just intelligence: the question is never whether the one assistant is smart enough, but how context, state, execution, and authority are distributed. Compressed to a line: deictic input, durable rooms, and service-level agents are the real AI/UX substrate. Point at the thing; land in a room that remembers; let the services that can actually act show up as themselves. Everything else is an assistant asking you to trust a voice.

Essay — the spine. The synthesis where the research meets the architecture; companion to Against the Middleman. ← All essays