Dr. Dave Gilbert · 30 years of situated UX research

Once software starts acting for you, one question matters: who are you talking to?

Everyone agrees the conversation is becoming where computing happens. The consensus then collapses everything into it — every service speaking through one assistant's voice, every app squeezed into a widget in the chat window. I'm betting against both collapses: services present as themselves, accountable, in the conversation — and real websites beside the chat, coordinated with it through a shared grammar.

Five research probes are in the field now — Suminar, Mem·Sum, Mail·Sum, Draft·Sum, and Photo·Sum — built to learn, not to launch.

Today's answer is a ventriloquist — one voice performing many hidden principals. The better answer is a conversation where every voice has an owner you can name — one you can see, question, and hold to account.


The first bet

One assistant, or the conversation?

The consensus model hides a multi-agent system behind the scenes and puts a single assistant in front: the sub-agents do the work, the assistant speaks for them, and you only ever talk to it. The bet here is the opposite — that the agents worth trusting are the ones you can address yourself: service and brand agents present as themselves, accountable for what they say, in one conversation with you and with each other.

The Forum — the architecture, in full →

ONE ASSISTANTthe assistantyouone voice; no one else you can addressTHE CONVERSATIONyouhotelairlineyour assistant
The first bet, drawn: one voice in front of agents you can't address — or named agents in one conversation, talking to you and to each other.
The second bet

A widget in the chat, or the website beside it?

There is a second collapse under way, and a second bet against it — independent of the first, and at least as important. The assistants with real adoption — ChatGPT, Claude, Perplexity — are becoming where people meet software, and the industry's first answer is to squeeze the world's services into the chat itself: apps as iframes and widgets embedded in the message stream. My bet is that this will never be rich enough to satisfy. A chat message is a fine place for an answer and a terrible place for an application to live.

The labs are already gesturing at the better geometry: OpenAI and Anthropic now ship fully functional, Chromium-based browsers inside their desktop apps — the chat on one side, real websites in tabs beside it. The shell is right, and the seam has started to fill. In late August the first native lane opened: WebMCP, an experimental browser API drafted at the W3C by Google and Microsoft engineers, lets a page register tools an agent can call, and OpenAI now honors those “site tools” inside ChatGPT's built-in browser. A transport can hand the assistant a page's verbs. It says nothing about the conversation between the two surfaces: what was selected, what is staged, what was approved, what actually happened, and what the page may do to the chat. Co-location is not coordination, and neither is a tool list. That seam — the grammar, not the transport — is where my work concentrates.

What has to develop across it is an interaction grammar: two surfaces working the same live objects — the chat reasons and authors; the site shows, stages, and authorizes; typed acts keep the two in common ground. Most of the technical pieces already exist. And the grammar cannot assume a desktop app, because that is not where the users are: people meet their assistants overwhelmingly through the web versions of ChatGPT, Claude, and Perplexity, so what I'm developing depends on no rich client at all — a browser tab of chat, a real application beside it, an ordinary protocol between them. Web-first is not web-only: the family's newest probes run the same grammar inside the desktop clients, because that is where Word and Lightroom live. The probes into what this future UX needs to be began as the Companion sidebar and have grown into Mail·Sum's one-window app, reached through two doors at once; The Grammar of the Second Surface is the vocabulary, with its first increments already running.

A WIDGET IN THE CHATTHE CHATyou“book the flight?”the assistantthe airline appin an iframe…the thread scrolls onsqueezed into the message streamTHE WEBSITE BESIDE ITYOUR CHATyouyour assistantreasons · authorsthe service's sitethe pagethe live objectshows · authorizesorient · pointreferences backcommitview moves freely; commitment is authorized on the site
The second bet, drawn: the app squeezed into the message stream — or two surfaces on one live object, where view flows freely and commitment passes the gate on the site's edge.
TWO ENTRANCES, ONE KERNELYOUR CHATyouyour assistantreasons · authorsthe service's kernelrecords · receipts · authorizesTHE BROWSERthe open pageWebMCP · the live-page APIfrom inside the browser showing itMCP · the service APIfrom any host, anywheresame records, same gates; the page is an entrance, not a second brain
The two doors, drawn: ordinary MCP reaches the service's kernel from any host, anywhere; WebMCP reaches the tools the open page registers, only from inside the browser showing it. Same records, same gates — the page is an entrance, not a second brain.
The bets, in the field

Five probes into computing after the app

I didn't stop at the argument. The ·Sum family is a set of research probes — small working products, each built to test one dimension of what computing looks like once the app is displaced by the conversation you already live in. What you read: Suminar (live, in beta) is the first bet, running: the chat you already use becomes the forum, and each scholarly source joins it as a genuinely separate agent — its own memory, its own calls, its own context, its own voice — not one backend ventriloquizing many names. What you remember together: Mem·Sum (live, in beta) is one shared Sum — pages, people, evidence — for one to five people, reached through the assistants they already use. What you send: Mail·Sum (in private alpha) is a real email address feeding an owner-confirmed private Sum — already sending mail: the chat authors, the site authorizes. It is the first member to get the family's next surface: one installed window, a progressive web app on desktop and mobile, a rail beside a content pane — reached through two doors, an ordinary MCP connector from any assistant or WebMCP inside ChatGPT's built-in browser. What you write: Draft·Sum (in private alpha, Mac) is a co-editor inside Word — a thin add-in pane and a local kernel, no cloud component at all; every edit your assistant proposes lands as a native tracked change you accept or reject in Word. What you see: Photo·Sum (in private alpha) connects Lightroom Classic and Photoshop to your own assistant through the desktop apps' built-in browsers — again local, again no cloud: the assistant adjusts, the receipt records, Lightroom shows. And beside all five, the instrument planes — a one-window app, a pane inside Word, Lightroom itself — are the second bet's probes: the two-surface UX between chat and the place the work already lives, in the field. This site is called After the App for a reason — this is what after the app looks like.

The probes are the study — not a startup being launched, and free is the method, not a promotion. This is how I do UX research now: build the instrument, put it in people's hands, and learn from what they choose to share. The operator surface is content-blind by construction, so the beta learns only what users volunteer. The argument ships, and the shipping argues back (the case study).

The chat authors, the site authorizes. Mail·Sum's one-window app inside ChatGPT's built-in browser, the WebMCP door: the member points at the open message with “this” and asks for a staged reply — no handle, no subject line — and eighteen seconds later the app has turned to the draft, version one, the send gate unpressed. Click or tap to view full screen.
The local probes, where no cloud is involved. Left: Draft·Sum — “tighten this paragraph by about a third” resolves from the selection in Word, and twenty-seven seconds later the tightened text sits in the document as a native tracked change, accept or reject, with the pane keeping the record. Right: Photo·Sum — “this photo” is whatever is on the easel, and thirty seconds later Shadows sits at +12 and the white balance at 5200K in Lightroom's own panel, one named, undoable step. Click or tap either to view full screen; the surfaces' lineage is The Companion.

Three arenas where it plays out

Same shape everywhere — a conversation of agents you can address directly, not a single assistant. These are the three concrete places it lands.

The Individual

You, as yourself, meeting the world's brands and services in the open — across every surface you use, web and apps alike.

Enter →

The Home

AI as domestic infrastructure — household administration handled by a configured appliance, not a chat toy or a butler.

Enter →

The Workplace

Agents inside the organization: the forum architecture deployed at work, with visible roles and accountable handoffs.

Enter →

Start here

Three refusals, one argument — the triptych at the front of the spine.

The spine · Essay

Against the Everything Assistant

One voice that sees and does everything is the wrong dream. Point at the thing; land in a room that remembers; let named services act and tell the truth about it.

The spine · Essay

Against the Middleman: Why Agents Should Speak to You Directly

Why the future isn't a hidden swarm behind one assistant, but branded agents that join the conversation as themselves.

The spine · Essay

Who Are You Talking To?

The ventriloquist problem: when one assistant speaks for many hidden brands, you can no longer tell whose voice you are hearing — or who benefits.

All the essays, newest first, with reading paths →

The practice behind the essays

Thirty years of one question

None of this is a pundit's bet. These essays are the latest application of a research practice held steady since 1995 — from a conversation-analytic study of an early agent-based PDA, through museum galleries, retail aisles, and an aging-in-place startup, to instrumented research on today's AI. One lens throughout: how people make sense of, break down with, and repair their understanding of intelligent machines. The case studies carry the primary evidence — the 1995 transcripts, the research programs, the shipped decks. And lately the practice ships: the Sum family above is the same method, run as live probes.

Read the practice →

The vocabulary, mapped

A connected system of concepts

The work is the conceptual apparatus — named, interlinked, and usable. Every term is defined and linked to where it is argued.

Open the concept map →