PRISM: telemetry as a turn
The showpiece, and a collaboration. Patrick Squire conceived PRISM and built the physical lab and the SentiMeter hardware. I built the other half — the methodology, and the pipeline that turns a recorded session into analyzable data and hands it to AI agents to read. Together they make the lens of 1995 multi-track, and turn a participant’s own hand into evidence.
PRISM records a usability session as a set of parallel, time-aligned tracks: each person’s speech, the state of every visible screen, the study’s task progression, and — the move that changes everything — a physical feedback controller the participant operates while they work. Its canonical output is a single plain-text transcript, the PRISM Notation, in which every track is interleaved in time, so an analyst or a model can read any one channel alone or see exactly how the channels line up at any instant.
[00:04:00.0] speaker_moderator: introduces the effort knob
[00:04:15.3] knob_effort: 45
[00:04:22.1] speaker_participant: okay so that's like how much work it takes
(1.2)
[00:05:47.0] pad_return: ●
The SentiMeter, and one ontological decision
The SentiMeter is a tabletop MIDI controller — continuous-value knobs and momentary pads — sitting within reach of the participant’s non-dominant hand. Each study configures its own dimensions: a knob might track frustration, or “this works well,” or “this feels made for me.” The participant sets the knobs during reflection and taps a pad to flag a moment.
Everything hinges on one decision about what those knob turns are. They are not measurements of an internal state. They are non-verbal turns in the conversation — the same status the 1995 study gave a hand tracing a frozen progress bar in the air. A knob overshot and walked back, micro-adjusted, then held: that is breakdown and repair made visible as motion, often before the participant can put it into words. This is Hutchins in the lab — the knob is a material representation the participant and moderator jointly manipulate as part of the act of reflection, not an instrument that reads a feeling off of them.
When a participant turns a knob, it is not data about a feeling. It is an act of placement — analyzable exactly the way a gesture or a phrase is analyzable.
The gap between word and hand
That decision is what buys the lab its central finding. People in usability sessions are polite, and they blame themselves; someone learning an unfamiliar device expects friction and reads the system’s failures as their own fault. So the spoken track and the physical track come apart — and the gap between them is the data. Picture a participant who calls a jarring reset a “minor glitch” with an easy shrug while the frustration knob, under their own hand, climbs toward its ceiling. The words are the polite account; the knob is the breakdown the account is smoothing over. Standard testing hears only the words. PRISM keeps both, and treats the discrepancy as the thing most worth explaining.
Two disciplines keep this honest, and they are the same ones the 1995 study insisted on. The transcript is primary data — every analysis is a derivative document that cites it and never edits it. And categories emerge from analysis, not before it: the transcript does not pre-tag a stretch as “calibration” or “frustration.” Even the navigational markers are held to plain descriptive prose, and their authorship is logged — because a marker is the one place an earlier reader’s interpretive frame can leak into the record. That representational humility — the instrument shapes what it captures, so keep its seams visible — runs through all of this work.
Four things fall out of treating the controller as a turn rather than a rating. It breaks completion bias — a task can finish while the dials show the person no longer trusts, identifies with, or would tolerate the device. It separates what a single score blurs together — frustration, “works well,” “made for me,” and trust routinely move in opposite directions, and that divergence is often the finding. It turns a private feeling into a public artifact the participant can look at, puzzle over, and revise — which hands the moderator a concrete next question: why that number, why now, what changed? And it keeps stakeholder claims close to the evidence, because the product language comes out of the sequence rather than a pre-baked survey category.
The pilot taught us something we had not designed for: the dials' deepest product is justification talk. A knob constrains the participant into articulating a position — what would push the score up, what would push it down — and that articulation, not the number, is where the detail lives. The apparatus also keeps a deliberate asymmetry: pads are spontaneous, dials are prompted at sub-task boundaries — so when a participant reaches for a dial unprompted, the reach itself is a signal. The lab's success criterion follows from all of this. It is not whether a participant completes tasks efficiently. It is whether the session exposes the moments where a first-time user's mental model fails against the device — and what resources, or absences, shape their repair.
Why it’s the same lens
PRISM converges six traditions on purpose — conversation analysis, Suchman’s situated action, Hutchins’s distributed cognition, Goffman’s frames, Winograd and Flores on breakdown, Goodwin & Goodwin on gesture. These are the foundations the 1995 Magic Cap study stood on — kept not out of habit, but because three decades of fieldwork kept confirming them — and extended here with something the 1995 study lacked: a real-time feedback device as a new class of non-verbal participant. The 1995 paper concluded that usability is a function of the resources available for the repair of breakdown. PRISM is an instrument built to catch those moments of breakdown and repair as they happen, across more than one channel at once.
PRISM Atlas: the half the agents read
A multi-track session is only as good as what you can do with it, and that is the half I built. A pipeline ingests the recorded session and produces the canonical artifact — a single plain-text transcript in a modified conversation-analysis notation that time-aligns the spoken track (transcribed and speaker-separated), prose descriptions of what the participant is physically doing, the visible screen state, the study’s task progression, and the real-time SentiMeter readings lifted straight off the controller’s display. One readable file, every channel interleaved in time.
Then comes the part that closes the loop. A study — its sessions, materials, and notes — is packaged as a PRISM Atlas bundle and opened in an agentic environment: Google Antigravity, Claude Code, or OpenAI Codex. The analysis is done with the AI, in plain English, against the transcript. But it is governed analysis: the bundle ships with a written agent contract — an AGENTS.md file — that holds the agent to the same discipline the method demands. The session transcript is primary data and is never edited; no analyst taxonomy is imposed in advance; a knob turn is treated as participation, not a measurement; and categories are made to emerge from the data rather than declared before it. The one interpretive layer an agent adds — navigational markers — is held to descriptive prose and stamped with its own provenance.
That is this site’s whole argument, practiced on my own research. The analysis is distributed cognition across a human, a set of AI agents, and a structured record. The agent works under an inspectable contract with visible seams rather than as an opaque oracle. And the discipline that keeps a conversation-analytic study honest — honor the primary data, distrust the smooth summary, let the categories arrive last — turns out to be exactly what keeps an AI analyst honest too.
What it means for AI
The gap between word and hand is the trust problem for AI. A user will tell you an assistant is “fine” while their confidence quietly withdraws; they will accept an AI’s claim that it has done something — and only later, viscerally, register that it didn’t. Those are precisely the moments a polite transcript misses and a second, bodily channel catches. The design conclusion is the one I keep arriving at from every direction: build for repair, make the system’s state an inspectable contract rather than a confident surface, and never trust the smooth report over the evidence of what the person actually did.
Status: the lab has recorded its pilot corpus, and we are preparing to use it to test new product concepts. Patrick and I intend to write PRISM up together; this page is the methodology, not that paper.
Case study · 2026. PRISM — concept, physical lab, and SentiMeter hardware by Patrick Squire; methodology, analysis pipeline, and the PRISM Atlas agent framework by Dave Gilbert. ← The Practice