Applied AI: the same lens, today
I’ve done contract UX research at Google, on and off since 2018 — most recently on AI-mediated experiences across surfaces like Gemini, Pixel, and Chrome. The same situated, interactional lens, turned on the place it now matters most. I author and present the research; the internal particulars stay internal, so what follows is the method and the stance.
The work has two faces. I author and present the formal research — the reports and readouts that turn ambiguous AI-interaction evidence into decisions a product team can act on. And I stay embedded: in the room with PMs, designers, and researchers, participating in reviews and helping steer products through ambiguity rather than handing down a verdict from outside. Think of it as an advocate for the user’s situated reality at the table where the decisions get made. The shape of the work, without the internal particulars: dozens of studies conducted or managed since 2018, about half a dozen of them on AI-mediated experiences.
Before the AI turn: Meet for Telehealth
The Google work did not begin with AI. In April 2020, with COVID closing patients' homes to clinicians, I studied real telemedicine onboarding sessions — April 7 through 27 — with Encore Healthcare, the post-acute respiratory care provider whose platform story runs through this site's Evermind arc. The question was brutally practical: Google Meet nominally “does telehealth,” so why was onboarding consuming clinician time and losing patients?
The findings lived in the mismatch between nominal capability and the real workflow surface. Patients on phones could not find Meet invitations inside the Meet app. Neither side held the implicit knowledge of how the suite's parts fit together. And the two roles needed opposite documents: the clinician had to schedule, invite, script, support, fall back to the phone bridge, and keep every session logged for auditing and billing — while the patient needed only the shortest possible path to a working call. So the design output was role-separated onboarding, with the phone bridge doing double duty as the fallback and the auditable channel. That deck is public: Meet for Telehealth: User Roles and Onboarding (PDF).
Nothing in it is AI. Everything in it is the lens: the finding is never in the feature list; it is in the gap between what the product nominally does and what the situated workflow — accounts, invitations, literacy, clinician time, billing — actually requires. Five years later the same provider's platform carries an AI layer I designed. The method that got there started with watching onboarding calls during a pandemic.
Slowing the scene down
The recurring move is simple to state and hard to do. A team sees a clean surface: the demo works, the satisfaction number isn’t catastrophic, the AI’s answer sounds fluent. But a satisfaction score that looks fine routinely hides three different stories — people who never found the feature, people rating a different feature, and people who succeeded yet wanted something the feature doesn’t do — and the work is to pull them apart before anyone trusts the number. So I slow the scene down and ask what the user actually understood — what context they believed was in scope, whether the interface made it legible who was acting, and, crucially, what the team can responsibly claim from the evidence in front of them. The drama is always the same: the gap between a fluent AI surface and a situated person who still has to decide whether to trust it.
Every commitment from the rest of this arc shows up here, unchanged. The surface lies: a satisfaction score or a smooth interaction can hide confusion, mistrust, or an unearned belief that the system did the right thing — the same gap between word and reaction that PRISM was built to catch. Breakdown is where the truth is: the strongest evidence comes when a user accepts a confident surface but has misread what happened or where control sits. And the hardest failures are frame failures — the user genuinely cannot tell whether they’re using a local tool, a personal assistant, an information layer, or an autonomous actor, which is the 1995 “scenes to screens” problem wearing a 2026 outfit.
The interesting evidence is in the mismatch between surface fluency and situated action.
Four questions for any AI feature
Under the method sits a short checklist I bring to any AI-mediated experience — four questions, each a different way the thing can fail. Is it legible (can the user tell who is acting and what is in scope)? Valuable (does it do real work the user already wanted done)? Trustworthy (does the surface earn the confidence it asks for, or merely assert it)? And contestable (when it’s wrong, can the user see that and push back)? Most AI-UX failures are a failure of one of the four — and they are failures of frame, evidence, trust, and recourse, not of model quality.
Where the arc points
This is the case study the rest of the site grows out of. The essays here are not speculation laid over a research career; they are this research career, pointed forward. Designing AI for repair rather than one-shot intent capture, treating the interface as an inspectable contract, locating intelligence in the coordination rather than the model, and noticing why ambient AI keeps losing — each is the same lens, in a present-tense room, with real products and real users on the other side of the glass.
Case study · contract UX research at Google, on and off since 2018. Public products named; internal specifics kept internal. ← The Practice