Interaction design · Speculative research
UC Berkeley Master's thesis · 2024–2025
Designing an AI-led experience across voice, touch, and body
What makes us human? Future CAPTCHA began as an authentication problem and became a research-led interactive test about the qualities people use to describe being human. I designed the research, interaction choreography, voice behavior, interface, and working iOS prototype, then observed the experience with public visitors at exhibition scale.
Before I explain it, experience it.
About three minutes on the exhibition floor: the evaluator speaks, a visitor answers out loud, moves, and receives a score. Sound on, the voice is half the interface.
Recorded at the thesis exhibition, UC Berkeley, December 2024.The tester's voice and their taps happen off-frame, so what you are watching is the system's side of the conversation: the prompts, the listening state, the transcript coming back, and the score it settles on.
I started with authentication. I ended up questioning the thing we were trying to authenticate.
My thesis began with telephone fraud. I spent months on it, mapping who gets targeted, why shame keeps victims silent, and where an intervention could sit. Then my attention shifted. The version of the problem that kept pulling at me was no longer a stranger on the phone but a synthesized voice that sounded like someone you love.
Where the question came from



Each detector eventually loses, which left me with a question underneath the whole effort.
What exactly are we trying to prove when we prove that someone is human?
That question moved the thesis from a real-world problem to a philosophical one: voice scam, to authentication, to humanness. Seven years of design education and practice had trained me to look for the practical solution first. For a master's thesis I wanted the opposite, the bigger and slower question that shapes how designers work with emerging technology.


What makes us human? Whose definition of human?
A quick interview would have forced people to answer an abstract question on demand. I wanted three kinds of evidence: what people notice in their own lives, what a broader community intuitively associates with humanity, and how different disciplines frame the human-machine boundary.
Participants documented moments when they felt deeply human, with context, action, and feeling. The mundane answers became the most revealing: morning coffee, guacamole, being overwhelmed, small rituals, procrastination, relationships.
14-person diary study
Three days to a week, so the thought could develop inside ordinary life instead of on the spot.
Campus-wide survey
I moved the question rather than assuming one community represented Berkeley: Soda Hall, Sutardja Dai, Life Sciences, Moffitt.
Interviews + literature
Social science researchers, ML engineers, and a review of how media frames the boundary.
01 / 07

Research findings became interaction requirements.
I resisted forcing the research into a neat definition too early. Across diaries, interviews, survey responses, and media research, recurring clusters formed around five qualities, and each one had to become something a person actually does in the room.
A screen full of survey questions would have contradicted all of it. If embodiment mattered, participants needed to move. If emotional complexity mattered, fixed-choice responses were not enough. If unpredictability mattered, the system needed to observe behavior rather than ask people whether they were unpredictable.
What makes us human?






















Multimodal design became attention design.
Before building the prototype, I physically mapped the sequence of prompts, inputs, measurements, feedback, and final reflection. I was designing the rhythm of attention, not just individual screens: listen, respond, move, receive feedback, continue.
Not every human response should use the same input. The screen is the stimulus and touch is the answer when reaction should happen before explanation. Voice carries expression when the response is a value judgment. The body becomes the input when the quality being tested is embodiment, and once movement began, voice mattered more than text, because looking back at instructions would break the interaction.
Visual affinity
Screen is the stimulus, touch is the answer, voice is the acknowledgment.
Value response
Screen holds context; the visual layer only reports that the system is listening.
Embodiment
Continuous guidance moves into voice so nobody has to look back at the screen.


The evaluator needed authority without becoming inhuman.
I moved away from the bright, friendly, often feminine assistant archetype. The premise was closer to an airport checkpoint: the system was evaluating you. I worked through Google Cloud's voice catalog, listening to hundreds of them, before choosing a lower-pitched male voice that balanced authority, warmth, dry humor, and audibility in a noisy room. Casting was only half of it: the same voice at default settings still sounded like an assistant, so I slowed the delivery and dropped the pitch until it read as an evaluator.
01 · Casting the voice
Warm and eager. It made the test feel like a product feature instead of an evaluation, and it disappeared under room noise.
Lower pitch, slower rate, flat delivery. Closer to an airport checkpoint than an assistant, and it still carried across a crowded room.
Waveforms are illustrative. Attach the two exported clips to turn this into an A/B listen.

02 · Writing the line so it can be heard
You have a name as an individual. Are you a real human? This is a mandatory test identifying humans among machines.
Okay okay, [600ms] well...... You have a name as an individual. [700ms] So.. Jade, are you uh... real human? [800ms] This is a mandatory test identifying humans among machines, or, something in the middle.
Pauses and fillers live in the script itself, so every visitor hears the same performance.
03 · Who is holding the turn
Voice and text finish together. Listening opens after a deliberate gap, never on the final syllable.
Silence after the prompt
Nothing on screen while the microphone was live. People asked out loud whether it was listening, then answered too quietly to register.
The transcript is the feedback
The recognizer streams its transcript and the screen renders it as it arrives. People corrected themselves without being told to: they saw a garbled line, stepped closer, said it again.
A liked photograph gets "Nice..", a liked AI image gets "Interesting choice..", a rejected AI image gets "Okay," and a rejected photograph gets "dislike... noted." Neutral returns nothing at all.
Their job is character, not information. That asymmetry made the system feel as if it were watching rather than recording, and a flat delivery kept it from steering the next answer.
I wanted to read emotion from the voice itself. That failed.
The original plan sent raw voice to Hume AI, a vocal emotion model. In exhibition conditions it was unreliable, so the working system transcribes speech on device and interprets the transcript semantically instead.
This was not an equivalent replacement. Vocal emotion asks how the person sounded; semantic emotion asks what the person expressed. I kept that limitation visible in the framing rather than pretending the system could measure something it no longer could. I changed the signal, not the interaction intent.
SFSpeechRecognizer / AVFoundation
GPT-4, context and semantic emotion
Google Cloud TTS, scripted with SSML pauses
Vision Framework / ARKit
The final interface had to work without me explaining it.
This was not a lab prototype. An academic committee and hundreds of visitors encountered it while other people were talking, waiting, watching, and moving around them. No formal onboarding, ambient noise, a social cost to speaking in public, shifting attention, and spectators who made the interaction partly performative.
So the exhibition became a usability study in the wild. I stood to the side and watched, and these are the things I could not have learned in a quiet room.
Eight things I watched happen
- 01Spoke over the prompt, or waited too longnothing marked when listening began
- 02Missed what the evaluator saidaudio alone does not survive a loud room
- 03Asked a neighbor to repeat itlong prompts need a second pass
- 04Assumed it was brokensilence during processing read as failure
- 05Mumbled with an audience behind themquiet answers under-registered
- 06Turned back to read instructionsreading stopped the movement being measured
- 07Waited for someone else to go firstno proof the tracking was live
- 08Compared results out loudthe argument was the point, so it stayed
Believable enough to engage. Questionable enough to discuss.
The goal was never an accurate measurement of humanity. I considered the experience successful when participants stopped treating the machine as an authority and started debating the assumptions behind its judgment: why did it think that, could I fool it, would another person get the same result, and why do we need to prove that we are human at all.


Recognition
A web version, rebuilt so you can try it here.
This is where those observations landed. The exhibited build used a tablet, a cast voice, capacitive touch, and head tracking, and the browser has none of that, so the port trades fidelity for access: the emotion reading is coarser, a pulse stands in for capacitive sensing, a cursor for head tracking, and your device speaks with whatever voice it has.
What it does carry is the script, the images, the scoring, and every interface change the exhibition argued for. Sound on. Five steps, about three minutes.
The evaluator's line appears on screen as it is said, so a missed word is never a lost instruction.
An explicit prompting, listening, captured, processing sequence, and the microphone opens after the prompt rather than on its last syllable.
The transcript streams while you speak, so you can see what was heard and say it again if it came out wrong.
One thing to do at a time instead of a paragraph to retain.
The one thing I did not fix.
Deliberate. The premise is an airport checkpoint, a test administered by someone standing in front of you, and you do not get to ask an officer to say it again from the top. A replay button would have made it a media player instead.




The test was never going to measure humanity accurately.It was built to open up the discussion, and the lineeach person drew between human and machinerevealed their own definition, not the system's accuracy.
