UX Research
The most expensive mistake in software isn't building something badly. It's building the right thing badly enough to fix, or building the wrong thing perfectly. UX research is the set of methods for reducing the second risk, and the striking thing about it is how cheap the high-value techniques are: five people, an hour each, a working prototype, and you'll find most of the severe problems in a flow.
The single most important principle is the say/do gap. People are unreliable narrators of their own behavior — not dishonest, but genuinely poor at predicting what they'll do, remembering what they did, and identifying why. "Would you use this?" gets you a polite yes from people who never will. "Show me how you did this last time" gets you the truth. Nearly every methodological rule in the field is a consequence of that one fact.
You don't need a research team to do this well. You need to talk to users regularly, watch rather than ask, and resist the urge to explain your interface while someone is failing to use it.
TL;DR
- Watch what people do, don't ask what they'd do. Stated preference and actual behavior diverge reliably.
- Five participants find roughly 85% of usability problems in a given flow — run more, smaller studies.
- Generative research (interviews, contextual inquiry) explores problems; evaluative (usability tests) validates solutions.
- Never lead. "What are you thinking?" not "Was that confusing?"
- Don't help. The silence while someone struggles is the finding.
- Qualitative tells you why; analytics tells you how many. Both, in that order.
- Test with a prototype, not a finished build — research is cheap before implementation and expensive after.
- Findings that don't change a decision were entertainment.
Quick Example
A usability test session structure — 45 minutes, and it works for almost anything:
The hardest discipline in that script is the ten seconds of silence. The instinct to rescue someone struggling with your interface is strong, and it destroys the finding — you'll never know whether they'd have found it, and they'll spend the rest of the session waiting for hints.
Core Concepts
Generative vs. evaluative
Teams overwhelmingly do evaluative research and skip generative, which is how you end up with a well-tested product nobody needed. Generative work is less comfortable — it can invalidate the roadmap — and it's where the leverage is.
The say/do gap
People answer hypotheticals with their aspirational self, remember frequency badly, and are reluctant to criticize something you made in front of them. The countermeasures are structural: ask about past behavior rather than future intent, observe rather than ask, and make it socially easy to be negative ("I didn't design this").
Usability testing
The highest value-per-hour method available, and the one most teams should do more of.
The five-participant figure comes from Nielsen's finding that usability problems follow a diminishing-returns curve: five people surface roughly 85% of the issues in a flow, and the sixth through fifteenth mostly re-find the same ones. The correct response is not "test with more people" but "run three studies of five instead of one study of fifteen" — you fix problems between rounds and find the next layer.
Interviews
Anchor every question to a specific past event. Concrete memories are far more accurate than generalizations, and they surface the workarounds, spreadsheets, and side channels that reveal what the product is actually missing.
The most productive habit is asking "why" several times past the point of comfort, and then being quiet. Most of the useful material in an interview arrives in the silence after someone thinks they've finished answering.
Analytics as a complement
Analytics is excellent at telling you that something is wrong and useless at telling you what. Session recordings sit between the two — useful for spotting rage clicks and dead ends, and lacking the "what were you thinking?" that makes a finding actionable. See Product Analytics.
Other methods worth knowing
Surveys are the most commonly misused: they're good for measuring the prevalence of a known thing and bad for discovery, because you can only ask about what you already thought of.
Synthesis
Raw sessions are not findings. The step between them is where most research value is lost.
Report severity with evidence: "4 of 5 participants could not find billing; two gave up entirely" is actionable in a way that "navigation could be improved" is not. Attach a clip if you have one — thirty seconds of someone struggling ends an argument that a slide deck cannot.
Best Practices
Do it regularly, not as a phase
A standing weekly slot with one or two users beats a big study once a quarter. Continuous contact means you catch problems while they're cheap and build genuine intuition about your users rather than a document about them.
Have the team watch
Engineers and PMs watching a session live learn more in an hour than they will from any report, and they stop arguing about what users want because they just saw it. This is the single highest-leverage practice in the whole topic — a recording nobody watches is a report nobody reads.
Test the earliest thing that's testable
Paper sketches, a Figma prototype, a clickable mock. Research on an unbuilt design changes a design; research on a shipped feature changes a backlog. The cost of acting on a finding rises by an order of magnitude at each stage.
Never lead the witness
"Was that confusing?" plants confusion. "What are you thinking?" doesn't. Similarly, don't explain how something works during a task — you can't un-teach it, and you've lost the finding.
Recruit people who resemble your users
Testing with colleagues finds layout bugs and nothing else, because they know the domain, the vocabulary, and the product. Five actual users beat fifty internal ones. Where recruiting is genuinely hard, a service (UserTesting, Respondent) or your own support and sales channels is worth the cost.
Record, with consent, and clip the moments
A 45-minute recording is unwatchable; a 30-second clip of someone hunting for a button is persuasive. Get explicit consent, store recordings according to your privacy obligations, and delete them on a schedule.
Ask about the last time, not the usual time
"Usually I check it every morning" is a self-description. "Last Tuesday I forgot until the client emailed" is data. Specific recent events are far more accurate than habitual summaries.
Write down what would change your mind
Before the study, state what result would cause you to change direction. This prevents the very common outcome where research confirms whatever the team already planned, because the findings that contradicted it were reinterpreted.
Common Mistakes
Asking about hypothetical future behavior
Helping during the task
Treating five participants as statistically insufficient
Only running evaluative research
Research that doesn't reach a decision
Recruiting the easiest people
FAQ
How many participants do I actually need?
Five per user segment for usability testing, which finds roughly 85% of problems in a flow. If you serve genuinely distinct segments — administrators and end users, say — that's five of each. For quantitative claims ("what percentage of users do X?") you need statistical sampling, which is a different method with different requirements. The most common error is spending on a large qualitative study instead of running three small ones.
We have no budget or researcher. What's the minimum?
Watch five people use your product for 30 minutes each, once a month. Recruit from your existing users via email or in-app prompt, offer a modest incentive, use a video call with screen sharing, and follow the script above. This costs a few hours and finds more than any amount of internal debate. Everything else in this topic is a refinement on that.
Moderated or unmoderated testing?
Moderated when you need to understand why — you can probe, follow unexpected threads, and ask what someone expected. Unmoderated when you need speed, volume, or geographic spread, and you already know what you're measuring. A common effective pattern is unmoderated for breadth on a specific task, then moderated follow-ups on the surprises.
How do I convince my team to act on findings?
Have them watch. A clip of a real user failing at a task settles arguments that a written finding won't, because it removes the possibility of imagining a more competent user. Beyond that: rank findings by severity, tie them to metrics the team already cares about, and pick one clear win first so the process demonstrates value.
Does research slow down shipping?
It changes what you ship, and the net effect is almost always faster. A week of research that prevents a quarter of building the wrong thing is not a delay. What genuinely does slow teams down is research theater — long studies producing reports that don't map to decisions. Keep studies small, frequent, and tied to a specific pending decision.
Can AI replace user research?
It can accelerate the mechanical parts — transcription, first-pass clustering of notes, drafting discussion guides, and summarizing sessions. What it cannot do is be your user. Synthetic-user tools that simulate participants produce plausible responses ungrounded in any real person's context, constraints, or workarounds — which is precisely the information you're doing research to obtain. Use it on the artifacts, not as the source.
Related Topics
- Prototyping — What you put in front of participants
- Design Systems — Where findings become reusable patterns
- A/B Testing — Quantitative validation at scale
- Product Analytics — Finding where to look
- Conversion Optimization — Acting on funnel findings
- Accessibility — Testing with users who use assistive technology
- Figma for Developers — Where design decisions get recorded