Part 1
Where AI adoption breaks down
AI products fail because people stop using them.
That usually happens for three reasons.
1. They don’t know when to trust it
So they check every answer, and every check adds friction. Eventually, using the AI takes as much time as doing the work themselves, and they stop relying on it.
2. It doesn’t fit the way they work
The output is correct, but they still have to reformat it, copy it into another tool, or finish part of the task manually. Whatever time the AI saved, the workflow gave back.
3. It got something wrong early, and they never gave it another chance
People form opinions about AI quickly. A few bad experiences are often enough to lose their trust. Traditional software usually gets a second chance after a rough release. AI rarely does.
All three failures have the same root cause: product teams tend to evaluate whether the AI works, but users evaluate whether it’s worth using.
At Akraya we call these the Three Pillars of AI Adoption Risk. Together, they explain why so many AI products that work technically never become part of people’s everyday workflow.
In this guide, we’ll explain how UX research for AI products helps uncover these risks before launch, which research methods answer which questions, how researchers measure trust, and when it makes more sense to build an in-house research team versus bringing in a UX research partner.
Our framework
The three things that determine whether people use AI
Most AI adoption problems can be traced back to three questions. The important part is that all three can be identified in UX research long before they appear in your product metrics.
We developed the framework at Akraya from research inside hi-tech, software product, and internet companies, and we call it the Three Pillars of AI Adoption Risk:
- Can people depend on it?
- Does it fit the way they already work?
- What happens after it gets something wrong?
Each question uncovers a different kind of adoption risk, and each leaves clear signals that UX researchers can observe well before users start abandoning the product.
1. Functional reliability: can people depend on it?
People only rely on AI when they believe it’s dependable. That’s harder than with traditional software because the same prompt can produce different answers, and an answer that looks right can still contain subtle mistakes.
How unreliability shows up
People check the AI’s work more than they would check their own.
- Copying answers into another tool to verify them.
- Running the old manual process alongside the AI.
- Hesitating before acting on a recommendation.
Why reliability matters
Verification is measurable. Researchers can track how often people check the AI before they trust it. If that number falls over time, trust is growing. If it rises after mistakes, trust is eroding.
2. Workflow fit: does it fit the way people already work?
An AI can produce the right answer and still lose the user. What decides adoption is how much work it takes to get that answer into the job. With traditional software, a correct output is usually the end of the task. With AI, it’s often the middle of one. On a device, the same problem shows up physically: an answer that arrives while someone is running, driving, or wearing gloves has to fit that moment or it goes unused.
How poor workflow fit shows up
People try the AI once or twice, then go back to their old process.
- Reformatting the AI’s output before it can be used.
- Switching to another tool to finish the task.
- Completing part of the work manually anyway.
- Using the feature less each week without ever complaining about it.
Why workflow fit matters
This risk never appears in an accuracy benchmark. Researchers can only find it by observing people in their own environment, working on their own tasks, which is why workflow problems tend to surface after launch instead of before it.
3. User trust: what happens after it gets something wrong?
Every AI makes mistakes. What decides adoption is what people do after the first one. Traditional software usually gets a second chance after a rough release, but AI often doesn’t, because one wrong answer can change how someone judges every answer that follows.
How lost trust shows up
Behavior changes noticeably after a single error.
- Verifying every answer after one bad experience.
- Narrowing use to low-stakes tasks only.
- Saying they wouldn’t use the tool for real work.
- Abandoning the feature without reporting a problem.
Why user trust matters
Trust breaks in a single moment and rebuilds over weeks. A one-hour session can’t capture either. Researchers have to follow the same people across repeated use, before and after the AI gets something wrong.
Akraya’s UXR Decision-Confidence Framework
Each risk needs a different method, so we scope the work around whichever pillar is most exposed rather than running all three at once. The framework is how we package the result, and it is written for the people who have to act on it rather than for the people who commissioned it.
Running these questions before launch is what makes them worth asking. At that point every one of the risks is still fixable. After launch, all three cost significantly more to address.
Self-check
Signs your AI product needs UX research
Most product teams don’t need convincing that UX research matters. The harder question is whether their product needs it right now.
Usage spiked at launch, then flattened
The launch numbers looked strong, but engagement dropped off within a few weeks. Usually that means people tried the feature, decided it wasn’t worth the trouble, and went back to the process they already trusted.
The technical metrics are healthy but adoption isn’t moving
Accuracy, latency, and response quality are all where they should be, but usage still isn’t growing. When the model is performing and the product isn’t, the limiting factor sits somewhere your model metrics can’t reach.
Nobody in the room can explain why people stopped
Analytics show you where the drop-off happened, but they don’t show you the moment someone decided the AI wasn’t worth relying on. When roadmap discussions start running on competing theories, you have a research question rather than a disagreement.
Users are going quiet instead of complaining
Support tickets and feedback stay low while usage declines. Silence is easy to read as satisfaction, but more often it means people stopped expecting the feature to get better. Abandonment rarely generates a ticket.
Product decisions are outrunning your capacity to answer them
Research requests keep arriving, the backlog keeps growing, and product managers have started making AI decisions without waiting for evidence. At that point the bottleneck is throughput rather than research quality.
If one of these signs describes your product, there is usually a single study that answers it. If three or four describe it, the problem is wider than one study can close, and what the product needs is a research plan rather than a project.
Part 2
What is UX research for AI products?
UX research for AI products studies whether people can rely on an AI system whose behavior changes over time, and what has to be true for them to keep relying on it.
Answering that means understanding four things at once:
- The person. What they’re trying to accomplish, what they already know, and how much a mistake would cost them.
- The AI. What it’s good at, where it struggles, how consistent its answers are, and where it runs — a model on a wearable behaves differently from one in the cloud.
- The interface. How the AI presents an answer, what people can change, and how easily they can correct it.
- The workflow. Where the AI sits in the larger process, who reviews its work, what happens when something goes wrong, and who owns the outcome.
The workflow is the one teams overlook. Most AI features that fail have very little wrong with the model, but the output doesn’t fit the way people work. Sometimes an answer is useful and takes too long to verify. Sometimes it’s technically correct and still can’t be used in the next step.
That connection is what makes a finding usable. “Users don’t trust the summaries” gives a product team nothing to build against.
Compare it with this: “Analysts stop trusting summaries when conflicting evidence is left out. Every summary should surface contradictory signals, and we can track how often known conflicts appear in the output.” Product and engineering can design, build, and test against that.
What makes it different
Why AI products need different UX research methods
Traditional UX research was built for deterministic software, where the same input produces the same output every time. AI is probabilistic, so the same prompt can return different answers, the path through a task rarely runs straight, and people can’t see the logic behind what they’re given.
That changes what UX researchers need to measure.
It’s no longer enough to ask whether people can complete a task. Researchers also need to understand whether people trust the AI, how they respond when it makes mistakes, and how much uncertainty they’re willing to tolerate before they stop using it.
That’s why UX research for AI products goes beyond traditional usability testing or model evaluation. A product can score well on both and still fail because people never make it part of their daily work.
The difference comes down to five key areas:
1. The same action can produce different results
In a normal usability test, two people doing the same task see the same screen, but with AI they may not.
When one participant has a smooth session and the next one struggles, researchers have to rule out the model before they look at the design. The same prompt may simply have produced two different answers.
That’s why every session has to capture exactly what the model produced, along with the model version, settings, and timing. Without that record, there’s no way to explain why two people had different experiences.
2. People invent their own rules for how AI works
When people can’t see how a system works, they fill in the blanks themselves.
Some assume the AI is searching the web. Some think it remembers every conversation they’ve had. Some think it already knows their company’s policies. Those assumptions are usually wrong, but they still drive how people use the product.
Researchers call these mental models, and surfacing them explains misuse and mistrust that looks inexplicable from the outside.
3. Trust is easier to observe than to ask about
Ask someone whether they trust an AI and you’ll get a thoughtful answer, but it won’t tell you what they actually do at their desk.
Researchers watch instead of asking:
- Do people accept the good recommendations?
- Do they catch the bad ones?
- Do they verify every answer regardless?
- Do they come back after the AI gets something wrong?
Those behaviors say more than any survey question.
4. People want to know where an answer came from
Most people don’t need to follow the model’s reasoning. They need enough context to judge whether an answer is believable.
Knowing a recommendation came from company policy, a customer record, or a résumé is more useful than a confidence score, because a source can be checked and a percentage gives someone nothing to act on.
Researchers call this provenance, and it has a significant effect on whether people trust an answer.
5. The product keeps changing
Models get updated, prompts get rewritten, and retrieval improves. On hardware, a firmware release can change an on-device model in the field, on devices people already own. AI products rarely sit still.
Findings expire faster than they do for traditional software, which changes how studies get built. Researchers date findings to a specific model version, and they design studies that can be re-run rather than written up once.
All five differences point back to the same limit.
That’s why human research matters more as models improve rather than less. Adoption comes down to how people react to the system, and performance data doesn’t describe that.
Timing
Why fixing adoption after launch is so expensive
The same finding might take a design change before launch — or three months of engineering after it.
The cost of fixing the product goes up
Questions that are inexpensive to answer during design become expensive to fix once customers are already using the product.
Take a simple question like:
Before launch, you can answer it with a prototype and a handful of research sessions.
After launch, the same finding can mean redesigning screens, rewriting workflows, updating documentation, retraining users, and helping existing customers adapt to the change.
The question didn’t become harder. The cost of acting on the answer did.
You only get one first impression
Launch is the one moment when the largest number of people are willing to try your product.
If those early users decide the AI isn’t trustworthy or isn’t worth the effort, many of them never come back — even if the product improves dramatically a few months later.
The AI keeps getting better. The people who would have benefited most have already stopped paying attention.
The three adoption failures we see most often
Across AI products, we see the same three adoption failures repeatedly.
Ghost Tools
Features that work technically but solve problems users never cared about, or fail on the tasks they actually wanted help with. The result is significant engineering effort and little real adoption.
The Trust Gap
A capable AI that people never fully rely on because they don’t know when to trust it and when to verify its answers.
The Sunk Cost Trap
Continuing to invest in a poorly adopted feature because too much time and money have already been spent to reconsider the original assumptions.
A week of discovery research before development can reveal these problems while they’re still inexpensive to fix. After launch, the same findings often require months of engineering effort and a long process of rebuilding user trust.
Types of research
Which type of UX research does your AI product need?
Nine types of UX research cover most of what AI products need. Which of them you need right now depends on where the product sits and which adoption risk is already showing.
Almost nobody needs all nine at once. They sort cleanly by lifecycle stage, and the stage indicates which question is live at that point in the product rather than how often research should run.
Before you build
Before anyone writes code, the question is whether AI belongs in the experience at all.
01Discovery research
Discovery research asks whether AI belongs in this experience at all.
Researchers map how the work gets done today, where people struggle, and which parts of the task still need human judgment. A common method is a Wizard of Oz study, where participants believe they’re using AI while a researcher writes the responses behind the scenes, which makes it possible to study the experience before the technology exists.
These studies work best when some responses are wrong, slow, or outside what the AI can do, because that is where people reveal how they would treat the real thing.
The most useful finding is often a negative one: that the workflow doesn’t have the problem the AI was meant to solve.
If you skip it
You risk building a good solution to a problem the workflow doesn’t have.
While you’re designing and building
Once the product takes shape, the questions turn to whether people can use it and whether they’ll trust it. This is the last stage where the answers are still cheap to act on.
02AI usability testing
AI usability testing asks whether people can complete real tasks with the product, and where they get stuck.
Because the same prompt can return different answers, every session records exactly what the model produced alongside what the participant did. Without that record, two people having different experiences is unexplainable.
Sessions have to cover real input rather than the easy cases: common requests, hard ones, ambiguous ones, and the ones that cause the most damage if the AI gets them wrong.
What usually surfaces has less to do with the interface than with comprehension, and whether people understood the answer well enough to act on it.
If you skip it
Users find the avoidable problems after launch instead of before it.
03Trust and explainability research
Trust and explainability research asks a simple question: do people know when they can trust the AI?
Researchers watch when people accept the AI’s answer, when they stop to verify it, and whether the explanations help them decide what to do.
They also study what happens after the AI makes a mistake. Do people keep using it, or do they lose confidence and start checking every answer?
These questions matter because people don’t adopt AI just because it’s accurate. They adopt it when they know when they can rely on it. Trust and explainability research shows where that confidence breaks down before it becomes an adoption problem.
If you skip it
Adoption slows and the cause stays invisible.
04Human-AI interaction research
Human-AI interaction research asks where the work should divide between the person and the AI.
Researchers study what the AI handles on its own, when it should stop and ask, and which decisions people expect to keep. The answers vary more than teams expect, because tolerance for automation depends on what a mistake would cost the person making it.
Agents raise the stakes. An agent interprets intent and acts across systems with little visibility into how it decided, so the research shifts from validating an interface to validating decisions and outcomes: did it read the intent correctly, did it hold up across a multi-step task, and could the person see what it did and stop it.
Gartner expects more than 40% of agentic AI projects to be cancelled by 2027, and attributes most of that to users not trusting or understanding the systems rather than to the technology failing. More in UX Research for AI Agents.
If you skip it
People either lean on the AI too heavily or avoid it altogether.
05Accessibility and human factors research
Accessibility research asks whether the product works for people who use it differently than the team imagined.
AI introduces barriers traditional software doesn’t have, particularly in conversational interfaces, dynamic content, and generated responses. The work spans screen readers, voice interfaces, cognitive accessibility, and assistive technology.
On hardware the same question becomes physical: whether a device fits the range of bodies expected to use it. That is human factors research, and it runs on anatomical data rather than assumptions about an average user.
If you skip it
Some people can’t use the product at all, and the problems get far more expensive to address later.
After you launch
Launching tells you the feature exists. This stage establishes whether anyone built it into their work.
06Adoption research
Adoption research asks what people actually did after launch, and why.
Researchers look at where usage drops, what brings people back, and what sustains long-term use, measured against activation, retention, and feature utilization.
The gap it closes is between knowing that usage fell and knowing what caused it. Analytics show the first. Research explains the second.
If you skip it
Usage plateaus and nobody can explain it.
07Benchmarking and measurement
Benchmarking asks how the experience compares to where it was, and whether it is improving.
Trust, confidence, and reliance can’t be measured by asking about them directly, so researchers turn them into behaviors that can be observed and counted: how often people verify the AI’s work, how often they accept its recommendations, and whether they return after a mistake.
This work has to happen before launch. Without a baseline there is no way to tell whether a release moved trust, damaged it, or changed nothing.
If you skip it
Conversations about progress run on opinion.
08Mixed-methods research
Mixed-methods research asks both halves of the question at once: what people did, and why they did it.
Automated evaluations and product analytics answer the first. Interviews, usability sessions, surveys, and observation answer the second. Run separately, each produces findings the other can’t explain.
Combined, they connect a number to a reason, which is usually what a product team needs before it can act on either.
If you skip it
You learn what happened or why it happened, rarely both.
Throughout the product’s life
Models improve, prompts change, data sources get added, and user behavior shifts, so research has to move with the product.
09Research operations
Research operations asks how a team keeps research running without starting over each time.
It covers participant recruitment, research repositories, insight management, governance, and repeatable workflows. For products with several AI surfaces it also covers how all of them get reviewed the same way, so results can be compared across them.
The less visible part is what happens to findings. Most AI problems don’t belong to a single team — one issue can involve the model, the interface, the product copy, and an internal policy — so research operations also determines whether a finding reaches whoever can act on it.
If you skip it
Every study starts from scratch and useful findings stay in reports.
Choosing the right research matters as much as doing the research
Knowing the different types of UX research is the easy part. The harder part is choosing the one that answers the question you’re actually trying to solve.
Take a common problem:
That could mean people don’t trust the AI. It could mean the AI doesn’t fit into their workflow. Or it could mean users were never taught when to trust the AI and when to double-check it.
Those are three different problems, and each requires a different research study. If you investigate the wrong one, you can spend months fixing the wrong problem.
That’s why every engagement starts with scoping. Before deciding which research to run, we answer a few simple questions:
- What decision are you trying to make?
- What do you already know?
- Which of the Three Pillars of AI Adoption Risk do the current signals point to?
- Can one focused study answer the question, or do you need an ongoing research program?
There’s a practical reason this matters. If a research firm only offers usability testing, most problems will look like usability problems.
Akraya offers the full range of UX research services for AI products, so we can recommend the research that best answers your question rather than the research we happen to sell.
Cadence
UX research for AI products is continuous
Launching an AI product changes the research questions rather than ending them.
Before launch, UX research asks whether people will adopt the AI. After launch, it asks whether they still are.
Unlike traditional software, AI products don’t reach a point where the user experience stays relatively stable. The product evolves, the AI evolves, and the people using it evolve too. Each major release, model update, or workflow change can affect how much people trust the AI and whether they continue making it part of their daily work.
That’s why adoption has to be measured continuously rather than once.
In practice, an ongoing UX research program includes:
- Measuring how each major release affects trust and adoption.
- Studying the live product to understand where people stop using it and why.
- Following the same users over time to see how confidence changes with continued use.
- Re-testing significant changes instead of assuming a new model, prompt, or feature automatically improves the experience.
Research operations make this practical. Participant panels, research repositories, and repeatable workflows allow each study to build on the last instead of starting from scratch.
The point is making sure the product people are using today is still the product they want to keep using tomorrow.
Part 3
How UX research for AI products is actually done
Researching AI products builds on traditional UX research methods, but applying those methods to AI requires a different process.
Every study follows the same three stages.
1. Designing the study
Most of a study’s value is determined before the first participant ever uses the product.
Researchers define the real-world scenarios to test, design tasks that reflect how people actually interact with AI rather than idealized prompts, identify the behaviors that matter most, and establish how success will be measured.
This is also where the biggest methodological decisions are made, including how trust will be evaluated, how AI errors will be introduced and observed, and how the study will distinguish genuine adoption problems from normal AI behavior.
2. Running the study
Once participants begin using the product, the focus shifts from asking questions to collecting evidence.
Researchers observe how people actually interact with the AI: when they trust it, when they verify its answers, where they hesitate, what they edit, and what changes after the AI makes a mistake.
Alongside participant behavior, the AI itself is documented so every finding can be tied back to the exact experience each participant had.
3. Validating the findings
Before the results can inform product decisions, researchers need confidence that the findings will hold up.
That means accounting for variables such as model versions, prompts, retrieval systems, data sources, and participant experience levels, so that differences between participants reflect how people actually behave rather than differences in the AI itself.
The goal is findings that remain reliable after the study ends rather than observations that only applied to one version of the product.
Each of these stages involves dozens of methodological decisions that shape the quality of the research. The method-by-method version of all of it is in How to Conduct UX Research for AI Interfaces.
Success
How the success of an AI product is measured
UX researchers check three things to establish whether an AI product is succeeding, and a product can pass one while failing the others.
- Adoption that survives past the first week. Launch usage means people were curious. Usage in week six means the product earned a place in their work.
- Calibrated trust. People should rely on the AI where it earns it and question it where it doesn’t. Measuring that is involved enough that it needs its own approach, covered below.
- Workflow integration that removes steps. A successful AI product shortens the path between a task and its outcome. If people are reformatting output, hopping between tools, or redoing steps by hand, the product is adding work whatever the model metrics say.
All three need a baseline to mean anything, which is the case for benchmarking before launch. Without knowing how people worked before the AI arrived, there’s no way to show leadership the product is winning, or to prove the next release improved it.
These also map to numbers leadership already watches. Time saved per task, shorter cycle times, and fewer abandoned workflows are what adoption and calibrated trust look like on a business dashboard.
Of the three, trust is the one that resists measurement, which is why it has a set of methods built specifically for it.
Measurement
How UX researchers measure trust in AI products
Trust is measured by watching what people do rather than by asking whether they trust the AI.
Asking directly tells you surprisingly little. Two people can both answer “yes” to “do you trust this AI?” for completely different reasons, and the same person may answer differently after a single good or bad interaction.
That’s why UX researchers measure behavior instead.
The behaviors that reveal trust
Trust shows up in the small decisions people make while using the product. Five behaviors are especially important:
- Acceptance. How often people act on the AI’s recommendation instead of changing or ignoring it.
- Verification. How often people check the AI’s work before relying on it, and whether that changes over time.
- Error detection. Whether people notice incorrect or misleading answers. This matters most in high-stakes products, where acting on a bad recommendation has real consequences.
- Recovery. What people do after the AI makes a mistake. Do they stop using it, start checking every answer, or regain confidence after a few successful interactions?
- Predictability. Whether people understand the AI well enough to know when they can rely on it — and when they should double-check it.
Together, these behaviors provide a much more reliable picture of trust than simply asking people how they feel about the AI.
How researchers measure those behaviors
No single research method captures all five behaviors, so most studies combine several approaches.
- Trust calibration tasks. Participants complete realistic tasks while the quality of the AI’s responses is deliberately varied. This shows when people trust the AI, when they become skeptical, and whether that shift happens at the right time.
- Diary studies. Researchers follow the same participants over several weeks to understand how trust changes with repeated use — when people rely on the AI, when they verify it, and what ultimately makes them keep using it or abandon it.
- Critical incident interviews. Rather than discussing the product in general, researchers focus on specific moments: the last time the AI got something wrong, or the first time someone stopped trusting it. Those conversations reveal exactly what built confidence — or broke it.
- Behavioral analytics. Researchers analyze what people do at scale: accepting recommendations, editing AI output, asking the AI to try again, abandoning workflows, or returning to manual work. Combined with interviews and observational research, that data explains what changed and why.
The goal is the right amount of trust
Most teams try to increase trust. The real goal is calibrated trust.
Trust an unreliable AI too much, and people act on bad recommendations. Trust a reliable AI too little, and they ignore good ones and keep doing the work themselves.
Calibrated trust means people understand what the AI does well, where its limits are, and when they should verify its work. UX research measures whether that balance exists and identifies what needs to change when it doesn’t.
Pitfalls
Common mistakes when researching AI products
Nearly all of these come from carrying habits that worked on traditional software into a product that behaves nothing like it. Each one has a tell.
Waiting until after launch
If your first research study begins because the numbers are dropping, users have already made up their minds. AI products rarely get a second chance.
Relying only on analytics
If you can point to exactly where usage falls off but have three competing theories about why, you’re missing the half that matters. Analytics record the decision. Research explains it.
Treating research on AI products like traditional software research
If your study plan is a task list with pass-or-fail criteria, it won’t hold up against an AI that answers two participants differently. Nobody judges an AI product the way they judge a checkout flow.
Expecting hiring to keep pace with demand
If your research backlog holds items older than the current roadmap, hiring alone won’t close that gap. Product managers make AI decisions with or without evidence.
Assuming one study is enough
If your most recent findings describe a model version that’s no longer in production, you’re deciding on history. The product moved and the research didn’t.
Testing the product the team already understands
If participants come from adjacent teams, or the sessions run on examples the AI handles well, the study will mostly confirm what the team already believed. That’s the most expensive kind of clean result.
Compliance
About privacy, compliance, and responsible UX research on AI products
Researching AI products carries obligations that traditional UX research doesn’t.
A study collects AI conversations, handles sensitive company data, tests systems nobody has announced, and produces evidence that product, legal, and compliance teams may lean on later. Done properly it protects participants, protects the company, and produces findings that survive scrutiny.
Protecting participants and their data
Participants need to know exactly what they’ve signed up for: when they’re talking to AI, what happens to what they type, whether the conversation gets stored, and whether it might be used to improve the product later. Consent forms work best when they answer those questions outright rather than falling back on a generic research agreement.
Beyond consent, three practices do most of the work:
- Collect as little sensitive information as the study allows.
- De-identify findings before they’re shared.
- Keep personal information, confidential documents, source code, and proprietary business data out of AI tools unless the organization has approved the tool and knows how the data is stored.
Supporting compliance efforts
The NIST AI Risk Management Framework and the EU AI Act expect organizations to demonstrate transparency, human oversight, and explainability, along with a real path for someone to challenge an important decision the AI made. UX research is how an organization shows those safeguards work on actual people.
It answers questions technical testing can’t:
- Can people tell they’re interacting with AI?
- Do they understand the explanations they’re given?
- Can they recognize when a recommendation needs review rather than automatic acceptance?
- Can they find and use the controls built for human oversight?
Research sits alongside legal reviews, security assessments, and algorithmic audits without replacing any of them. What it adds is evidence of how the safeguards behave in practice, which legal, product, and risk teams can’t source anywhere else.
Researching confidential AI products
Plenty of studies involve products nobody has announced, and the research process has to guard that information as carefully as the product team does. Depending on the project that means participant NDAs, prototypes stripped of company branding, tight access control on recordings and transcripts, and research data stored only in approved systems.
Part 4
Why companies invest in UX research for AI products
Building AI products is already expensive. Engineering, product, design, infrastructure, and the models themselves all pull from the same budget, so research has to earn its place against them.
Every team makes assumptions about how people will use an AI feature, what they’ll trust, where they’ll struggle, and whether it fits their workflow. Some of those assumptions hold and some don’t. Research finds the bad ones while they’re still cheap.
The return shows up in four places.
Better ROI
An AI feature only returns anything once people adopt it, trust it appropriately, and build it into their day. Research finds the friction holding adoption down, which produces a roadmap for improving the experience instead of a guess about what to build next.
The model spend is already committed. Adoption is what decides whether it pays.
Engineering time spent on the right problems
Engineering time is the most expensive thing most product organizations have. Research sorts what deserves building, what needs redesigning before development starts, and what shouldn’t be built at all.
Work avoided is worth as much as work shipped.
Product decisions backed by evidence
Product teams make hundreds of decisions a year and can’t test all of them. Without research, the riskiest of those decisions rest on assumptions nobody ever checked.
Research puts evidence behind the ones that carry the most risk, which lets a team move faster rather than slower.
A case for the next AI investment
Building the product is only half the job. Eventually leadership asks the harder question.
Model accuracy and benchmark scores don’t answer that. What executives want to know is whether people are actually using the AI, whether it’s saving them measurable time, and whether adoption is still growing.
UX research produces those numbers along with the explanation behind them: where people trust the AI, where they abandon it, how it changed the way they work, and what’s holding wider adoption back. That’s what connects an AI product to results a CFO recognizes.
The next investment decision turns on what the last one returned. Whether the technology works is assumed by then.
Build or buy
Build your own UX research capability or bring in a partner?
Most companies end up doing both. The real question is which work makes sense to own, and which is better handled by a specialist.
When building in-house is the right choice
If UX research is going to be a core part of how your company builds AI products, investing in an internal team is the right long-term decision.
Internal researchers develop deep knowledge of your customers, your product, your domain, and your organization. That context compounds over time, making every future study stronger than the last.
For companies building AI products continuously, that institutional knowledge becomes a competitive advantage.
Why building that capability is hard
Recognizing the need is the easy part. Building the team is harder.
AI products are being developed faster than most organizations can hire researchers. Hiring takes time, experienced AI researchers are in short supply, and many organizations already have more research requests than their teams can handle.
The most specialized roles are even harder to find. Researchers with experience in conversational AI, accessibility, probabilistic systems, research operations, and AI evaluation remain scarce, and many companies don’t need enough of that work to justify hiring them full-time.
Even mature research organizations run into this. Demand usually grows faster than headcount.
When bringing in a partner makes sense
External research partners are most valuable in three situations.
- You need answers faster than you can hire. An important product decision can’t wait for the next hiring cycle.
- You need expertise your team doesn’t have. Some studies require specialized experience that most internal teams don’t encounter often enough to build themselves.
- Your research backlog keeps growing. Your team already knows what needs to be studied. They simply don’t have the capacity to do it all.
The different ways to get outside help
Not every external partner solves the same problem.
| Option | What it gives you | What to consider |
|---|---|---|
| Large consultancies | Scale, global reach, and experience with large enterprise programs. | Higher costs, longer onboarding, and teams that are often staffed with generalists. |
| Boutique UX research agencies | Deep expertise in a specific area of research. | Limited capacity for multiple concurrent studies or large international programs. |
| Staffing firms | Researchers who join your team quickly. | Your organization still owns the study design, methodology, analysis, and recommendations. |
| Small agencies | Flexibility, close collaboration, and lower costs. | Enterprise procurement, compliance, or participant operations may be more limited. |
| UX research-as-a-service | Specialist researchers who scale with demand without adding permanent headcount. | The quality depends entirely on choosing a partner with a proven methodology and delivery record. |
Capacity or ownership?
Cost is the obvious difference between these options. The one that matters more is who owns the outcome.
A contract researcher adds capacity, but your team still owns the research question, the study design, the analysis, and the recommendations.
A research partner owns the problem from end to end. They help define the research, recruit participants, conduct the study, analyze the findings, and deliver recommendations your team can confidently act on.
There’s another advantage that’s easy to overlook. When an important launch decision is being debated, evidence gathered outside the delivery team often carries additional credibility. The discussion stays focused on the findings instead of becoming a debate about who produced them.
For most organizations, the strongest model combines both approaches. An internal team provides product knowledge, continuity, and long-term direction. A trusted research partner adds specialized expertise and flexible capacity whenever demand grows beyond what the internal team can support.
Part 5
Meet Akraya, a UX research-as-a-service partner
Akraya helps enterprise product teams build better AI products through UX research.
We work with Fortune 500 organizations on the questions that have to be answered before a product ships. Will people trust this AI? Will they adopt it? Does it fit the way they already work? What has to change before launch?
Akraya has worked inside enterprise product organizations for 25 years, first as a staffing provider and now as a managed UX research partner. The research practice came out of that access: years spent next to the product, design, and engineering teams at some of the largest technology companies in the world, learning how they ship and what slows them down.
Two ways to work with us
Embedded UX research
An embedded researcher owns research strategy across a program: they set the research direction, decide what to study next, carry context from one study into the next, and keep the whole effort pointed at the decisions the product team has to make. This is the model when the gap is research strategy rather than a single answer.
Rapid UX research
A scoped study owned end to end and delivered in one to three weeks, measured from brief rather than from contract to kickoff. You bring the question. We design the study, recruit participants, run the sessions, analyze the findings, and return a recommendation, without adding anyone to your meeting load. This is the model when capacity is the gap rather than strategy.
Both operate under a PMO-backed structure with defined SLAs, which is what turns a delivery record into something a procurement team can evaluate rather than take on faith.
What you gain working with Akraya
A roadmap that keeps moving
Research shouldn’t be the thing standing between an idea and a release. We answer the critical questions fast enough that development keeps moving.
Less costly rework
The earlier you find out people don’t trust, understand, or adopt an AI feature, the cheaper the fix. We surface those problems before they turn into months of engineering rework.
Higher AI adoption
A high-performing model guarantees nothing about the product. We find the usability, workflow, and trust barriers keeping people from using AI in their daily work, and tell you what to fix first.
More of the questions answered
As more teams build AI products, research demand outruns hiring. Instead of handing your team more people to direct, we take individual questions and run them through to a decision. Your researchers stay on product strategy while more questions get answered properly.
Built for enterprise AI
A lot of what we study is confidential, complex, and built by large cross-functional teams. The practice is designed for that environment: secure, rigorous, and fast enough to keep up with how modern product teams actually work.
That work spans hardware, enterprise software, AI platforms, and wearables. Enterprise AI rarely lives in one of those categories alone, and a product that combines a device, an app, and a model needs researchers who can study all three.
Researchers are matched to the research context rather than drawn from a general bench. That includes PhD-level human factors and quantitative psychology specialists, with principal-level human-technology interaction leadership embedded in every hardware engagement. Akraya is also a Google Cloud partner.
Trusted by Enterprises Worldwide
Where to start
You don’t need a research plan before getting in touch. Most engagements begin with a scoping call, and the point of that call is working out what you actually need.
What we cover:
- Where the product is now — pre-launch, recently shipped, or live with adoption stalling.
- What decision is waiting on evidence, and when it has to be made.
- What has already been measured, and what those signals point at.
- Which of the three adoption risks looks most exposed.
What you get back is a recommended first study: the question it answers, the method, the timeline, what lands on your desk at the end, and who runs it. If an ongoing program fits your situation better than a single scoped study, we’ll say so.
And if research isn’t what you need right now, we’ll tell you that too.
FAQ
Frequently asked questions
What is UX research for AI products?
UX research for AI products studies whether people can rely on an AI system whose behavior changes over time, and what has to be true for them to keep relying on it. It covers usability, but it also measures trust, workflow fit, and adoption, because an AI feature can be accurate and still go unused.
How is UX research for AI products different from traditional UX research?
The methods overlap, but the questions differ. Traditional UX research asks whether people can complete a task. Research on AI products also asks whether they trust the system appropriately, understand its limits, know when to verify its output, and manage to build it into their work. It also requires capturing exactly what the model produced in each session, because the same prompt can return different answers to different participants.
How do you measure user trust in an AI product?
Trust is measured through behavior rather than self-report, because asking someone whether they trust an AI produces answers that don’t predict what they do. Researchers track five behaviors: acceptance of recommendations, verification before acting, error detection, recovery after a mistake, and predictability. Methods include trust calibration tasks, multi-week diary studies, critical incident interviews, and behavioral analytics.
How do you measure whether an AI product is successful?
Three things, measured against a pre-launch baseline: adoption that survives past the first week, calibrated trust, and workflow integration that removes steps rather than adding them. Those map to numbers leadership already tracks, including time saved per task, shorter cycle times, and fewer abandoned workflows.
Our AI model performs well in benchmark testing. Do we still need UX research?
Usually yes. Benchmarks measure technical performance and say nothing about whether people trust the AI, understand it, or choose to use it. Plenty of products post excellent benchmark numbers and still struggle with adoption, because the experience is the limiting factor rather than the model.
We’re still pre-launch. Is it too early to start UX research?
Pre-launch is usually the best time to start. Many AI experiences can be tested before the model is finished using methods such as Wizard of Oz studies, where a researcher generates the responses behind the scenes. Catching usability and workflow problems before development wraps costs a fraction of catching them after launch.
How long does a UX research study take?
A focused evaluative study typically delivers findings in one to three weeks, measured from brief to findings rather than from contract to kickoff. Discovery research and longitudinal studies take longer because they explore broader questions or track behavior over time. Participant recruitment is often the biggest factor in the timeline.
What is UX research as a service?
UX research as a service treats research capacity as something an organization scales up and down against demand, the way engineering teams already treat cloud infrastructure, rather than something hired permanently or bought one project at a time. In practice it means a specialist team takes a research question end to end — designing the study, recruiting participants, running it, and delivering recommendations — under a defined structure procurement can evaluate.
Why bring in an outside partner if we already have an internal UX research team?
The constraint is usually capacity or specialized expertise rather than capability. Headcount approvals take months, many organizations are running freezes on research roles, and the profiles AI research needs most are scarce. An outside partner takes individual questions end to end, which keeps internal researchers on product strategy instead of executing studies they don’t have room for.
Can UX research for AI products scale across multiple product teams and regions?
Yes, though consistency becomes the constraint. Research across multiple teams needs shared methods and shared measurement, or findings can’t be compared. International studies also need local recruitment and in-language moderation, because expectations, mental models, and trust in AI vary by market.
How do you research confidential or unreleased AI products?
Through unbranded studies, participant nondisclosure agreements, controlled access to study materials, and data-handling procedures built for sensitive environments. Research can also be conducted inside client systems where the engagement requires it, so no data leaves the client environment.
What should we look for in a UX research partner for AI products?
Evidence they have researched AI products rather than only traditional software. Look for a clear method for measuring trust and adoption, an approach that accounts for changing model behavior, and the ability to work directly with product, design, engineering, and data science. Ask who will actually run the study, because the experience of that person shows up in the findings.
What should a UX research proposal include?
The business decision the research will inform, the recommended methodology and why it was chosen, how participants will be recruited, what you receive at the end, the timeline, and who is running it. If any of those are vague, the proposal can’t be evaluated against what you’re trying to decide.
Would a good UX research partner ever recommend not building a feature?
Yes. UX research exists to reduce risk rather than to validate ideas. Sometimes the most valuable finding is that a feature won’t solve the problem it was built for, or that something simpler would serve users better.