Generative research explores user needs and problems before a solution exists, using methods like interviews and ethnography to answer "what should we build?" Evaluative research tests and validates an existing concept, prototype, or product, using methods like usability testing and A/B testing to answer "how well does this work?" Most research programs use both at different project stages.

Summary:
|
All research requests come in one of two forms: either the team hasn't yet decided what to build or they have built something and would like to know whether it works.
They are different questions and each of them requires its own kind of research. It is one of the most frequent and most costly errors in the product process to carry out a usability test without first agreeing on the problem that the product is meant to solve, or to conduct discovery interviews three weeks before a launch which no one is going to put off.
The guide explains the differences between generative and evaluative research, the methods associated with each, how to decide between them, and the way in which the best teams combine both.
What Is Generative Research?
Generative research is carried out before any solution has been developed in order to investigate user needs, behaviours, and problems; it is also referred to as exploratory, discovery, or foundational research, the various terms being used interchangeably.
The main question is this: which problem should we aim to solve and why do people act the way they do?
Generative research doesn't test an idea; it comes up with new ones. If you enter the situation without having a hypothesis or with one that you are deliberately trying to disprove, you will leave with a clearer understanding of the problem area than the team possessed before. This clearer understanding covers the aspects that no one had originally thought to ask about, and it is often in those areas that value lies.
The output is not a score. It is understanding: personas, journey maps, opportunity areas, jobs-to-be-done, and a set of prioritized problems worth solving. This is the stage where asking the right research question matters most, because a badly framed generative study produces confident answers to a question the business was never asking.
What Is Evaluative Research?
Evaluative research is used to test and validate something that is already available—such as a concept, a prototype, a live feature, or a complete product—and it is also referred to as evaluation, summative, or assessment research.
The main question is this: how effective is our solution and does it satisfy the need that we believe it meets?
Evaluative research starts with a hypothesis. This checkout flow reduces abandonment. Users will understand this navigation label. This concept is more appealing than that one. The study exists to support or reject that claim.
The output consists of performance data and validated recommendations, including task success rates, error counts, preference results, and a specific list of changes to make. It is this level of specificity that is important. It is through evaluative research that a design debate is turned into a decision.
Generative vs Evaluative Research: Key Differences
Timing: Generative research happens before design decisions are made. Evaluative research happens after something tangible exists to react to. If there is nothing to show a participant, the study is generative whether or not anyone called it that.
Approach: Generative work is open-ended exploration, and the researcher follows what the participant finds important. Evaluative work is hypothesis-driven, and the researcher holds the structure constant so results can be compared.
Outcomes: Generative research produces personas, journey maps, and opportunity areas. Evaluative research produces metrics, benchmarks, and validated recommendations.
Risk covered: Generative research reduces the risk of building the wrong thing. Evaluative research reduces the risk of building the right thing badly. Both risks are real, and only one of them is visible in a launch review.
The business case for running both is stronger than most research teams argue. McKinsey's study of 300 publicly listed companies over five years found that top-quartile design performers saw 32 percentage points higher revenue growth than industry peers, and one of the four behaviors distinguishing them was continuous listening, testing, and iteration with end users rather than research treated as a phase.
The cost of skipping the generative half shows up later. An analysis of 12,000 FMCG product launches found that 76 percent did not survive a year of sales, and 45 percent did not last 26 weeks. Most of those products were usable. They solved problems that were not urgent enough, which is a generative research failure rather than an execution one.
Methods Used in Generative Research
In-depth interviews - The default generative method, and the most flexible.
Ethnographic and field studies - Observing behavior in context, where people do the thing rather than describe it. The fundamentals are covered in this guide to ethnographic research.
Diary studies - Capturing behavior over time rather than in a single session, which surfaces patterns a one-hour interview never reaches. Worth reading alongside this guide to diary studies.
Open card sorting - Letting participants group and label content their own way, which reveals mental models rather than testing yours.
Open-ended surveys - Useful for breadth when you need to know the range of a problem before choosing where to go deep.
Generative methods are predominantly qualitative, because the goal is depth and discovery rather than measurement. The broader framing in this overview of qualitative and quantitative research is a useful companion, with one caution: qualitative is not a synonym for generative. Plenty of evaluative work is qualitative too.
Methods Used in Evaluative Research
Usability testing: The most common evaluative method, moderated or unmoderated. This walkthrough of conducting usability testing covers the mechanics.
A/B testing: Quantitative comparison at scale, once you have enough traffic and a specific hypothesis.
Tree testing: Validating whether an information structure works without visual design in the way. See this explanation of tree testing.
Closed card sorting: Testing a category structure you have already defined, the mirror image of the open version described in this guide to card sorting.
Closed-ended surveys: Measuring satisfaction, preference, or comprehension across a larger sample.
Heuristic evaluation: An expert review against established principles, useful as a cheap first pass before involving participants.
The type of method used in evaluative research depends on the stage and may be either qualitative or quantitative; the testing of early prototypes is qualitative and on a small scale, while the post-launch validation is quantitative and on a large scale.
When to Use Generative Research
Use generative research when:
The problem space is unclear: You know a metric is bad but not why, or you know the segment exists but not what it wants.
You are entering a new market, segment, or category: Assumptions carried over from existing users are the most expensive kind, because they feel like knowledge. Multi-market discovery also needs more participants than a single-market study: research revisiting the standard saturation benchmarks found that 16 or fewer interviews covered themes within homogeneous sites, but 20 to 40 were needed for themes cutting across all sites.
Stakeholders disagree about user needs: When three people hold three different mental models of the user, no amount of evaluative testing resolves it. They will each read the results differently.
You are planning a roadmap rather than a release: Opportunity areas come from generative work.
The tradeoff is real. Generative research costs more upfront in time, recruitment, and analysis, and it does not produce a clean number for a slide. What it buys is a lower chance of building something well-crafted that nobody needed.
Choosing the right approach within that space is its own decision, and this overview of UX research methods covers how to match method to question once you know which half of the process you are in.
When to Use Evaluative Research
Use evaluative research when:
Something already exists to test. A concept, sketch, prototype, or live product. If you can put it in front of a participant, evaluative research applies.
You have a specific hypothesis or design decision to settle. Two navigation options, a pricing display, a new onboarding flow.
You need to measure rather than explore. Benchmarking a current experience, tracking whether a redesign improved anything, quantifying a known problem.
The timeline is short. Evaluative studies are faster to run and faster to analyze.
The tradeoff is scope. Evaluative research can only tell you about what you tested. It will tell you that version B outperformed version A, and it will not tell you that a third approach nobody sketched would have beaten both. That limitation is exactly the gap generative research fills.
How to Decide: A Quick Decision Framework
Work through four questions.
1. Does a solution exist yet? No means generative. Yes means evaluative. This resolves most cases on its own.
2. Can you state a testable hypothesis? If yes, you are ready for evaluative work. If your hypothesis is "users probably want something better here," you are not, and that vagueness is a generative signal.
3. What decision does this research feed? Roadmap and strategy decisions need generative input. Design and implementation decisions need evaluative input.
4. What is the cost of being wrong? High-stakes, hard-to-reverse decisions justify generative investment first. Low-stakes and reversible decisions can start with a quick evaluative study.
When in doubt, and when research credibility inside the organization is still being built, start with a small evaluative study. It produces a visible result quickly, which tends to earn the trust required to propose a larger generative investment later. Leading with a six-week discovery project in a team that has never funded research is a hard first ask.
Two pitfalls to watch, one on each side:
Generative research that rushes to solutions. The moment an interview turns into "would you use a feature that does X," it has stopped being generative and become a weak evaluative study of an idea you have not built.
Evaluative research that over-relies on numbers. A conversion drop tells you what happened, not why. Quantitative evaluative work without a qualitative layer produces precise measurements of a mystery, and teams that hit this usually need generative follow-up rather than more dashboards.
Combining Generative and Evaluative Research
Most mature programs run both. Three patterns cover most of what works in practice.
Sequential: generative first. Discovery defines the problem, design responds to it, evaluative testing refines the solution. The classic order, and the right default for new products or major initiatives.
Reverse: evaluative first. You have a live product and a metric that is underperforming. Evaluative testing identifies where users struggle, then generative research explains why. Common in established products, and often the most persuasive way to justify discovery budget, because the evaluative finding creates the question.
Parallel: continuous programs. Quarterly generative studies run alongside sprint-based evaluative testing. Discovery keeps the roadmap grounded while evaluation keeps the current release honest. This is where the practice covered in this look at generative research becomes an operating cadence rather than a project type.
One practical note on combining them: keep the artifacts separate. A journey map built from generative work and a usability benchmark from evaluative work answer different questions, and merging them into a single "research findings" document usually means one gets read as the other. That confusion is how a directional insight ends up cited as a measured result six months later.
Running Both Generative and Evaluative Studies Without Switching Tools
Most teams end up with a split stack: one tool for discovery interviews, another for usability testing, and a third for surveys. The cost is not just licensing. It is that findings live in separate systems and never get connected.
Decode's AI Moderator supports both sides of that split. It runs open-ended discovery conversations, adapting follow-up questions in real time based on what a participant says, which is what generative in-depth interviews require. It also runs structured, hypothesis-driven sessions where consistency across participants is the point, which is what evaluative work requires. Longitudinal formats like diary studies sit on the same platform, as do evaluative methods including preference testing and A/B testing.
The relevant background reading is this guide to AI moderated interviews, which covers where adaptive automated moderation fits and where a human moderator remains necessary.
Behavioral measurement adds something specific to generative work, which is worth calling out because it runs against the usual assumption that measurement belongs to the evaluative half. Facial coding at 90 plus percent accuracy and eye tracking at 96 percent accuracy record where attention went and how someone reacted during an exploratory session, not just what they said. In a discovery interview, that surfaces the moment a participant hesitated over a topic they then talked around. Entropik supports 150 plus global brands across 70 plus languages with 17 patents behind the underlying technology, so multi-market discovery and multi-market validation can run on the same ux testing platform.
For teams evaluating options across the category, this roundup of user experience testing platforms is a useful comparison point, and the question worth asking is whether a tool handles open-ended discovery as well as it handles structured testing.
Frequently Asked Questions
1. What is the difference between generative and evaluative research?
Generative research explores problems before a solution exists and answers what to build. Evaluative research tests an existing concept or product and answers how well it works.
2. When should you use generative research instead of evaluative research?
When the problem space is unclear, when entering a new market or segment, when stakeholders disagree about user needs, or when the decision at stake is strategic rather than tactical.
3. Can generative and evaluative research be combined in one study?
Partly, but be careful. A session can open with exploratory questions and move to a prototype, but analysis should keep the two separate, since one produces hypotheses and the other tests them.
4. What methods count as generative research?
In-depth interviews, ethnographic and field studies, diary studies, open card sorting, and open-ended surveys. Mostly qualitative and mostly open-ended.
5. What methods count as evaluative research?
Usability testing, A/B testing, tree testing, closed card sorting, closed-ended surveys, and heuristic evaluation. A mix of qualitative and quantitative depending on stage.
6. Is evaluative research the same as usability testing?
No. Usability testing is one evaluative method. Evaluative research also covers concept testing, preference testing, A/B testing, and post-launch measurement.
7. How many participants do you need for generative versus evaluative research?
For generative interviews, the landmark study on saturation found that saturation occurred within the first 12 of 60 interviews, with the main metathemes present by six. Later work distinguishes code saturation at around nine interviews from meaning saturation at 16 to 24, so plan for the higher range when depth of understanding matters. Evaluative usability testing typically needs five to eight participants per segment, while quantitative evaluative work needs a sample sized for statistical power.
The Practical Takeaway
The distinction is simpler than the terminology suggests. If nothing exists yet, the research is generative. If something exists and you want to know how well it works, the research is evaluative.
The failure to avoid is running one while the team needs the other: a beautifully executed usability test on a feature nobody asked for, or a discovery sprint on a decision that shipped last Tuesday. Match the research to the decision it feeds, and run both across the year rather than choosing a side. Teams that hold that balance are the ones whose user experience testing programs keep informing the roadmap instead of only validating it, which is also why user interviews remain the backbone of both halves.


