Three families. Nineteen sources.
Every source is obtained one of three ways.
Asked
— you put questions to people
Observed
— you read the signals people and markets leave behind
Assembled
— you gather what already exists
Every source gets a Quantitative read and a Qualitative read. There is no quant-only or qual-only source; only sources read with one eye. The Pass column marks the purpose each serves best: I Build (insight), M Prove (measurement), B both. The Lens column tags each source to the Assessment lenses it feeds.
Nineteen sources, two reads each.
| Source | Quant read | Qual read | What AI scales | Pass | Lens |
|---|---|---|---|---|---|
| Asked | |||||
| Survey | Frequencies, scales, cell comparisons | Open-ends | Open-end coding at n=150+ | B | 5, 7 |
| Interview | Theme frequency per cell | Meaning, verbatims | AI-moderated, 30–50 per cell | I | 1, 2, 3, 5, 6 |
| Focus group | In-room concept scores | Group language, dynamics | Transcript synthesis; moderation stays human | I | 5 |
| Message / concept test | Preference, credibility ratings | Why it lands, or feels defensive | Variant volume | B | 5 |
| Internal expert perspectives | Structured internal polls, Delphi rounds | SME interviews, institutional memory | Cross-interview synthesis | I | 4, 7 |
| Observed | |||||
| Media monitoring | Volume, share of voice, sentiment | How the story is told, by whom | Full-corpus reading | B | 6 |
| Social listening | Volume, reach, sentiment | Actual language; the say-do gap | Thread and community reading | B | 1, 5, 6 |
| Search | Query volume, trends | What people are actually asking | Long-tail clustering | B | 1, 6 |
| LLM audit | Citation share, position | What models say, and cite | Multi-model, multi-prompt runs | B | 6 |
| Owned / web | Traffic, engagement | Paths, behaviour | Session synthesis | M | 6, 7 |
| Paid performance | Reach, CPM, conversion | Creative resonance by segment | Variant testing | M | 6 |
| Customer service | Contact volume, CSAT, resolution | Transcripts, complaint language | Every transcript read | B | 5 |
| Regulatory | Filings, comment counts, approval timelines | Comment content, decision language, testimony | Docket-scale reading | B | 2 |
| Financial performance | Results, share price, multiples vs peers | Earnings-call language, analyst questions | Transcript reading across peers and time | M | 3 |
| Assembled | |||||
| Secondary / literature | Published data | Prior findings, gaps | Rapid review | I | Any |
| Syndicated trackers | Benchmarks | Category narrative | Cross-study reconciliation | M | 5, 7 |
| Expert analysis & ratings | Ratings, rankings, scores, awards | The reasoning in the reports | Cross-report synthesis | B | 1, 3 |
| Competitive / peer | Every row above, on peers | Their story vs yours | Same grid, run on peers | B | 1 |
Placement notes: customer service, regulatory and financial performance sit in Observed because they are signal streams over time, which is what makes them measurement-ready. Expert analysis & ratings sits in Assembled because it already exists in report form. Internal expert perspectives sits in Asked because you have to go and get it. The Observed family is POET's evidence layer plus search and LLMs (cross-link to Paper No. 2).
Markdown edition: the Q2 Grid template
The grid as a research-plan template — fill the Quant and Qual columns, count the cells, size each cell, run the five tests.
Not a menu. A checklist.
The grid is not a menu; it is a checklist. The more of it a project reads, the more immersive the insight and the more impactful the measurement, for three reasons:
- Triangulation.
- An insight that appears in all three families — asked, observed, assembled — has been found three different ways. That is the cheapest confidence you will ever buy, and it raises t (transferability) without a single extra interview.
- Immersion.
- Each source adds a dimension: what people say, what they do, what already exists about them. Together they produce a picture you can walk around in, which is what personas need.
- Measurement.
- Every source read for Build becomes a baseline for Prove. Coverage now is measurement later.
Coverage was always desirable and rarely affordable, because someone had to read it all and hold it in one head. That is the constraint AI removes: it reads the whole grid together and keeps the sources in one frame, which is what “making sense of all of it together” means in practice. The human role moves to designing the coverage and judging the synthesis.
Simple measure for a project
grid coverage = sources read ÷ sources relevant
A campaign research plan that reads four of twelve relevant rows has covered a third of what it could know.
Sample size is how you buy probability.
The reframe. The classic justification for small qualitative samples is saturation. The empirical literature is narrow: saturation is reached within 9–17 interviews or 4–8 focus groups, mainly with homogeneous populations and narrow objectives.S8 Finer cut: code saturation (the range of issues) around 9 interviews; meaning saturation (fully understanding them) around 24.S9 Focus groups: four to identify the issues, more to understand them.S10 But small samples were never only about saturation. They were also about analyst throughput — Kvale's “1,000-page question”: you cannot analyse data you cannot read.S11 AI removes the throughput constraint. The question changes from “when can I stop” to “what job is the qual doing.”
Three jobs, sized per cell
| Job | Claim | Qual per cell | Basis |
|---|---|---|---|
| Discover | “These are the things people say” | 9–12 | Code saturation |
| Understand | “This is why, and what it means” | ~24 | Meaning saturation |
| Compare | “Group A raises X more than B” | 30–50 | Emerging practitioner normS12, S13 |
| Quantify | “38% raised X” | Approaches quant n | Content analysis; needs power |
Quant per cell: ~100 for a directional comparison; ~385 total for ±5 at 95%.
Starter numbers: quantitative thresholds
Margin of error at 95% confidence (approximate, 1/√n; exact figures inS24, S25):
| n | ± at 50% |
|---|---|
| 100 | 10.0 |
| 200 | 7.1 |
| 400 | 5.0 |
| 500 | 4.5 |
| 1,000 | 3.1 |
| 2,000 | 2.2 |
Doubling from 1,000 to 2,000 buys about one pointS26. The margin applies to the total sample only; every subgroup has a wider oneS26.
Starter targets by population
ScaleQ2 defaults; adjust to the decision
| Population | Starter n | Why |
|---|---|---|
| Market-wide / general population | 1,000 | ±3.1; the point where extra sample buys little |
| Single market, region or segment | 400–500 | ±4.5–5 |
| B2B decision-makers | 200 | ±7; the practical ceiling for a hard-to-reach professional universe |
| Niche stakeholders (policymakers, regulators, investors, analysts) | 25–50 (e.g. 35) | Treated as representative of a small universe; report as counts and themes, not percentages with a margin |
| Employees — ~10% of the population as a rule of thumb; coverage of every market, role and demographic matters more than the total | ~10% | Response-rate benchmarks 65–85%S27, S28 |
| Any reported subgroup | 100 to compare; 25 to report at all | Comparison needs ±10 or better; 25 is the ScaleQ2 floor for anonymity and stability (platform norms suppress below 5–10S29, S30; Q2 sets a higher bar) |
Two rules over the table. Coverage before count: a sample that misses a market, a role or a demographic is wrong at any n. Never report below n=25. It protects anonymity and it keeps a chart honest about its width.
Total = cells × per-cell minimum. Count the cells first. Four usage states × 30 = 120 interviews to compare them. Three persuadability tiers × two audiences × 10 = 60 to discover.
Norms with real support
- Human first, then scale: ~10 well-designed interviews, review, then 50 or 100.S14
- Read the raw edges: AI synthesis is the starting point; read 3–5 transcripts yourself.S14
- Report code saturation and meaning saturation separately.S9
- Scale does not fix recruitment. 500 AI interviews of a convenience sample is a bigger convenience sample.
- Quantify themes only with a sample built for it; below that, report prevalence within the sample and say so.S13
- Synthetic respondents pretest instruments; they do not produce findings.
- Rule of thumb for the room: scale a cell until the next ten interviews change nothing; then stop and spend on recruiting the cell you are missing.
What doesn't scale
Human moderation on contested or emotionally complex topics. Executive interviews. Ethnography. Anything where the interviewer is the instrument.
The paper names its own limits so the “why not 500?” critique is answered before it is asked.
Sources cited on this page.
- [S8]
Hennink, M. & Kaiser, B. (2022). Sample sizes for saturation in qualitative research: a systematic review. Social Science & Medicine 292. https://pubmed.ncbi.nlm.nih.gov/34785096/
- [S9]
Hennink, Kaiser & Marconi (2017). Code saturation versus meaning saturation. Qualitative Health Research 27(4). https://pmc.ncbi.nlm.nih.gov/articles/PMC9359070
- [S10]
Hennink, Kaiser & Weber (2019). What influences saturation? Focus group sample sizes. Qualitative Health Research 29(10). https://pmc.ncbi.nlm.nih.gov/articles/PMC6635912/
- [S11]
Kvale (1996) throughput argument, as summarised in User Intuition (2026). https://www.userintuition.ai/posts/ai-qualitative-research-at-scale/
- [S12]
Perspective AI: n=30 directional, n=50 per segment, n=100+ concept tests (vendor guidance). https://getperspective.ai/blog/ai-moderated-interviews-how-they-work-when-to-use-them-and-what-they-replace
- [S13]
Merren: AI-moderated 20–100+, 15–30 per segment; larger samples stabilise thematic analysis (vendor guidance). https://merren.io/blog/sample-size-qualitative-research
- [S14]
Koji: start with ~10, review, scale to 50–100; read raw transcripts (vendor guidance). https://www.koji.so/docs/ai-moderated-interviews
- [S24]
Forum Research, margin-of-error table by sample size and observed proportion. https://forumresearch.com/tools-margin-of-error.asp
- [S25]
Sample size by population and margin (PMC table). https://pmc.ncbi.nlm.nih.gov/articles/PMC5723800/table/t3
- [S26]
AAPOR, Margin of Sampling Error / Credibility Interval explainer. https://aapor.org/wp-content/uploads/2023/01/Margin-of-Sampling-Error-508.pdf
- [S27]
Employee survey response-rate benchmarks (65–85%; by company size). https://leadx.org/articles/what-is-a-good-employee-engagement-survey-participation-rate/
- [S28]
Employee survey response-rate benchmarks (70–85%; <60% a warning). https://www.culturemonkey.io/employee-engagement/employee-engagement-survey-benchmark-data/
- [S29]
Gallup: aggregated reporting with minimum response thresholds. https://www.gallup.com/workplace/692474/workplace-employee-surveys.aspx
- [S30]
Platform norms: minimum reporting group size 5–10. https://heartcount.com/employee-engagement/survey-response-rate/
Prove it. Then bring it to life.
Coverage now is measurement later. Five tests decide whether the insight can leave the room.