The grid

Three families. Nineteen sources.

Every source is obtained one of three ways.

01 / 03

Asked

you put questions to people

02 / 03

Observed

you read the signals people and markets leave behind

03 / 03

Assembled

you gather what already exists

Every source gets a Quantitative read and a Qualitative read. There is no quant-only or qual-only source; only sources read with one eye. The Pass column marks the purpose each serves best: I Build (insight), M Prove (measurement), B both. The Lens column tags each source to the Assessment lenses it feeds.

The Q2 Grid

Nineteen sources, two reads each.

SourceQuant readQual readWhat AI scalesPassLens
Asked
SurveyFrequencies, scales, cell comparisonsOpen-endsOpen-end coding at n=150+B5, 7
InterviewTheme frequency per cellMeaning, verbatimsAI-moderated, 30–50 per cellI1, 2, 3, 5, 6
Focus groupIn-room concept scoresGroup language, dynamicsTranscript synthesis; moderation stays humanI5
Message / concept testPreference, credibility ratingsWhy it lands, or feels defensiveVariant volumeB5
Internal expert perspectivesStructured internal polls, Delphi roundsSME interviews, institutional memoryCross-interview synthesisI4, 7
Observed
Media monitoringVolume, share of voice, sentimentHow the story is told, by whomFull-corpus readingB6
Social listeningVolume, reach, sentimentActual language; the say-do gapThread and community readingB1, 5, 6
SearchQuery volume, trendsWhat people are actually askingLong-tail clusteringB1, 6
LLM auditCitation share, positionWhat models say, and citeMulti-model, multi-prompt runsB6
Owned / webTraffic, engagementPaths, behaviourSession synthesisM6, 7
Paid performanceReach, CPM, conversionCreative resonance by segmentVariant testingM6
Customer serviceContact volume, CSAT, resolutionTranscripts, complaint languageEvery transcript readB5
RegulatoryFilings, comment counts, approval timelinesComment content, decision language, testimonyDocket-scale readingB2
Financial performanceResults, share price, multiples vs peersEarnings-call language, analyst questionsTranscript reading across peers and timeM3
Assembled
Secondary / literaturePublished dataPrior findings, gapsRapid reviewIAny
Syndicated trackersBenchmarksCategory narrativeCross-study reconciliationM5, 7
Expert analysis & ratingsRatings, rankings, scores, awardsThe reasoning in the reportsCross-report synthesisB1, 3
Competitive / peerEvery row above, on peersTheir story vs yoursSame grid, run on peersB1

Placement notes: customer service, regulatory and financial performance sit in Observed because they are signal streams over time, which is what makes them measurement-ready. Expert analysis & ratings sits in Assembled because it already exists in report form. Internal expert perspectives sits in Asked because you have to go and get it. The Observed family is POET's evidence layer plus search and LLMs (cross-link to Paper No. 2).

Markdown edition: the Q2 Grid template

The grid as a research-plan template — fill the Quant and Qual columns, count the cells, size each cell, run the five tests.

Download the template (.md)
Coverage compounds

Not a menu. A checklist.

The grid is not a menu; it is a checklist. The more of it a project reads, the more immersive the insight and the more impactful the measurement, for three reasons:

Triangulation.
An insight that appears in all three families — asked, observed, assembled — has been found three different ways. That is the cheapest confidence you will ever buy, and it raises t (transferability) without a single extra interview.
Immersion.
Each source adds a dimension: what people say, what they do, what already exists about them. Together they produce a picture you can walk around in, which is what personas need.
Measurement.
Every source read for Build becomes a baseline for Prove. Coverage now is measurement later.

Coverage was always desirable and rarely affordable, because someone had to read it all and hold it in one head. That is the constraint AI removes: it reads the whole grid together and keeps the sources in one frame, which is what “making sense of all of it together” means in practice. The human role moves to designing the coverage and judging the synthesis.

Simple measure for a project

grid coverage = sources read ÷ sources relevant

A campaign research plan that reads four of twelve relevant rows has covered a third of what it could know.

How much is enough

Sample size is how you buy probability.

The reframe. The classic justification for small qualitative samples is saturation. The empirical literature is narrow: saturation is reached within 9–17 interviews or 4–8 focus groups, mainly with homogeneous populations and narrow objectives.S8 Finer cut: code saturation (the range of issues) around 9 interviews; meaning saturation (fully understanding them) around 24.S9 Focus groups: four to identify the issues, more to understand them.S10 But small samples were never only about saturation. They were also about analyst throughput — Kvale's “1,000-page question”: you cannot analyse data you cannot read.S11 AI removes the throughput constraint. The question changes from “when can I stop” to “what job is the qual doing.”

Three jobs, sized per cell

JobClaimQual per cellBasis
Discover“These are the things people say”9–12Code saturation
Understand“This is why, and what it means”~24Meaning saturation
Compare“Group A raises X more than B”30–50Emerging practitioner normS12, S13
Quantify“38% raised X”Approaches quant nContent analysis; needs power

Quant per cell: ~100 for a directional comparison; ~385 total for ±5 at 95%.

Starter numbers: quantitative thresholds

Margin of error at 95% confidence (approximate, 1/√n; exact figures inS24, S25):

n± at 50%
10010.0
2007.1
4005.0
5004.5
1,0003.1
2,0002.2

Doubling from 1,000 to 2,000 buys about one pointS26. The margin applies to the total sample only; every subgroup has a wider oneS26.

Starter targets by population

ScaleQ2 defaults; adjust to the decision

PopulationStarter nWhy
Market-wide / general population1,000±3.1; the point where extra sample buys little
Single market, region or segment400–500±4.5–5
B2B decision-makers200±7; the practical ceiling for a hard-to-reach professional universe
Niche stakeholders (policymakers, regulators, investors, analysts)25–50 (e.g. 35)Treated as representative of a small universe; report as counts and themes, not percentages with a margin
Employees — ~10% of the population as a rule of thumb; coverage of every market, role and demographic matters more than the total~10%Response-rate benchmarks 65–85%S27, S28
Any reported subgroup100 to compare; 25 to report at allComparison needs ±10 or better; 25 is the ScaleQ2 floor for anonymity and stability (platform norms suppress below 5–10S29, S30; Q2 sets a higher bar)

Two rules over the table. Coverage before count: a sample that misses a market, a role or a demographic is wrong at any n. Never report below n=25. It protects anonymity and it keeps a chart honest about its width.

Total = cells × per-cell minimum. Count the cells first. Four usage states × 30 = 120 interviews to compare them. Three persuadability tiers × two audiences × 10 = 60 to discover.

Norms with real support

  • Human first, then scale: ~10 well-designed interviews, review, then 50 or 100.S14
  • Read the raw edges: AI synthesis is the starting point; read 3–5 transcripts yourself.S14
  • Report code saturation and meaning saturation separately.S9
  • Scale does not fix recruitment. 500 AI interviews of a convenience sample is a bigger convenience sample.
  • Quantify themes only with a sample built for it; below that, report prevalence within the sample and say so.S13
  • Synthetic respondents pretest instruments; they do not produce findings.
  • Rule of thumb for the room: scale a cell until the next ten interviews change nothing; then stop and spend on recruiting the cell you are missing.
What doesn't scale

What doesn't scale

Human moderation on contested or emotionally complex topics. Executive interviews. Ethnography. Anything where the interviewer is the instrument.

The paper names its own limits so the “why not 500?” critique is answered before it is asked.

Sources

Sources cited on this page.

  1. [S8]

    Hennink, M. & Kaiser, B. (2022). Sample sizes for saturation in qualitative research: a systematic review. Social Science & Medicine 292. https://pubmed.ncbi.nlm.nih.gov/34785096/

  2. [S9]

    Hennink, Kaiser & Marconi (2017). Code saturation versus meaning saturation. Qualitative Health Research 27(4). https://pmc.ncbi.nlm.nih.gov/articles/PMC9359070

  3. [S10]

    Hennink, Kaiser & Weber (2019). What influences saturation? Focus group sample sizes. Qualitative Health Research 29(10). https://pmc.ncbi.nlm.nih.gov/articles/PMC6635912/

  4. [S11]

    Kvale (1996) throughput argument, as summarised in User Intuition (2026). https://www.userintuition.ai/posts/ai-qualitative-research-at-scale/

  5. [S12]

    Perspective AI: n=30 directional, n=50 per segment, n=100+ concept tests (vendor guidance). https://getperspective.ai/blog/ai-moderated-interviews-how-they-work-when-to-use-them-and-what-they-replace

  6. [S13]

    Merren: AI-moderated 20–100+, 15–30 per segment; larger samples stabilise thematic analysis (vendor guidance). https://merren.io/blog/sample-size-qualitative-research

  7. [S14]

    Koji: start with ~10, review, scale to 50–100; read raw transcripts (vendor guidance). https://www.koji.so/docs/ai-moderated-interviews

  8. [S24]

    Forum Research, margin-of-error table by sample size and observed proportion. https://forumresearch.com/tools-margin-of-error.asp

  9. [S25]

    Sample size by population and margin (PMC table). https://pmc.ncbi.nlm.nih.gov/articles/PMC5723800/table/t3

  10. [S26]

    AAPOR, Margin of Sampling Error / Credibility Interval explainer. https://aapor.org/wp-content/uploads/2023/01/Margin-of-Sampling-Error-508.pdf

  11. [S27]

    Employee survey response-rate benchmarks (65–85%; by company size). https://leadx.org/articles/what-is-a-good-employee-engagement-survey-participation-rate/

  12. [S28]

    Employee survey response-rate benchmarks (70–85%; <60% a warning). https://www.culturemonkey.io/employee-engagement/employee-engagement-survey-benchmark-data/

  13. [S29]

    Gallup: aggregated reporting with minimum response thresholds. https://www.gallup.com/workplace/692474/workplace-employee-surveys.aspx

  14. [S30]

    Platform norms: minimum reporting group size 5–10. https://heartcount.com/employee-engagement/survey-response-rate/

Contact

Prove it. Then bring it to life.

Coverage now is measurement later. Five tests decide whether the insight can leave the room.