Sources#
Summary#
An AI-moderated interview is an LLM conducting a qualitative interview by voice: it holds a topic guide and a per-block objective and decides, turn by turn, whether to probe, clarify or move on. Deng, Liu, Toubia and Jain (Columbia Business School / Wharton, arXiv 2609.29143, 2026-09-24, empirical) test it in consumer research with a design that has one clean contrast and one confounded one:
- AI vs static (randomized, same batch, same platform, same guide). Static participants saw the fixed question for each open-ended block and recorded one spoken answer, with no probing. The AI (GPT-5.1, temperature 0, 300-token turns, gpt-4o-transcribe in, tts-1 out) got the same block's objective and probed adaptively. This contrast holds the guide, the medium and the recruitment fixed and varies only whether follow-ups adapt to the answer.
- AI vs human (not randomized). 24 interviews by trained GBK Collective moderators, recruited later as a separate sample with a Qualtrics intake survey, scheduled live slots, $12 instead of $10, and "unexpectedly high no-show rates" that cut the preregistered 30 to 24. Sample, selection, medium (live human vs browser voice agent) and moderator identity all move together here.
The headline: adaptive probing is worth a lot over standardization, human moderators add affect but not content, and richer interview content does not make better predictive twins.
Interviews: what adaptive probing buys over a fixed guide#
From prose, Tables 6 and 7 (verified against the PDF):
| Measure | AI (N=139) | Static (N=154) | Human (N=24) |
|---|---|---|---|
| Speaking time (min) | 8.1 | 4.3 | 11.3 |
| Words | 952 | 609 | 1,404 |
| Turns | 30.5 | 8.3 | 32.4 |
| Needs per interview (bootstrap, N=24) | 7.48 [6.58, 8.38] | 5.01 [4.38, 5.62] | 7.08 [6.54, 7.62] |
| Unique needs, full $5K budget | 238 | 188 | 109 |
| Cost per unique need | $19 | $27 | $36 |
- Breadth and depth beat static in three of four blocks (purchase journey breadth 7.29 vs 5.90, brand awareness 3.32 vs 1.73, brand perception 5.88 vs 5.52, all p<.001). The advantage survives length control: +1.05, +1.50 and +0.28 themes after controlling for block word count (Appendix G), so it is not just longer answers. The exception is the one block the AI was told to cap at three questions, where nothing differed.
- Enjoyment and comfort are unchanged (5.65 vs 5.59, 6.37 vs 6.27, p=.70/.39). 3.7× as many turns did not cost participant experience.
- Against humans, AI covers more themes in three of four blocks (7.29 vs 6.08, 3.32 vs 1.79, 5.88 vs 5.12) at comparable depth, while humans draw more words. At a matched 24 interviews the two recover about the same number of needs (111.5 vs 109.0). Appendix I adds the caveat the headline drops: AI exceeds human in only 77% / 61% of bootstrap draws, "no reliable AI-human ordering". "Matches human" is the defensible reading and "beats human" is not.
- AI's coverage is less dispersed than the humans'. Breadth SDs are 0.91 vs 1.61 (purchase journey), 0.80 vs 0.88, and 0.35 vs 1.51 (brand perception). That is the Controlled Variance: AI's Edge as Reduced Dispersion signature in a second task, but on the confounded contrast and n=24.
- What humans keep is affect. Speech-emotion scores are higher with a live human on valence (0.528 vs 0.471), arousal (0.517 vs 0.359) and dominance (0.562 vs 0.449), all p<.001. AI and static are indistinguishable on all three. The authors read this as rapport and say sensitive or affect-laden topics may still need humans. Because the human arm differs in medium and sample, a live-call effect can't be separated from a live-human effect.
Measurement caveats the paper does not dwell on#
- The breadth coder scores against the AI's own brief. Themes are "objective-derived": each block's codebook decomposes the moderator objective, and the AI moderator was given that same objective as its goal and told to probe until it was met. The static fixed question enumerates largely the same items (Table 3), so the checklist was visible in both arms. The difference is that only the AI was instructed to keep pursuing it, and the metric counts exactly that pursuit. The length-controlled regression and the blind pairwise LLM-judge check (AI preferred on every dimension in every block, Appendix H) point the same way but don't use an independent codebook.
- The coder is not validated against humans. Five gpt-4o coders at temperature 0.7 plus a majority-vote reviewer; the only reliability number is coder-vs-reviewer agreement (98.5–99.9%), which is self-consistency, not accuracy. The coder prompt also tells the model how to treat "AI-moderated data" versus "audio-only data", so it is not blind to condition.
- Customer needs are LLM-extracted and clustered (gpt-4o, MiniLM embeddings, cosine distance < 0.60 → 252 needs). The cluster threshold sets the count.
- Internal inconsistency: Table 3's note gives the purchase-journey block eight themes, Table 7 gives max breadth 9.
Digital twins: richer input, no better prediction#
Each AI and static participant got a twin: gemini-3.1-flash-lite-preview conditioned in context on either demographics only (Demo) or demographics + transcript + individual-difference scales (Full), then asked to predict the person's held-out reactions to two mailers (images), two TV commercials (video) and two claim-ranking sets. Replicated with gpt-5-mini on the non-video tasks, with the same pattern.
- Transcripts beat demographics. Full > Demo on all four pooled metrics for static twins; for AI twins on τ-b and cosine, marginal on accuracy (.743 vs.731, p=.086), not on claim-ranking τ (p=.21).
- AI-moderated transcripts do not beat static ones. Accuracy, τ-b and τ differences are non-significant for both persona types, and so are all condition×persona interactions. Static twins are slightly better on cosine similarity. The 56% more words and 3.7× more turns that AI moderation elicited carried no additional predictive signal for these tasks.
- Twins are more deliberative than their humans. On an LLM-judged 0 (System 1) to 4 (System 2) scale, twins run +0.87 on mailers (2.18 vs 1.31), +0.63 on commercials, +0.25 on claims, all p<.001, and GPT-5-mini twins show the same gap. Humans react to what they see; twins reason from interview context about the person. Reasoning fidelity predicts accuracy on every task. Training-data relevance predicts it on claims and mailers but not commercials.
- Similar-task calibration helps a little. Adding the person's answer to the other stimulus of the same type (+1Val) lifts rating accuracy.743 →.787 (p<.001) but leaves rank agreement flat (τ-b.326 →.342, p=.78). Even giving the twin the person's own reasoning about the held-out stimulus (2Val, an upper bound, not a prediction) only reaches.815 accuracy and τ-b.376 on multimodal stimuli.
The practical split the authors draw: for need discovery, AI moderation substitutes for human moderation; for prediction, a richer transcript is not the missing input. Calibration data from closely related tasks helps modestly. What seems missing is the fast, affective reaction that neither the transcript nor the twin's reasoning style captures.
Scope: market research is not hiring#
This is the second AI-interviewer study in the vault, after Controlled Variance: AI's Edge as Reduced Dispersion (Jabarian & Henkel's 67,056-applicant hiring RCT), and the transfer runs one way only.
- Different objective. A market-research interview tries to surface as many needs as possible from someone who has no stake in the outcome. A hiring interview elicits evidence for a decision about the person talking, who has every reason to perform. Probing that draws out more customer needs has no equivalent pressure to resist.
- No screen-out function. Eligibility here was settled by screeners before the interview (at intake, for the human arm). Neither moderator could end an interview on a disqualifier, and the paper records no aborts. It therefore has nothing to say about the abort/screen-out channel in the hiring result.
- What does transfer is the decomposition. Jabarian & Henkel's AI arm differs from human recruiters in two things at once: it follows the guide more consistently, and it still adapts per applicant. This paper's randomized AI-vs-static contrast pulls those apart for one task. Pure standardization (static: identical fixed questions, zero discretion, zero aborts) is the weakest arm on every content measure. The AI's gain comes from adaptive probing within a fixed guide. So "controlled variance" is not a benefit of standardization as such; the benefit comes from the pairing of low cross-interview dispersion and within-interview responsiveness.
- A caution that also transfers, weakly. More elicited content did not raise the predictive value of the record for twins. That was a prediction task on marketing stimuli, not a human recruiter reading a transcript, so it does not contradict the hiring RCT's retention gains. It does mean "the AI collects more information" doesn't by itself imply the information predicts better.
Evidence and conflicts#
empirical, preregistered (AsPredicted #284225 for the two AI-platform arms; #290640, added later after peer feedback, for the human arm). The breadth/depth, affect and prosody comparisons are labeled exploratory in the preregistration table, and the length controls, LLM-judge check and §4.2–4.3 twin-error analyses are not preregistered. The twin comparisons (Demo vs Full, static vs AI) and the budget-matched need coverage are.
- Industry partners: an unnamed AI-interview startup supplied the platform for both AI and static arms; GBK Collective ran the human arm and is the source of the $300/h moderator price; a North American building-products brand supplied the stimuli. Toubia co-authored the HBR piece on scaling qualitative research with AI that the paper cites for the startup funding figures, and co-authored Peng et al. 2025 ("digital twins as funhouse mirrors"), whose System-2 distortion this paper reproduces.
- One category. Window and door replacement, a high-involvement, deliberative purchase, among Prolific homeowners with income ≥ $75k. The authors expect different results for low-involvement, symbolic or sensitive topics.
- Cost figures are market prices, not measured costs, and the human arm's no-shows (still paid the intake fee, still costing moderator time) inflate its per-interview cost in a way a production operation might schedule around.
Connections#
- Controlled Variance: AI's Edge as Reduced Dispersion — the hiring-interview counterpart. This page supplies a randomized split that page lacks (adaptive probing vs a fixed guide under identical conditions) and a second instance of lower AI dispersion in coverage, and it has no bearing on that page's abort/screen-out question. See Scope above
- Adaptive Probing, Not Standardization: Splitting the Information Side of the AI-Interview Effect — uses this page's randomized AI-vs-static contrast to split the hiring RCT's "information collection" channel. Computed from Table 7, the static arm is also the more dispersed in outcome (breadth CV 0.24 vs 0.12 on purchase journey), so a uniform procedure did not give uniform information. It also reads the pre-interview screener as the design that would remove the hiring study's abort confound
- Diversity Calibration Under SFT — the population-level version of the fidelity problem. That page finds a fine-tuned model's spread of survey answers converges to the human target. The twins here are in-context personas of individuals, never fine-tuned, and they miss in a specific direction: more deliberative than the person on every task and in both model families
- Configurable Human Participation — HAS-Bench's "human" is an LLM user-simulator. The twin results here are direct evidence on what such a simulator gets wrong about a specific person: it reasons more analytically (+0.25 to +0.87 on a 0–4 scale) and draws on biographical context where people react to the stimulus
- LLM-Assisted Grey-Literature Theory Building — the other end of the qualitative pipeline. There LLMs do the mechanical coding of existing text and humans keep synthesis. Here an LLM does the collection (the interviewing) and LLMs also do the coding, and the only thing the human moderator still measurably adds is affect
- Anthropic — cited as a large-scale user of AI-moderated interviews (the Anthropic Interviewer program, listed as an adjacent instrument on the AEI page)
Open Questions#
- Does the AI-over-static breadth advantage survive a codebook written independently of the moderator objective, so the AI is not scored against the checklist it was told to cover? The length-controlled and blind pairwise checks point the same way, but neither uses an independent codebook.
- Is the human arm's affect advantage (valence, arousal, dominance, p<.001) a property of a human moderator, or of a live synchronous call? A randomized design with a live human over the same browser voice channel, or a lower-latency full-duplex AI, would separate them.
Sources#
- AI-Moderated Interviews for Market Research and Digital Twins Calibration — Yuting Deng, Jingxuan Liu, Olivier Toubia (Columbia Business School) & Naman Jain (Wharton), AI-Moderated Interviews for Market Research and Digital Twins Calibration, arXiv 2609.29143 (2026-09-24; 53pp, 36 tables;
empirical). §2.1 design and the non-randomized human arm; §2.2 + Table 3 fixed question vs moderator objective; §3.1 + Tables 6–8 volume, breadth/depth, affect; §3.2 + Table 9 + Appendix I need coverage and cost; §4.1 + Table 10 Demo vs Full, static vs AI twins; §4.2 + Tables 11–12 training-data relevance, reasoning fidelity, System-1/2 gap; §4.3 + Table 13 similar-task ladder; Appendix C coder design; Appendix G length controls; Appendix H pairwise LLM judge; Appendix L GPT-5-mini replication. Parse check: Tables 6, 7, 8 and 9 are clean, and Table 7 was verified cell by cell againstpdftotext -layout(its block-label column is misaligned in the markdown but the condition labels and values are intact). Tables 4 and 5 have multi-value cells only in the metric-definition rows. Table 11's R²/Observations rows are welded ("0. 17 277") and are cited nowhere. Figures 1–2 were not opened; every count quoted comes from prose.
Cited by 8
- Adaptive Probing, Not Standardization: Splitting the Information Side of the AI-Interview Effect×6
The coder scores against the AI's own brief. The breadth codebook is the moderator objective broken…
- Controlled Variance: AI's Edge as Reduced Dispersion×3
ai moderated interviews market research digital twins — Deng, Liu, Toubia & Jain, arXiv 2609.29143…
- Configurable Human Participation
Ai Moderated Interviews — direct evidence on how an LLM standing in for a specific human diverges…
- Diversity Calibration Under SFT
Ai Moderated Interviews — the individual-level version of the silicon-sampling question, with no…
- LLM-Assisted Grey-Literature Theory Building
Ai Moderated Interviews — the collection end of the same pipeline, handed to an LLM. Here LLMs code…
- AI Economics & Labor
Ai Moderated Interviews — Deng, Liu, Toubia & Jain (arXiv 2609.29143): a preregistered…
- Open Questions Dashboard
Controlled Variance: How much of the +12% is controlled variance in information collection versus…
- Open Questions Backlog
Ai Moderated Interviews ×2 (oldest 4d) — Does the AI-over-static breadth advantage survive a…
Related articles
- Controlled Variance: AI's Edge as Reduced Dispersion
Jabarian & Henkel (arXiv 2607.28222): a pre-registered natural field experiment randomizing 70,884 job applicants betwe…
- Procedural Value in AI Decisions
Wang, Sturgis & de Kadt (LSE/Cornell, arXiv 2609.16390): a preregistered paired-profile conjoint on 1,919 US job seeker…
- Open Questions Backlog
Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…
- The Solo-Authorship Rebound
Matsui (arXiv 2607.10780): across 300M+ OpenAlex works and 26 fields, the decades-long decline in solo-authored papers…
- What the Instrument Can Resolve: Two Headline Numbers and Their Missing Denominators
Answers two open questions as one: the +12% AI-interview offer effect and ATLAS's 22.6% classifier accuracy are both ra…
