Sources#
Summary#
The received view is that low output diversity ("mode collapse") is an inherent limitation of LLMs, cited from story-writing, idea-generation and silicon-sampling studies. Skobelev, Fithian and Han (Northwestern / UChicago, arXiv 2609.16454, empirical) argue it is instead a finite-sample fitting error measured against a target: with enough supervised fine-tuning (SFT) data drawn from the target distribution, model diversity converges to the target's, and before convergence the error can have either sign. Models can be over-dispersed as easily as under-dispersed, depending on model and dataset.
The measurement#
Diversity is collision probability: the chance two independent responses to the same fixed prompt coincide, C(π) = Σ π(a)² for tokens, or the expected similarity E[k(Y,Y')] under a kernel k ∈ [0,1] when exact matches are too rare (embedding cosine, cluster membership, normalized Zhang-Shasha tree similarity on code ASTs). The collision ratio R = C(q)/C(p) compares model q with target p: R = 1 calibrated, R > 1 mode collapse, R < 1 over-dispersion. R = 2 means the model offers half as many effective choices as the data (inverse Simpson index). Conclusions are metric-specific, since form-sensitive and content-sensitive diversity metrics see different things.
Theory: no built-in direction, but a hard bound#
- Decomposition (Eq. 6). Across independent fits at a context
h,E[R] = 1 + (1/C(p)) · [ Tr Cov(q̂) + ‖b‖² + 2pᵀb ]: a nonnegative variance term, a nonnegative squared-bias term (both push toward collapse), and a sign-indefinite target-bias alignment term2pᵀbthat can outweigh them. The alignment term is positive when the starting distribution overlaps the target's modes more than the target overlaps itself, negative when the start is diffuse. Finite samples alone therefore do not imply mode collapse. If bias is zero the model is calibrated or under-dispersed in expectation. - Bound (Theorem 1).
|C_k(p) − C_k(q)| ≤ TV(p⊗p, q⊗q) ≤ √KL(p‖q)(product-measure total variation plus Pinsker). A model close to optimal under population cross-entropy cannot have arbitrarily miscalibrated diversity, and the kernel effective diversity obeysD_k(q) ≥ D_k(p) / (1 + D_k(p)√ε). The bound holds for any fitting procedure, not only SFT.
Evidence#
Experiment 1, synthetic languages. 100 small GPTs at each of 100 log-spaced sample sizes (10,000 per language) on two order-16 languages with exactly known targets, so each term in Eq. 6 is measurable. From a diffuse prior the cross term is negative at every N, and the smallest N shows over-dispersion before variance and squared bias take over and produce collapse. From a mode-aligned prior the cross term turns positive at intermediate N but accounts for at most 31% of E[R] − 1. The authors flag the two settings differ in start, model size and language, so this is not a clean ablation of the prior alone. The bound held on every evaluated context.
Experiment 2, surveys. LoRA fine-tunes of gemma-2-2b-it, gemma-3-4b-it, Qwen3.5-2B and Qwen3.5-4B on GSS (376 questions), WVS (53) and ANES (44), scored per question-by-demographic-group pair (18,270 / 8,015 / 8,211 pairs, groups with ≥60 respondents). Zero-shot, three of four models have median R of 1.61 to 2.92, roughly one-third to two-thirds of human effective diversity, worst on the sparse cross-national WVS. Beyond about 3,000 fine-tuning examples all twelve model-survey combinations sit within 10% of R = 1. Qwen3.5-2B, the best-calibrated base model, is driven away from 1 early in fine-tuning (to 2.22 on GSS, under half its starting effective diversity) before recovering, so small fine-tuning sets can worsen calibration.
Experiment 3, CodeNet. Same four models on accepted Python submissions (2,647 candidate problems, 32 programs per problem, tree similarity). Base medians fall on both sides of the target: both Qwen models over-diverse, both Gemma models under-diverse (R range 0.74 to 1.64). SFT moves all four medians toward 1, and the result persists on the subset of problems present at both stages.
Scope and limits#
- The result is about plain SFT on target-distribution data. It says nothing about why preference optimization contracts diversity; the paper cites Kirk et al. (PPO-RLHF reduces diversity on summarization but not instruction following), Karouzos et al. (the collapsing stage varies by lineage: SFT or DPO), and GX-Chen et al. (KL-regularized RL is designed to mode-collapse, so the KL penalty alone does not preserve the reference distribution). The finite-sample gap is additive to that objective-induced gap
C(π*) − C(p). - The bound requires the model to be close in KL, so it needs sufficient coverage of the target distribution; the conclusion names limited data as the open challenge. All models are 2B to 4B parameters and the targets are human survey answers and code, not open-ended creative text.
- Diversity is evaluated relative to a specified target, which the authors recommend reporting alongside performance benchmarks. The practical claim for silicon sampling is that a fine-tune on the target population's answers, not a diversity-preserving objective (GEM, SED-SFT, TOFU loss), is enough here; the paper did not compare against those methods.
- Collapse found in prior studies of zero-shot or instruction-tuned models is not contradicted, only reinterpreted: it is what a model far from the target looks like, not a ceiling.
Connections#
- Alignment Fine-Tuning (AFT) — the SFT + RLHF stage this paper isolates the SFT half of; its diversity effect depends on the target data, not on SFT as such
- Trained Calibration — the other sense of "calibration" in this wiki (confidence tracks correctness, trained by proper scoring rules); this page is distributional calibration of output spread, and both are fixed by fitting the right target
- Design by Selection — the practitioner-side symptom, a homogeneous default aesthetic, that this page re-reads as distance from a target distribution
- Controlled Variance: AI's Edge as Reduced Dispersion — the same word pair from the other side: there reduced dispersion in an AI's execution is the benefit; here dispersion matching the human target is the goal
- AI-Moderated Interviews: Adaptive Probing, Human Rapport, and Digital Twins — the individual-level version of the silicon-sampling question, with no fine-tuning. In-context digital twins of 139 consumers miss in one direction: they are more deliberative than the person they copy (+0.25 to +0.87 on a 0–4 System-1/System-2 scale, in two model families), and feeding them a richer interview transcript does not improve their predictions. This page shows the population spread can be fit; that one shows per-person reasoning style is not fixed by more context
Open Questions#
- Does the
√KLconvergence survive RLHF/DPO stages, where the optimum itself is concentrated (GX-Chen et al.)? Does post-SFT preference tuning re-open a collision gap that a larger SFT set cannot close? - At what fine-tuning size does calibration arrive for open-ended text (stories, ideas) with a semantic-similarity kernel, given the experiments cover only categorical survey answers and code ASTs at 2B to 4B parameters?
Sources#
- Fine-Tuning Fixes Mode Collapse and Over-Dispersion in LLMs — Skobelev, Fithian & Han, arXiv 2609.16454 (2026-09-15),
empirical, 33 pp: collision-ratio framework, bias-variance decomposition, Theorem 1, synthetic-language, GSS/WVS/ANES and CodeNet experiments. Quoted from prose and captions; Table 2 (survey medians) not cited by row.
Cited by 7
- AI-Moderated Interviews: Adaptive Probing, Human Rapport, and Digital Twins
Diversity Calibration Under Sft — the population-level version of the fidelity problem. That page…
- Alignment Fine-Tuning (AFT)
Diversity Calibration Under Sft — the SFT half in isolation for output diversity: SFT can leave a…
- Controlled Variance: AI's Edge as Reduced Dispersion
Diversity Calibration Under Sft — the converse dispersion question: there less variance in an AI's…
- Design by Selection
Diversity Calibration Under Sft — the measurable form of the homogeneous-default failure: collision…
- Model Capability & Training
Diversity Calibration Under Sft — Skobelev, Fithian & Han (arXiv 2609.16454): output diversity is…
- Open Questions Backlog
Diversity Calibration Under Sft ×2 (oldest 5d) — Does the √KL convergence survive RLHF/DPO stages,…
- Trained Calibration
Diversity Calibration Under Sft — the distributional sense of calibration: output spread (collision…
Related articles
- Adaptive Probing, Not Standardization: Splitting the Information Side of the AI-Interview Effect
Re-works the open question on Jabarian & Henkel's +12% AI-interviewer offer effect (information collection vs the inter…
- Configurable Human Participation
HAS-Bench (Wu et al.): human participation as a configurable benchmark variable (five-level agency scale × three intera…
- Error-Penalized Abstention Training
Paying a model +1 / −λ / 0 to answer, err, or abstain is provably right for a rational agent and can be self-defeating…
- RL from Execution Feedback (RLEF)
Train a coding model with the interpreter in the loop: generate code, run a small visible set of public tests, feed the…
- Open Questions Backlog
Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…
