H
Howardism
Plate IIEvals & BenchmarksHOWARDISM

Construct or Roster? What BEI's Entanglement Graph and IRT's Ability Ranking Are Measuring

Two #oq/now items answered as one: both read shared structure off a model × item matrix while conditioning on something fitted from the same roster, so what the structure means depends on who is in the roster. (1) BEI's top tier is a tier effect, not a family effect. Kuai's own Table C.2 is 10 of the 15 pairs among six 2023–24 open-weight checkpoints, with cross-family Llama↔Qwen pairs at ranks 2, 3, 6 and 10, so vintage, weakness and base-vs-instruct status are collinear and cannot be separated on that roster. Frontier panels whose difficulty is conditioned from outside the roster (Kohli, Hossain) keep the dependence and do not organise it by vendor. (2) No page tests θ̂ against accuracy as a predictor of an external outcome. ATLAS's convergence row compares θ̂ with θ̂ on one calibration population. The only out-of-sample latent-vs-raw comparison is Zhu's, one level up: a single factor loses badly to a crude mean (R² 0.110 vs 0.771) and three factors win narrowly (0.808). Both questions stay open, retagged #oq/source, with the settling experiment specified

Article metadata
Publication details
Published:October 1, 2026
Filed:Essay
Domain:Evals & Benchmarks
Reading:18 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Construct or Roster? What BEI's Entanglement Graph and IRT's Ability Ranking Are Measuring

Sources#

Question#

Two #oq/now items, answered as one synthesis because they share a failure mode:

  1. Cross-Model Error Entanglement: Is the BEI graph measuring entanglement or vintage? Its top 10 pairs are 2023–24 Llama and Qwen models. Difficulty is fitted from the other 16 models, so a uniformly weak pair looks co-failing everywhere. Would the intra-family signal survive a single-release-window roster, or residualisation on release date?
  2. Item Response Theory for LLM Benchmarks: Does the ability ranking θ̂ predict anything downstream that accuracy does not? The validity evidence is internal and convergent only. The falsification test: rank models by θ̂ and by accuracy on benchmark A, then see which ranking predicts benchmark B.

Short answer#

Both instruments read shared structure off a models × items matrix, and both condition on a quantity fitted from the same roster. BEI conditions on difficulty d_t, estimated from the other 16 models. IRT conditions on item parameters b, a, c, estimated from the calibration population. So the structure each one finds is defined relative to the roster, and three roster variables in the vault can masquerade as the construct: vintage (Zhu), family (Zheng & Yang) and capability (Li).

  • Q1: a tier effect, not a family effect, and the roster cannot say which tier variable drives it. The paper's own top-10 table does not have the intra-family shape the question assumes. It is a clique of six old open-weight checkpoints, and cross-family edges sit among the intra-family ones. Vintage, weakness and base-vs-instruct status all coincide in those six models. The evidence from outside the roster points one way. Every frontier panel that conditioned on difficulty from outside itself kept its dependence, and none organised it by vendor. The head-on rerun is still unrun. Stays open; partial.
  • Q2: the vault has no test of it. Everything ATLAS offers is θ̂ compared against θ̂, calibrated on one population. The vault's only out-of-sample comparison of a latent score with a raw average is Zhu's, made at the level of models × benchmarks. There the dominant factor predicts held-out benchmarks far worse than a crude mean, and three factors predict them slightly better. That sets a prior, not an answer. Stays open; partial by analogy.

Neither question can be advanced further by synthesising pages already in the vault. Both are retagged #oq/source, with the settling experiment specified below.

The shared mechanism: the conditioning variable comes from the roster#

InstrumentWhat it conditions onWhere the conditioning variable comes fromRoster variable that can leak into the construct
BEI / CIG (Kuai et al.)task difficulty d_tfailure rate among the other M − 2 modelscapability and vintage: a weak pair has high fitted p_m(d) everywhere
3PL θ̂ (ATLAS)item b, a, c~4,000 Open LLM Leaderboard modelsvintage (the temporal holdout degrades ability MAE ~50%) and family (DIF)
Factor scores (Zhu)common factorthe 96-model gridvintage: F1 tracks release date at logistic R² = 0.505
Family DIF (Zheng & Yang)8–128-dim spectral abilitythe RouterEval populationshows family is a residual that survives the adjustment
Kish n_eff (Kohli)per-item difficulty100 human annotators (ChaosNLI entropy), not the panelnone from the panel: the conditioning is external
Ledoit-Wolf ρ̄ (Hossain et al.)item difficultya different bank (open-weight judges)none from the frontier roster: the conditioning is external
Measurement Layouts (Prunty et al.)item demandrubric annotation, never a response matrixmoves to the level intercept, which is catalogue-relative

The last three rows are the constructive pattern. An instrument escapes its roster only when its conditioning variable is estimated outside that roster. Prunty's case shows the escape is partial: the dependence on the catalogue moves to the capability level instead of disappearing. That pattern organises both answers below.

Q1 — What BEI's top tier actually is#

What Table C.2 shows#

The raw's Table C.2 (A Statistical Framework for Auditing Behavioral Dependence and Induced Bias in LLM Judges) lists the top 10 BEI pairs. Every one of them falls inside a set of six models: Llama-2-70b-hf, Llama-3-70B, Llama-3.1-70B, Qwen1.5-110B, Qwen1.5-72B-Chat and Qwen1.5-14B-Chat. Six models give 15 pairs, and the top 10 are 10 of those 15. The ranking, with family class:

RankPairBEIClass
1Llama-3-70B ↔ Llama-3.1-70B0.0525intra
2Llama-2-70b ↔ Qwen1.5-110B0.0398cross
3Llama-3-70B ↔ Qwen1.5-110B0.0352cross
4Qwen1.5-14B-Chat ↔ Qwen1.5-72B-Chat0.0333intra
5Llama-2-70b ↔ Llama-3-70B0.0326intra
6Llama-3.1-70B ↔ Qwen1.5-110B0.0325cross
7Llama-2-70b ↔ Llama-3.1-70B0.0290intra
8Qwen1.5-110B ↔ Qwen1.5-72B-Chat0.0266intra
9Qwen1.5-110B ↔ Qwen1.5-14B-Chat0.0245intra
10Llama-3.1-70B ↔ Qwen1.5-14B-Chat0.0211cross

Four of the ten edges are cross-family, and they are not in the tail. Llama-2-70b ↔ Qwen1.5-110B is second, above every intra-family pair except the first, which is two adjacent releases of the same base weights. So inside the clique, family does not order the edges; tier does. The paper's prose ("strong intra-family entanglement (e.g., within the LLaMA family)") overstates what its own table shows. The question's premise, that there is an intra-family signal whose survival is in doubt, inherits that overstatement. (Wiki reading of the paper's table. The paper does not make this comparison.)

Three roster variables are collinear in the clique#

  • Vintage. All six models were released in 2023–24. No closed-weight model appears, and neither do the roster's newest releases (GPT-5, GPT-oss-20B, Claude 4.6 Sonnet, Gemini-3.1-Pro).
  • Capability. They are the roster's weak tier. This is exactly where the question's worry bites: with d_t fitted from the other 16 models, a uniformly weak pair has a high fitted failure probability everywhere.
  • Post-training status. The paper's roster convention marks instruction-tuned variants with "Instruct" or "Chat" (§B.1). By that convention, four of the six are base checkpoints: Llama-2-70b-hf, Llama-3-70B, Llama-3.1-70B and Qwen1.5-110B. The one contrast in the roster that holds family and release fixed is Llama-3.1-70B base against Llama-3.1-70B-Instruct. The base appears in four of the top 10 pairs; the Instruct variant appears in none. That is suggestive only: absence from a top-10 list is not a measured value. But it is the only within-release comparison the paper supplies, and it varies post-training rather than family or date.

A single-release-window restriction, which is what the question proposes, would not separate these three on this roster. A 2024-H1 window still contains the Qwen1.5 trio and Llama-3-70B, and they are still the weak, mostly-base tier within that window. Its one cross-family pair, Llama-3-70B ↔ Qwen1.5-110B (0.0352), outranks the intra-family Qwen1.5-14B ↔ Qwen1.5-72B (0.0333). So a window restriction would keep a cross-family edge on top and still leave capability and base status unresolved.

What the other instruments add#

  • Dependence survives on rosters with no vintage spread, conditioned from outside. Kohli's nine current frontier judges show φ̄ = 0.391. The majority vote falls 22.0pp short of independence even after per-item difficulty is conditioned on human annotation entropy, which comes from outside the panel. The most-correlated pairs are cross-family, and the same-family excess is only +0.047. Hossain, Yousefi & Lim's GPT-5.6-sol / Claude Opus 5 / Grok 4.5 panel correlates at 0.42 across providers against 0.40 within Gemini. Adjusting for difficulty taken from a different bank moves it only from 0.465 to 0.467. So once the conditioning variable no longer comes from the roster, dependence is still present and still does not follow vendor lines. Both are different instruments (raw verdict φ, shrunk error ρ̄) on different tasks, not BEI.
  • Family-specific item behaviour exists beyond capability in the open-weight tier. Zheng & Yang find item × family residuals that survive an 8–128-dimensional ability adjustment, far richer than BEI's single d_t. They replicate across owner-disjoint halves (median ρ =.308–.589), and they survive an owner cap of 1 and a filter removing merges, distills and LoRAs. So "family" is a real variable in open-weight response matrices, not only a proxy for weakness. It is DIF, though, not co-failure. Its sign flips across benchmarks for every family. And the paper tests no vintage variable (no release-date control appears in the raw). It shows family can matter, not that it is what BEI's top tier measures.
  • Capability structures judge error too. Li finds a frozen judge's task-conditioned false acceptance rising with target capability (mean Spearman +0.82 across 35 SWE-bench agents), with no lineage variable involved. This bears on Kuai's judge-bias half, not on BEI's answerer-side graph. But it is the vault's clearest evidence that capability, one of the three collinear variables above, is a live confound on the same kind of pairwise matrix.

Verdict and the settling experiment#

On the current evidence, BEI's top tier is best read as a tier effect in which vintage, weakness and base-model status are indistinguishable. It is not an intra-family effect: the table interleaves cross-family edges with intra-family ones. Three independent frontier instruments with externally sourced difficulty keep dependence without any vintage spread, which is evidence against pure vintage. Nothing in the vault separates weakness from base status.

The experiment that would settle the question has three parts, and none of them is in the vault. Rerun BEI on a roster that, within one release window, crosses (a) family with (b) base vs instruct variants of the same weights, spanning (c) a capability range. Then residualise the pair residuals on the product of the two models' accuracies before ranking. The intra-family reading survives only if same-family pairs still exceed matched cross-family pairs after that. Retagged #oq/source.

Q2 — Does θ̂ predict anything accuracy does not?#

Every piece of evidence in the vault is internal or θ̂-against-θ̂#

EvidenceWhat is comparedWhy it is not the posed test
ATLAS split-half (Table 4)θ̂ on one half against θ̂ on the otherinternal reliability; accuracy is already ≥ 0.94
ATLAS cross-benchmark (0.439 → 0.904 on MMLU chemistry)rank(θ̂_A) vs rank(θ̂_B), against rank(acc_A) vs rank(acc_B)both sides are θ̂ calibrated on the same population, so a shared scaling artifact inflates both. The target is not independent of the scoring rule under test
ATLAS reordering examples (0.713 vs 0.714 → ranks 270 vs 2,612)θ̂ rank vs accuracy rankZheng & Yang show any principled reweighting flips 30.9–47.1% of sub-point pairs, and a random half-subtest flips 11.4–34.7%, so a reordering is expected and is not validity evidence
Prunty et al. (GPT-4o-mini: last on accuracy, fourth on inferred capability)inferred level vs accuracy"Nothing is validated against an observed outcome" (Cognitive Capability Profiling for Task Suitability)
She & Lin (MD-2PL 1.03pp MAE vs Rasch 1.37pp)IRT reconstruction of the same benchmark's full scorewithin-benchmark prediction; the target is accuracy on unadministered items, not an external outcome (Recurring Production Agent Evaluation)
Desai et al. 1PL ΔAUCshared vs separate trait, held-out item cellsa dimensionality diagnostic, not a θ̂-vs-accuracy comparison (Benchmark Convergent and Discriminant Validity)

The ATLAS convergence row is the closest to the test, and it carries a roster problem of its own. The raw says IRT separates models most in the tails. At the low end, accuracy compresses into a 0.10–0.15 band while θ spans about −3 to −1. The Open LLM Leaderboard population is dense in exactly that weak tail, and a Spearman over ~4,000 models puts most of its weight where most of the models are. So part of the 0.439 → 0.904 gain may be θ̂ spreading out a tail that accuracy ties at chance level. That would be a real gain in resolution, but it would sit where this roster is densest and could shrink on a frontier-only roster. Neither ATLAS nor anyone else reports the convergence gain by ability stratum. (Wiki inference from ATLAS's reported compression band and population. Untested.)

The one out-of-sample latent-vs-raw comparison, one level up#

Zhu's leave-one-benchmark-out ladder is the only place in the vault where a latent-structure score and a raw average compete to predict held-out measurements. Factors are re-estimated inside each fold:

Predictor of a held-out economic benchmarkOut-of-fold R²
mean of the other eleven standardised scores0.771
first factor only (74.5% of common variance)0.110
three factors0.808 (ΔMSE vs mean +0.037 [+0.019, +0.055])

Two readings carry over to θ̂:

  1. A dominant latent dimension can predict worse than the crude sum it was supposed to improve on. The factor-score projection discards the level information that the mean keeps. θ̂ within one benchmark does not share that exact failure, because it is a discrimination-weighted monotone function of the responses and keeps level. So Zhu's rung (iii) is a warning about what "latent" can cost, not a prediction that θ̂ loses.
  2. The latent reading won only with more than one dimension, and only narrowly. The vault's dimensionality evidence points the same way within benchmarks. Zheng & Yang's held-out selection picks K = 8–128 on five item banks, two of them ATLAS's own. She & Lin's multidimensional 2PL beats Rasch at every budget. A unidimensional 3PL θ̂ may discard exactly the structure an external target needs.

This is models × benchmarks on a 96-model frontier grid, not θ̂ against accuracy within a benchmark. It says that "a latent score beats the raw score out of sample" is not free and not ruled out. It does not say which way the posed test comes out.

The test, specified so a roster artifact cannot win it#

The question's own test is right. The vault adds four constraints, each from a page that shows the failure mode:

  1. Score benchmark B by accuracy, not θ̂. Otherwise both sides share calibration and scaling (the ATLAS convergence row's problem).
  2. Exclude near-ties or model them as noise. Zheng & Yang show that sub-point orderings flip under any reweighting, so neither scoring can "win" them.
  3. Residualise on release date and stratify by family before comparing the two predictors. Zhu: F1 tracks date at R² 0.505. Zheng & Yang: family DIF replicates. Without this, θ̂ can beat accuracy by tracking the roster instead of the construct.
  4. Report the comparison by ability stratum. That shows whether any θ̂ advantage lives only in the dense low-ability tail of an Open-LLM-Leaderboard-style population.

The data exist: ATLAS's response matrices on five benchmarks, and RouterEval's matrices, which Zheng & Yang used. That makes this a re-analysis, not a new collection, but nobody in the vault has run it. Retagged #oq/source.

Tensions and contradictions found#

  • Kuai et al.'s prose against their own Table C.2. The text calls BEI's leading structure intra-family ("strong intra-family entanglement (e.g., within the LLaMA family)"; Appendix C: "concentrated primarily among closely related model families"). The table has cross-family Llama↔Qwen pairs at ranks 2, 3, 6 and 10.
  • Item Response Theory for LLM Benchmarks's Connections bullet on Cross-Model Error Entanglement says that page finds excess co-failure "concentrated in exactly the intra-family fine-tune pairs that make up the Open LLM Leaderboard population". Two parts of that are inaccurate. Four of BEI's top 10 are cross-family. And Kuai's roster holds official base and chat releases, not community fine-tunes. The analogy to the leaderboard population still holds at the level of "older open-weight tier". Left for the next lint pass to reword.
  • Kohli against Hossain on the sign of the family gap (same-family +0.047 above cross-family, against cross 0.42 ≥ within 0.40). This is already recorded on Cross-Model Error Entanglement. Both are small and of the same order, and both support "vendor is not the decorrelating variable", so they are consistent in conclusion.

Sources#

§ end
Cited by 4
Related articles
  • Benchmark Score Redundancy

    Zeng & Papailiopoulos: an 84-model × 133-benchmark public score matrix is effectively rank-2, so BenchPress matrix comp…

  • Item Response Theory for LLM Benchmarks

    Score a benchmark with a 3PL IRT model instead of percent-correct and two separable things follow. Scoring: ability θ r…

  • Measuring Beyond Accuracy Saturation

    Princeton-led case study (arXiv 2606.26158): accuracy saturation is not benchmark saturation — re-instrument a saturate…

  • Evals & Benchmarks

    Map of Content for the evals-and-benchmarks domain — 36 concepts. The science of measuring models: benchmark validity,…

  • Open Questions Backlog

    Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…