Howardism · Vol. 03Plate I · No. 01
Evaluation, tagged.
Notes52TagEvaluationOldest8 May 2026Newest1 Oct 2026
Every article tagged evaluation, newest first.
- C01Verbalized-Confidence Soft Scoring for LLM JudgesLLM As A JudgeCalibrationVerbalized Confidence+2· 16′
- C02· 18′
- C03· 20′
- C04· 16′
- C05· 13′
- C06· 15′
- C07· 55′
- C08· 18′
- C09· 68′
- C10Unsanctioned Agent Message BoardsSecurityMulti AgentCovert Channel+3· 50′
- C11Skill LiftEvaluationAgent SkillsAblation+3· 25′
- C12AI-Assisted Error AnalysisEvalsEvaluationError Analysis+3· 16′
- C13· 16′
- C14· 17′
- C15GDPval BenchmarkEvaluationBenchmarksEconomic Impact+1· 39′
- S16· 34′
- C17· 51′
- C18· 44′
- C19· 60′
- C20· 45′
- C21· 53′
- C22· 28′
- C23· 52′
- C24· 35′
- C25· 73′
- C26· 48′
- C27· 21′
- C28· 40′
- C29· 127′
- C30· 39′
- C31· 26′
- C32· 52′
- C33· 14′
- C34· 20′
- C35· 7′
- C36· 17′
- C37· 24′
- C38· 18′
- C39· 22′
- C40· 46′
- C41· 86′
- C42· 25′
- C43· 41′
- E44· 4′
- C45· 58′
- C46· 29′
- C47· 67′
- C48Production-Sourced EvaluationEvaluationBenchmarksData Provenance+1· 32′
- C49· 35′
- C50· 26′
- C51· 65′
- C52· 27′