Sections

Search

ArXiv paper measures how much benchmark contamination actually inflates scoresGoogle's Gemini 3.8 TTS models are cheap and handle multi-voice dialogue, per Simon WillisonUniDataAgent: China Unicom's Ontology-Grounded Enterprise Q&A Agent Cuts Report Time from Days to MinutesOpenAI says Harvey uses GPT-6 Astra to produce more structured legal draftsArXiv paper proposes auditable LLM labeling for classroom talk
All stories

Models·

ArXiv paper measures how much benchmark contamination actually inflates scores

arXiv preprint LeakScale separates benchmark contamination's presence from its score impact, testing 262,144 generations.

What happened

> An arXiv CS.CL preprint (arXiv:2609.27176v1) proposes LeakScale, described as an interventional framework for estimating how much a benchmark score owes to exposure to evaluation material. The paper, from unidentified authors in the supplied evidence, argues provenance can establish contact but only a counterfactual can quantify performance attributable to that contact. [1]

How LeakScale works, as described

> The method creates fresh executable tasks that require private, family-specific information absent from, and non-derivable from, the public task. It controls access to that information and estimates the resulting control-adjusted change in executable accuracy. [1]

What the experiment found

> Across 2,048 unique families, two model families, two executable domains, and 262,144 generations, exposure improved accuracy in every model-by-domain combination, with gains ranging from +7.17 to +27.31 percentage points, according to the paper. These are the authors' reported results from their own framework, not an independent replication. [1]

Why it matters

> The paper frames two questions that are often conflated: whether benchmark contact occurred, and how strongly a reported score depends on it. LeakScale is presented as making the second question directly measurable. [1]

Sources

  1. ArXiv CS.CL (Computation and Language) · Reporting ·
    Beyond Overlap: Estimating the Causal Effect of Benchmark Exposure