What happened
> An arXiv CS.CL preprint (arXiv:2609.27176v1) proposes LeakScale, described as an interventional framework for estimating how much a benchmark score owes to exposure to evaluation material. The paper, from unidentified authors in the supplied evidence, argues provenance can establish contact but only a counterfactual can quantify performance attributable to that contact. [1]
How LeakScale works, as described
> The method creates fresh executable tasks that require private, family-specific information absent from, and non-derivable from, the public task. It controls access to that information and estimates the resulting control-adjusted change in executable accuracy. [1]
What the experiment found
> Across 2,048 unique families, two model families, two executable domains, and 262,144 generations, exposure improved accuracy in every model-by-domain combination, with gains ranging from +7.17 to +27.31 percentage points, according to the paper. These are the authors' reported results from their own framework, not an independent replication. [1]
Why it matters
> The paper frames two questions that are often conflated: whether benchmark contact occurred, and how strongly a reported score depends on it. LeakScale is presented as making the second question directly measurable. [1]
Sources
- Beyond Overlap: Estimating the Causal Effect of Benchmark Exposure
ArXiv CS.CL (Computation and Language) · Reporting ·