Sections

Search

ArXiv paper measures how much benchmark contamination actually inflates scoresGoogle's Gemini 3.8 TTS models are cheap and handle multi-voice dialogue, per Simon WillisonUniDataAgent: China Unicom's Ontology-Grounded Enterprise Q&A Agent Cuts Report Time from Days to MinutesOpenAI says Harvey uses GPT-6 Astra to produce more structured legal draftsArXiv paper proposes auditable LLM labeling for classroom talk
All stories

Research·

FrontierMath Erdős benchmark: GPT-6 Astra solves 3% of open conjectures, all others score zero

ArXiv CS.CL preprint introduces FrontierMath Erdős, a Lean-based benchmark of 68 open Erdős problems, and reports GPT-6 Astra resolved 3% with all other tested models at 0%.

A benchmark built from open problems

> ArXiv CS.CL researchers introduced FrontierMath Erdős (FME), a benchmark of 68 Erdős problems that were still open as of August 2026.> To count as solved, an AI must prove or disprove the conjecture in the Lean proof assistant.> The problems were chosen by the paper's second author from 652 open problems listed on erdosproblems.com, selected for mathematical interest and difficulty.> The paper argues that recent AI resolutions of open problems are isolated demonstrations rather than a systematic study of capability — FME is meant to close that gap by evaluating every model on the same fixed problems, autonomously and under the same budget. [1]

Results under a $300-per-problem budget

> Five AIs were evaluated with a budget of $300 per problem.> One model, GPT-6 Astra, scored 3%; every other model scored 0%.> Caveat: this is the paper's own evaluation protocol and scoring, and only five models were tested, so the 3% vs. 0% spread is a narrow snapshot rather than a broad ranking of AI math ability. The abstract does not break down per-problem results or error modes. [1]

Sources

  1. ArXiv CS.CL (Computation and Language) · Reporting ·
    FrontierMath Erd\H{o}s