Sections

Search

ArXiv paper measures how much benchmark contamination actually inflates scoresGoogle's Gemini 3.8 TTS models are cheap and handle multi-voice dialogue, per Simon WillisonUniDataAgent: China Unicom's Ontology-Grounded Enterprise Q&A Agent Cuts Report Time from Days to MinutesOpenAI says Harvey uses GPT-6 Astra to produce more structured legal draftsArXiv paper proposes auditable LLM labeling for classroom talk
All stories

Models·

ArXiv preprint proposes 'R3Con' harness to cut reliance on model scale

An ArXiv preprint proposes R3Con, a harness for building principled context representations, claiming it beats baselines on large-corpus reasoning benchmarks and lets smaller models outperform larger ones at lower cost.

What the preprint claims

> An ArXiv CS.CL preprint proposes design principles, drawn from the cognitive theory of relevance realization, for AI systems that build representations of very large contexts, and introduces R3Con, a harness meant to apply those principles systematically. The authors argue that existing approaches — graphs, textual memories, retrieval collections — succeed or fail according to how well they match those principles, and that representation design is currently largely ad hoc. — ArXiv CS.CL [1]

Reported results, with caveats

> Against nine state-of-the-art baselines on two recent benchmarks of reasoning over large document corpora, the authors report R3Con outperforming the strongest baseline by 20 and 8.4 percentage points. They also report that R3Con with 4B and 9B models beats all evaluated 35B baselines, and that R3Con with a 35B-A3B model beats Claude Code with Claude-Sonnet-5 at 3.7x lower cost. — ArXiv CS.CL, preprint claims; benchmark conditions and baselines are the authors' own, not independent evaluations. The paper states code is available on GitHub. [1]

Sources

  1. ArXiv CS.CL (Computation and Language) · Reporting ·
    Realize What Matters: Principled Context Representation for Large-Scale Reasoning