Search stories

All stories

OpenAI Releases Framework for Reporting Model MisalignmentVerbose Prompts Help Vision-Language Models Resist Image CorruptionOpenAI Announces Astra for Law for Firm WorkflowsResearchers Stress-Test Alignment Midtraining Across 110-Billion-Parameter AI ModelsOpenAI Research Studies How Workers Integrate AI Beyond Traditional RolesGoogle DeepMind Launches Institute to Widen Debate on AGIAI Evaluator Forum Releases AEF-1 Standard for Third-Party AI EvaluationsGood Start Labs Trains AI on Railroad Strategy Game to Boost Finance Benchmark ScoresAIUC Raises $40M Series A and Launches Insured Agent Standard AIUC-1
All stories

AI Evaluator Forum Releases AEF-1 Standard for Third-Party AI Evaluations

According to Latent Space, the AI Evaluator Forum has published AEF-1, establishing a baseline standard for independent third-party AI evaluations as frontier labs face growing calls for external oversight.

AEF-1 Baseline and Embedded Evaluators

The AI Evaluator Forum published AEF-1 to set baseline expectations covering access, conflicts of interest, funding relationships, recusal, and transparency. Alongside this framework, Anthropic committed to hosting embedded third-party evaluators such as METR, promising internal-style access including office desks, access badges, and company laptops to verify safety adherence and examine training pipelines. [1]

Industry Split Over Pacing and Governance

The release of the standard comes amid sharp disagreement across the industry over pacing progress versus control-first safety. Former Google DeepMind researcher Bilal Chughtai warned that progress may be outrunning alignment and advocated for pacing and transparency, while Cohere's Aidan Gomez argued against allowing a small group of Silicon Valley companies to become AI gatekeepers for governments. [1]

Sources

  1. 01