Gemini 4 Argon outperforms GPT-6 Astra, Claude Fable 5.1, and Claude Opus 5.5 on a majority of benchmarks in knowledge work, agentic coding, long-context tasks, and multimodal understanding, according to Google. In some categories, such as Harvey’s Legal Agent Benchmark and Vals Finance Agent v2, Argon’s lead is substantial. Yet the model trails in areas including FrontierSWE v2 (agentic coding), Terminal-bench 4.0, and Terminal-Bench Science 0.1, where it falls behind competitors by up to 10 points.
The announcement comes a day after Google CEO Sundar Pichai co-signed a commitment with executives from Anthropic, Meta, Nvidia, OpenAI, and SpaceX to “self-police” AI development, following a meeting with President Donald Trump. The commitment lacks an enforcement mechanism.
Google says it will gather feedback from early testers in its Fairwind Program and iterate on guardrails before expanding access. Paid API customers and AI Ultra subscribers will get access first, followed by developers, enterprises, and consumers.