GPT-6 Astra Low beats GPT-5.6 Sol High in benchmark
In a recent five-run benchmark of Codex agents, GPT-6 Astra Low outperformed GPT-5.6 Sol High, completing the task in 7 minutes and 55 seconds with only 33 tool calls, compared to 14 minutes and 24 seconds and 92 tool calls for its predecessor. This efficiency marks a significant advancement in agent performance, highlighting the improvements in the latest model. Martin Lanczi shared the details of this competitive performance. That's @martinlanczi.
Published · last aired