OpenAI quietly revised several performance benchmarks for its new GPT-6 Astra flagship model. These changes occurred just days after the model's September 3 launch.
Fortune first reported the post-launch edits. The adjustments generally improved Astra’s performance metrics against its predecessor and rival models.
One revision briefly lowered Astra’s hallucination rate from 4.2% to 2% before OpenAI reverted the figure. The company also modified a cybersecurity score based on a reasoning tier not yet commercially available.
OpenAI stated the changes reflect the best available performance estimates. The revisions have triggered industry scrutiny regarding the reliability of launch-day AI evaluations.