Date: August 7th, 2026 7:20 PM
Author: The Penis
How wide is the gap in iq between phd theoretical physicists and mathematicians and claude fable or open ai astra:
The gap is already very wide in the domains that matter most for the claim.
On the specific achievements
OpenAI’s internal Astra model (the one that produced the ten results) solved long-standing open problems in pure mathematics and theoretical computer science—non-sofic groups, Connes rigidity, improved sphere-packing bounds, several ErdÅ‘s problems, etc.—and supplied machine-checkable Lean proofs. These were not routine exercises; they had resisted serious human effort for decades in some cases. Claude Fable 5 (Anthropic’s Mythos-class model) has shown comparable strength on hard formal and scientific tasks, including at least one notable counterexample in algebraic geometry.
An average PhD in mathematics or theoretical physics is a competent specialist. Most never produce a result that resolves a multi-decade open problem, and many never work at that frontier at all. Solving even one such problem is typically the work of elite researchers (often near or at Fields/Abel level). Doing ten of them, with formal verification, puts Astra well into the extreme right tail of human mathematical ability—beyond the large majority of working research mathematicians, let alone the average PhD.
How to think about the size of the gap
Narrow formal reasoning and theorem-proving: the models are already superhuman relative to the average PhD and competitive with (or ahead of) the strongest human specialists in the subfields they attacked.
Breadth, taste, and research direction-setting: still more mixed. Humans still dominate long-horizon scientific taste, choosing which problems are worth solving, and integrating results into a broader intellectual program.
Everything outside high-end formal domains (physical intuition, embodied skills, social navigation, open-ended engineering under real-world constraints): humans retain large advantages for now.
So if we restrict the comparison to the kind of high-end mathematical and theoretical work these models just demonstrated, the gap is not incremental. It is the difference between a solid professional and a system that can generate novel, correct, field-advancing results that most PhDs will never approach. As the models improve further on agency, verification, and cross-domain synthesis, that gap is more likely to widen than shrink for the median PhD
(http://www.autoadmit.com/thread.php?thread_id=5890512&forum_id=2E#50052299)