Chapter 05
Science
Stanford presents science as a mixed zone: publication volume is rising quickly, domain models are emerging and selected pipelines are now AI-native, but replication and expert-level autonomous research remain hard.
Why it matters
This chapter matters because it shows where AI is becoming embedded in scientific method rather than only in paper-writing or search assistance.
How to apply it
Read it as a maturity map by discipline: where models help, where they automate and where they still fail to reproduce scientific work.
Core signals
Core signals
Scientific AI is expanding on some fronts and stagnating on others — and its institutional base remains anchored in academia and government rather than the industry-heavy model driving general AI.
AI for science is growing, but not uniformly
Publication growth and new foundation-model infrastructure are real, but benchmark results show that difficult end-to-end scientific tasks remain far from solved.
Scientific AI is still institutionally different from general AI
Stanford notes that most science models still come from academic and government settings rather than the industry-heavy pattern seen in general-purpose AI.
Selected data points
Selected data points
Condensed numbers and comparisons pulled from the official Stanford chapter materials.
| Signal | Value | Context |
|---|---|---|
| Natural-science AI publications | 80,150 | Natural sciences reached about 80,150 AI publications in 2025, up 26% from 2024. |
| AI share of research output | 5.8% to 8.8% | Depending on the field, AI now accounts for 5.8% to 8.8% of scientific output, up from below 1% in 2010. |
| AION-1 training base | 200M objects | AION-1 was trained on more than 200 million celestial objects across five major surveys. |
| FourCastNet 3 speed | <4 min / 8x-60x | FourCastNet 3 can generate a 60-day global forecast in under four minutes, 8 to 60 times faster than prior approaches. |
| PaperArena agent vs PhD baseline | 38.8% vs 83.5% | On end-to-end scientific research tasks, the best agent scored 38.8% versus 83.5% for PhD experts. |
Implications
- Scientific impact should be judged by workflow integration, not only model novelty.
- Field-specific benchmarks remain essential because generic benchmark success does not transfer cleanly.
- Academic and public institutions still matter disproportionately in AI-for-science progress.