[Video Credit: N-body simulation Francisco Villaescusa-Navarro]
Carolina Cuesta-Lazaro
Doubling time every 4 months
"AI can't do X"
9.11 > 9.9
3 b's in blueberry
Navier Stokes
88 hours / 10k agents
Huggingface
> 55 websites
+O(100) others?
(3 hours of compute on average)
LLMs for Dummies
Many Worlds: RLVR for Physics
An Optimistic Future, if we build it
1) Pretraining:
fill in the gaps
Learning a prior
2) Postraining:
Reinforcement Learning from Verifiable Rewards
Updating the prior to conform with evidence provided by the reward
X: Sequence of tokens. Basic units in a finite vocabulary (100k)
Language Model:
un
believ
able
+
+
unbelievable =
Neural Network Weights
Gradient descent
True
Reconstructed
["Joint cosmological parameter inference and initial condition reconstruction with Stochastic Interpolants" Cuesta-Lazaro, Bayer, Albergo et al NeurIPs 2024 ML for the Physical Sciences]
Physics students at Penn are
...
OVER-CAFFEINATED
RESILIENT
SMART
ATHLETIC
The entire internet
# parameters
# tokens
The biggest lesson that can be read from 70 years of AI research is that general methods that leverage computation are ultimately the most effective, and by a large margin. [...]
methods that continue to scale with increased computation even as the available computation becomes very great. [...]
We want AI agents that can discover like we can, not which contain what we have discovered.
TD-GAMMON
1992
2013
DQN
2016
AlphaGo
AlphaGo
ChatGPT
(RLHF)
2022
Reasoning (RLVR)
2025
Hide And Seek
2019
Agent
Environment
State
Action
Reward
Policy
Maximise!
Policy
Trajectory
State
Action
Evidence: Reward Model
Posterior: RL-ed model
p(behaviour|correct)
Maximise reward
Stay close to original LM
Variational Inference:
p(behaviour)
Prior: Original Language Model
Mathematics (and coding) are hard to solve but easy to verify
Navier Stokes: No human understood the proof before we knew it was correct
Accuracy on Maths Problems
# RL steps
Response Length
# RL steps
Empirical Models that are predictive
Predictiveness
Mechanistic theories that generalize
Conceptual Understanding
DiscoverPhysics
Lean? Verifiable Rewards?
Many Worlds: RL in simulated worlds
Propose Experiment
Data
Analyse
Hypothesis
arxiv:2605.26087
Experimentation, hypothesis generation, model selection...
Matt Wiemann
Lindsay Smith
Hypothesis
Simulate World
Invisible particles
Extra dimensions
Multi Species ...
Simulate World
Invisible particles
Extra dimensions
Multi Species ...
Propose Experiment
Simulate
Text: Conceptual Understanding
Trajectories (.csv)
Science Agent (LLM)
Science Agent (LLM)
Outputs
Python Code: Trajectory MSE
["DiscoverPhysics: Benchmarking LLMs for Out-of-the-Box Scientific Thinking" Wiemann, Smith et al (including CCL)]
Matt Wiemann
Lindsay Smith
22 Curated Worlds
Pavel Izmailov
Andrew Gordon-Wilson
Peter Melchior
def discovered_law(pos1, pos2, p1, p2, velocity2, duration, **params):
"""Particle 2 is accelerated toward particle 1 by a radial force
per unit along r^ that combines a static source term G*p1,
a radial-velocity term (analogous to an advective/retardation
coupling), and a centripetal-like tangential kinetic-energy
term, all scaled by 1/r and independent of p2."""
...
return trajectories
def fit_parameters():
return {
"G": {"init": 1.0, "bounds": [0.1, 5.0]},
...
}
Predictive Code
Rubric Based LLM judge
Mean Square Error on test trajectories
Agent Explanation
How Predictive the Theory is
Conceptual Understanding
You are an expert physicist tasked with discovering scientific laws in a simulated universe. Your goal is to propose experiments, analyse the data they return, and ultimately deduce the underlying scientific law. Note that the laws of physics in this universe may differ from those in our own. You can perform experiments to gather data but must follow the protocol strictly.
If F = G/r²: G = 80. If F = G/r: G = 8. Linear in 1/r!
m=1: a = 0.664 · m=2: a = 0.614 · m=4: a = 0.664 · m=8: a = 0.714 … That's strange."
Consistent pattern: small +y 'extra' force (~0.05) regardless of position … maybe the anchor itself moves.
a_y_extra ≈ +0.05 everywhere on probes. Spot on! So anchor too just experiences g_y = 0.05 … g = (0, 0.05) is a uniform external acceleration on every particle.
You are an expert physicist tasked with discovering scientific laws in a simulated universe. Your goal is to propose experiments, analyse the data they return, and ultimately deduce the underlying scientific law. Note that the laws of physics in this universe may differ from those in our own. You can perform experiments to gather data but must follow the protocol strictly.
Confirmed: a ∝ 1/r.
observed −5.18. WAY too small. Something is very wrong with my understanding!
artifact of integration accuracy, NOT the true law. Probably because of close encounters causing chaos.
I'll just submit my law. The functional form is solid.
You are an expert physicist tasked with discovering scientific laws in a simulated universe. Your goal is to propose experiments, analyse the data they return, and ultimately deduce the underlying scientific law. Note that the laws of physics in this universe may differ from those in our own. You can perform experiments to gather data but must follow the protocol strictly.
a × r seems roughly constant … not quite 1/r.
Let me try a × r²: … Not consistent.
So at large r, a ~ 1/r. At small r, a ~ 1/r².
Actually note the file path in the error message: 'extra_dimensions.csv'! This is a hint! The world might have an extra compactified dimension…
[Model starts fitting Yukawa potential, Bessel functions... At some point fitting tool errors.]
Conceptual Understanding
Trajectory Prediction
Random Seeds
Verifiable Rewards: Ground truth physics of the simulator
1) Can we generate a large enough dataset of good quality randomized worlds with interesting physics?
2) What should be the model's reward?
Predictive
easy, MSE
Conceptual understanding?
World Generator
World Solver
def simulate(
pos1,
pos2,
duration,
**params,
):
"Simulate Universe"
return trajectories
Convergence,
Re-implementation tests....
def discovered_law(pos1, pos2, p1, p2, velocity2, duration, **params):
"""Particle 2 is accelerated toward particle 1 by a radial force
per unit along r^ that combines a static source term G*p1,
a radial-velocity term (analogous to an advective/retardation
coupling), and a centripetal-like tangential kinetic-energy
term, all scaled by 1/r and independent of p2."""
...
return trajectories
def fit_parameters():
return {
"G": {"init": 1.0, "bounds": [0.1, 5.0]},
...
}
Running Experiment...
Reward
Solve the task
Hard, but Solvable
(Solver)
(Generator)
"This world consists of ..."
World Definition
Simulation Code
Attempting the task
Proposing a new task
Solver: Qwen3-4B
Generator: Claude Opus
(no grads)
Predictiveness: Trajectory MSE in test held out trajectories
Insight / Conceptual Understanding:
Rubric: World generator makes a point system akin to the one in our benchmark
Adversarial MSE: Find experiments that would disproof the solver's theory -> discourages effective theories with a large regime of validity
Vortex in Uniform wind
Attractor / Repulsor
Hidden lattice
Striped magnetic field
How Predictive
Conceptual Understanding
p(behaviour)
p(behaviour|correct)
Scientific Discovery
Binary Stars
Structured Reasoning
Coding
Observation
Question
Hypothesis
Testable Predictions
Gather data
Alter, Expand, Reject Hypothesis
Develop General Theories
[Figure adapted from ArchonMagnus] "AI can't do X"
What is a valuable contribution to the community?
What are our goals?
Work at a new level of abstraction
The end of hyperspecialization
The end of corporate academia
Community driven
Access to frontier compute concentrates
Breath without depth
Incremental work (5 minutes of Claude)
hyperspecialization
grants
Academia keeps doing same, disconnects form AI progress and becomes irrelevant
Longer horizons, more ambitous projects
Models capable of autonomos work
Resource allocation and credit assignment
PROMISES
RISKS
coding is solved
taste may be the scarce resource
safety concerns
live open science
p(behaviour)
p(behaviour|correct)