Discover Many Worlds:
RL for Scientific Discovery in Self Generated Worlds
[Video Credit: N-body simulation Francisco Villaescusa-Navarro]
Carolina Cuesta-Lazaro
Flatiron Institute
NYU
Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026
Doubling time every 4 months
"AI can't do X"

9.11 > 9.9
3 b's in blueberry
Navier Stokes
88 hours / 10k agents
Huggingface
> 55 websites
+O(100) others?
(3 hours of compute on average)
Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026
LLMs for Dummies

Many Worlds: RLVR for Physics

An Optimistic Future, if we build it

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

1) Pretraining:
fill in the gaps
Learning a prior
2) Postraining:
Reinforcement Learning from Verifiable Rewards
Updating the prior to conform with evidence provided by the reward
Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026
X: Sequence of tokens. Basic units in a finite vocabulary (100k)
Pretraining a Language Model
Language Model:

un
believ
able
+
+
unbelievable =
Neural Network Weights
Gradient descent

True
Reconstructed

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026
Same models we use to infer the random seed of our Universe!
["Joint cosmological parameter inference and initial condition reconstruction with Stochastic Interpolants" Cuesta-Lazaro, Bayer, Albergo et al NeurIPs 2024 ML for the Physical Sciences]
Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026
Pretraining: Next token prediction
Physics students at Penn are
...
OVER-CAFFEINATED
RESILIENT
SMART
ATHLETIC
The entire internet
Scaling Laws
Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026
# parameters
# tokens

The bitter lesson by Rich Sutton
The biggest lesson that can be read from 70 years of AI research is that general methods that leverage computation are ultimately the most effective, and by a large margin. [...]
methods that continue to scale with increased computation even as the available computation becomes very great. [...]
We want AI agents that can discover like we can, not which contain what we have discovered.
Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026


TD-GAMMON
1992

2013
DQN
2016
AlphaGo
AlphaGo
ChatGPT
(RLHF)
2022
Reasoning (RLVR)
2025
Hide And Seek
2019

Agent


Environment
State


Action

Reward
Policy
Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026
Maximise!
Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026
Policy
Trajectory
State
Action
Reinforce
Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026
Post-training for LMs

Evidence: Reward Model

Posterior: RL-ed model
p(behaviour|correct)
Maximise reward
Stay close to original LM
Variational Inference:
p(behaviour)
Prior: Original Language Model

Mathematics (and coding) are hard to solve but easy to verify
Navier Stokes: No human understood the proof before we knew it was correct
DeepSeek R1: RL with Verifiable Rewards
Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026


DeepSeek R1: RL with Verifiable Rewards
Accuracy on Maths Problems
# RL steps
Response Length
Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026
# RL steps
Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026
- Increase in thinking time
- Autonomously (unsupervised) develops advanced reasoning strategies: reflective reasoning and systematic exploration of alternative solutions
- Sudden increase in the use of the word "wait" during reflections, change in reasoning patterns
Teaching LMs to discover physics outside of their box
Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026
Empirical Models that are predictive
Predictiveness
Mechanistic theories that generalize
Conceptual Understanding




- Benchmark that tests the entire scientific method
DiscoverPhysics
- Can we train language models to be better phycists?
Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026
Lean? Verifiable Rewards?
Many Worlds: RL in simulated worlds
Propose Experiment


Data
Analyse

Hypothesis
Can pretrained models discoveyr new physics?
arxiv:2605.26087
Experimentation, hypothesis generation, model selection...

Matt Wiemann

Lindsay Smith
Hypothesis
Simulate World

Invisible particles
Extra dimensions
Multi Species ...
Simulate World

Invisible particles
Extra dimensions
Multi Species ...
Propose Experiment


Simulate
Text: Conceptual Understanding
Trajectories (.csv)

Science Agent (LLM)

Science Agent (LLM)

Outputs
Python Code: Trajectory MSE
["DiscoverPhysics: Benchmarking LLMs for Out-of-the-Box Scientific Thinking" Wiemann, Smith et al (including CCL)]

Matt Wiemann
Lindsay Smith

DiscoverPhysics: Benchmarking LLMs for Out-of-the-Box Scientific Thinking
22 Curated Worlds



Pavel Izmailov
Andrew Gordon-Wilson
Peter Melchior
def discovered_law(pos1, pos2, p1, p2, velocity2, duration, **params):
"""Particle 2 is accelerated toward particle 1 by a radial force
per unit along r^ that combines a static source term G*p1,
a radial-velocity term (analogous to an advective/retardation
coupling), and a centripetal-like tangential kinetic-energy
term, all scaled by 1/r and independent of p2."""
...
return trajectories
def fit_parameters():
return {
"G": {"init": 1.0, "bounds": [0.1, 5.0]},
...
}

Predictive Code
Rubric Based LLM judge
Mean Square Error on test trajectories
Agent Explanation
Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026



How Predictive the Theory is
Conceptual Understanding
Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026
Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

Adversarial Experiments to Test Generalization
You are an expert physicist tasked with discovering scientific laws in a simulated universe. Your goal is to propose experiments, analyse the data they return, and ultimately deduce the underlying scientific law. Note that the laws of physics in this universe may differ from those in our own. You can perform experiments to gather data but must follow the protocol strictly.

If F = G/r²: G = 80. If F = G/r: G = 8. Linear in 1/r!
m=1: a = 0.664 · m=2: a = 0.614 · m=4: a = 0.664 · m=8: a = 0.714 … That's strange."
Consistent pattern: small +y 'extra' force (~0.05) regardless of position … maybe the anchor itself moves.
a_y_extra ≈ +0.05 everywhere on probes. Spot on! So anchor too just experiences g_y = 0.05 … g = (0, 0.05) is a uniform external acceleration on every particle.

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026
You are an expert physicist tasked with discovering scientific laws in a simulated universe. Your goal is to propose experiments, analyse the data they return, and ultimately deduce the underlying scientific law. Note that the laws of physics in this universe may differ from those in our own. You can perform experiments to gather data but must follow the protocol strictly.
Confirmed: a ∝ 1/r.
observed −5.18. WAY too small. Something is very wrong with my understanding!
artifact of integration accuracy, NOT the true law. Probably because of close encounters causing chaos.
I'll just submit my law. The functional form is solid.


Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026
You are an expert physicist tasked with discovering scientific laws in a simulated universe. Your goal is to propose experiments, analyse the data they return, and ultimately deduce the underlying scientific law. Note that the laws of physics in this universe may differ from those in our own. You can perform experiments to gather data but must follow the protocol strictly.
a × r seems roughly constant … not quite 1/r.
Let me try a × r²: … Not consistent.
So at large r, a ~ 1/r. At small r, a ~ 1/r².
Actually note the file path in the error message: 'extra_dimensions.csv'! This is a hint! The world might have an extra compactified dimension…


[Model starts fitting Yukawa potential, Bessel functions... At some point fitting tool errors.]
Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026




Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026
Conceptual Understanding
Trajectory Prediction
Random Seeds

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026
Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026
Teaching LMs to discover physics outside of their box
Verifiable Rewards: Ground truth physics of the simulator
1) Can we generate a large enough dataset of good quality randomized worlds with interesting physics?
2) What should be the model's reward?
Predictive
easy, MSE
Conceptual understanding?
World Generator

World Solver
def simulate(
pos1,
pos2,
duration,
**params,
):
"Simulate Universe"
return trajectories

Convergence,
Re-implementation tests....
def discovered_law(pos1, pos2, p1, p2, velocity2, duration, **params):
"""Particle 2 is accelerated toward particle 1 by a radial force
per unit along r^ that combines a static source term G*p1,
a radial-velocity term (analogous to an advective/retardation
coupling), and a centripetal-like tangential kinetic-energy
term, all scaled by 1/r and independent of p2."""
...
return trajectories
def fit_parameters():
return {
"G": {"init": 1.0, "bounds": [0.1, 5.0]},
...
}

Running Experiment...
Reward
Solve the task
Hard, but Solvable
(Solver)
(Generator)
"This world consists of ..."
World Definition
Simulation Code
Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026
Asymmetric Self Play
Asymmetric Self Play

Attempting the task
Proposing a new task
The Solver Reward
Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026
Solver: Qwen3-4B
Generator: Claude Opus
(no grads)
Predictiveness: Trajectory MSE in test held out trajectories
Insight / Conceptual Understanding:
Rubric: World generator makes a point system akin to the one in our benchmark
Adversarial MSE: Find experiments that would disproof the solver's theory -> discourages effective theories with a large regime of validity
Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026


Self Generated Worlds
Vortex in Uniform wind
Attractor / Repulsor
Hidden lattice
Striped magnetic field

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

How Predictive
Conceptual Understanding

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

Evaluating Capabilities

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

Testing Out of Distribution

p(behaviour)
p(behaviour|correct)


Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026
Scientific Discovery
Binary Stars
Structured Reasoning
Coding
Testing Out of Distribution
Observation
Question
Hypothesis
Testable Predictions
Gather data
Alter, Expand, Reject Hypothesis
Develop General Theories
[Figure adapted from ArchonMagnus] The Scientific Method in > 2025
Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026
Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

"AI can't do X"

- "AI can't do X" based on 6 months ago models, too noisy to help us plan for the future.
Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026
We should pay attention to Mathematics
- Very disorienting for the community.
- Opportunity to redefine.
What is a valuable contribution to the community?
What are our goals?
Work at a new level of abstraction
The end of hyperspecialization
The end of corporate academia
Community driven
Access to frontier compute concentrates
Breath without depth
Incremental work (5 minutes of Claude)
hyperspecialization
grants
Academia keeps doing same, disconnects form AI progress and becomes irrelevant
Longer horizons, more ambitous projects
Models capable of autonomos work
Resource allocation and credit assignment
Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026
The Future of Physics: Promises and Risks
PROMISES
RISKS
coding is solved
taste may be the scarce resource
safety concerns
live open science

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

p(behaviour)
p(behaviour|correct)

Penn - 2026
By carol cuesta
Penn - 2026
- 78