Discover Many Worlds:

RL for Scientific Discovery in Self Generated Worlds

[Video Credit: N-body simulation Francisco Villaescusa-Navarro]

Carolina Cuesta-Lazaro

 

Flatiron Institute

NYU

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

Doubling time every 4 months

"AI can't do X"

9.11 > 9.9

3 b's in blueberry

Navier Stokes

88 hours / 10k agents

Huggingface

> 55 websites

+O(100) others?

(3 hours of compute on average)

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

1

LLMs for Dummies

Many Worlds: RLVR for Physics

2

An Optimistic Future, if we build it

3

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

1) Pretraining:

fill in the gaps

p(\mathrm{behaviour})

Learning a prior

p(\mathrm{behaviour}|\mathrm{success})

2) Postraining:

Reinforcement Learning from Verifiable Rewards

Updating the prior to conform with evidence provided by the reward

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

X: Sequence of tokens. Basic units in a finite vocabulary (100k)

Pretraining a Language Model

p_\theta \left(x \right)

Language Model:

un

believ

able

+

+

unbelievable =

Neural Network Weights

Gradient descent

True

Reconstructed

\delta_\mathrm{Obs}
\delta_\mathrm{ICs}
p(\delta_\mathrm{ICs}, \theta|\delta_\mathrm{Obs})

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

Same models we use to infer the random seed of our Universe!

["Joint cosmological parameter inference and initial condition reconstruction with Stochastic Interpolants" 
Cuesta-Lazaro, Bayer, Albergo et al 
NeurIPs 2024 ML for the Physical Sciences]

 

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

Pretraining: Next token prediction

\mathcal{L}(\theta) = -\,\mathbb{E}_{x \sim \mathcal{D}} \left[ \sum_{t=1}^{T} \log p_\theta \left(x_t \mid x_{\lt t} \right) \right]

Physics students at Penn are

...

OVER-CAFFEINATED

RESILIENT

SMART

ATHLETIC

The entire internet

Scaling Laws

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

L(N, D) = E + \frac{A}{N^{\alpha}} + \frac{B}{D^{\beta}}
E \approx 1.69
A \approx 406.4
\alpha \approx 0.34
\beta \approx 0.28

# parameters

# tokens

The bitter lesson by Rich Sutton

The biggest lesson that can be read from 70 years of AI research is that general methods that leverage computation are ultimately the most effective, and by a large margin. [...]

 

methods that continue to scale with increased computation even as the available computation becomes very great. [...]

 

We want AI agents that can discover like we can, not which contain what we have discovered.​

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

TD-GAMMON

1992

2013

DQN

2016

AlphaGo

AlphaGo

ChatGPT

(RLHF)

2022

Reasoning (RLVR)

2025

Hide And Seek

2019

Agent

Environment

State

Action

s_t
s_{t+1}
a_t \sim \Pi_\theta(s_t)

Reward

R_t

Policy

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

Maximise!

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

J(\theta) = \mathbb{E}_{\tau \sim \Pi_\theta}\left[ R(\tau) \right]
\qquad R(\tau) = \sum_{t} R_t, \qquad \tau = (s_0, a_0, s_1, a_1, \dots)

Policy

Trajectory

State

Action

\nabla_\theta J = \int R(\tau)\, \nabla_\theta p_\theta(\tau)\, d\tau
\nabla_\theta J = \nabla_\theta \int R(\tau)\, p_\theta(\tau)\, d\tau
\nabla_\theta p_\theta(\tau) = p_\theta(\tau)\, \nabla_\theta \log p_\theta(\tau)
= \int R(\tau)\, p_\theta(\tau)\, \nabla_\theta \log p_\theta(\tau)\, d\tau
= \mathbb{E}_{\tau \sim \Pi_\theta}\left[ R(\tau)\, \nabla_\theta \log p_\theta(\tau) \right]

Reinforce

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

Post-training for LMs

\mathbb{E}_{x \sim \Pi_\theta}\left[ R(x) \right] - \beta \mathrm{KL}(\Pi_\theta, \Pi_0)

Evidence: Reward Model

\exp \left( r(x) \right)

Posterior: RL-ed model

\Pi^\star (x) \propto \Pi_0(x) \exp \left( r(x) \right)

p(behaviour|correct)

Maximise reward

Stay close to original LM

\mathrm{KL}(\Pi_\theta, \Pi^\star)

Variational Inference:

\Pi_0 (x)

p(behaviour)

Prior: Original Language Model

Mathematics (and coding) are hard to solve but easy to verify

Navier Stokes: No human understood the proof before we knew it was correct

DeepSeek R1: RL with Verifiable Rewards

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

x, y > 1, \qquad \log_x\left(y^x\right) = \log_y\left(x^{4y}\right) = 10. \qquad \text{Find } xy.
xy = 25

DeepSeek R1: RL with Verifiable Rewards

Accuracy on Maths Problems

# RL steps

Response Length

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

# RL steps

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

  • Increase in thinking time
  •  Autonomously (unsupervised) develops advanced reasoning strategies: reflective reasoning and systematic exploration of alternative solutions 
  • Sudden increase in the use of the word "wait" during reflections, change in reasoning patterns

Teaching LMs to discover physics outside of their box

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

Empirical Models that are predictive

Predictiveness

Mechanistic theories that generalize

Conceptual Understanding

  • Benchmark that tests the entire scientific method

DiscoverPhysics

  • Can we train language models to be better phycists?

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

Lean? Verifiable Rewards?

Many Worlds: RL in simulated worlds

Propose Experiment

Data

Analyse

Hypothesis

Can pretrained models discoveyr new physics?

arxiv:2605.26087

Experimentation, hypothesis generation, model selection...

Matt Wiemann

Lindsay Smith

Hypothesis

Simulate World

Invisible particles

Extra dimensions

Multi Species ...

Simulate World

Invisible particles

Extra dimensions

Multi Species ...

Propose Experiment

Simulate

Text: Conceptual Understanding

Trajectories (.csv)

Science Agent (LLM)

Science Agent (LLM)

Outputs

Python Code: Trajectory MSE

["DiscoverPhysics: Benchmarking LLMs for Out-of-the-Box Scientific Thinking" Wiemann, Smith et al (including CCL)]

Matt Wiemann

Lindsay Smith

DiscoverPhysics: Benchmarking LLMs for Out-of-the-Box Scientific Thinking

22 Curated Worlds

Pavel Izmailov

Andrew Gordon-Wilson

Peter Melchior


def discovered_law(pos1, pos2, p1, p2, velocity2, duration, **params):
    """Particle 2 is accelerated toward particle 1 by a radial force 
    per unit along r^ that combines a static source term G*p1,
    a radial-velocity term (analogous to an advective/retardation
    coupling), and a centripetal-like tangential kinetic-energy 
    term, all scaled by 1/r and independent of p2."""
    ...
    return trajectories

def fit_parameters():
    return {
        "G":     {"init": 1.0, "bounds": [0.1, 5.0]},
        ...
    }

Predictive Code

Rubric Based LLM judge

Mean Square Error on test trajectories

Agent Explanation

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

How Predictive the Theory is

Conceptual Understanding

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

Adversarial Experiments to Test Generalization

You are an expert physicist tasked with discovering scientific laws in a simulated universe. Your goal is to propose experiments, analyse the data they return, and ultimately deduce the underlying scientific law. Note that the laws of physics in this universe may differ from those in our own. You can perform experiments to gather data but must follow the protocol strictly.

If F = G/r²: G = 80. If F = G/r: G = 8. Linear in 1/r!

 

m=1: a = 0.664 · m=2: a = 0.614 · m=4: a = 0.664 · m=8: a = 0.714 … That's strange."

 

Consistent pattern: small +y 'extra' force (~0.05) regardless of position … maybe the anchor itself moves.

 

a_y_extra ≈ +0.05 everywhere on probes. Spot on! So anchor too just experiences g_y = 0.05 … g = (0, 0.05) is a uniform external acceleration on every particle.

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

You are an expert physicist tasked with discovering scientific laws in a simulated universe. Your goal is to propose experiments, analyse the data they return, and ultimately deduce the underlying scientific law. Note that the laws of physics in this universe may differ from those in our own. You can perform experiments to gather data but must follow the protocol strictly.

Confirmed: a ∝ 1/r.

 

 

observed −5.18. WAY too small. Something is very wrong with my understanding!

 

 

artifact of integration accuracy, NOT the true law. Probably because of close encounters causing chaos.

 

I'll just submit my law. The functional form is solid.

 

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

You are an expert physicist tasked with discovering scientific laws in a simulated universe. Your goal is to propose experiments, analyse the data they return, and ultimately deduce the underlying scientific law. Note that the laws of physics in this universe may differ from those in our own. You can perform experiments to gather data but must follow the protocol strictly.

a × r seems roughly constant … not quite 1/r.

 

Let me try a × r²: … Not consistent.

 

 

So at large r, a ~ 1/r. At small r, a ~ 1/r².

   

Actually note the file path in the error message: 'extra_dimensions.csv'! This is a hint! The world might have an extra compactified dimension…

 

[Model starts fitting Yukawa potential, Bessel functions... At some point fitting tool errors.]

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

Conceptual Understanding

Trajectory Prediction

Random Seeds

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

Teaching LMs to discover physics outside of their box

Verifiable Rewards: Ground truth physics of the simulator

1) Can we generate a large enough dataset of good quality randomized worlds with interesting physics?

2) What should be the model's reward?

Predictive

easy, MSE

Conceptual understanding?

World Generator

World Solver


def simulate(
	pos1, 
    pos2, 
    duration, 
    **params,
  ):
	"Simulate Universe"

    return trajectories

Convergence,

Re-implementation tests....


def discovered_law(pos1, pos2, p1, p2, velocity2, duration, **params):
    """Particle 2 is accelerated toward particle 1 by a radial force 
    per unit along r^ that combines a static source term G*p1,
    a radial-velocity term (analogous to an advective/retardation
    coupling), and a centripetal-like tangential kinetic-energy 
    term, all scaled by 1/r and independent of p2."""
    ...
    return trajectories

def fit_parameters():
    return {
        "G":     {"init": 1.0, "bounds": [0.1, 5.0]},
        ...
    }

Running Experiment...

Reward

Solve the task

Hard, but Solvable

(Solver)

(Generator)

"This world consists of ..."

World Definition

Simulation Code

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

Asymmetric Self Play

Asymmetric Self Play

Attempting the task

Proposing a new task

The Solver Reward

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

Solver: Qwen3-4B

Generator: Claude Opus

(no grads)

Predictiveness: Trajectory MSE in test held out trajectories

Insight / Conceptual Understanding: 

Rubric: World generator makes a point system akin to the one in our benchmark

Adversarial MSE: Find experiments that would disproof the solver's theory  -> discourages effective theories with a large regime of validity

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

Self Generated Worlds

Vortex in Uniform wind

Attractor / Repulsor

Hidden lattice

Striped magnetic field

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

How Predictive

Conceptual Understanding

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

Evaluating Capabilities

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

Testing Out of Distribution

p(behaviour)

p(behaviour|correct)

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

Scientific Discovery

Binary Stars

Structured Reasoning

Coding

Testing Out of Distribution

Observation

Question

Hypothesis

Testable Predictions

Gather data

Alter, Expand, Reject Hypothesis

Develop General Theories

[Figure adapted from ArchonMagnus] 

The Scientific Method in > 2025

p(\mathcal{M}|x)

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

"AI can't do X"

  • "AI can't do X" based on 6 months ago models, too noisy to help us plan for the future.

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

We should pay attention to Mathematics

  • Very disorienting for the community.
  • Opportunity to redefine. 

What is a valuable contribution to the community?

What are our goals?

Work at a new level of abstraction 

The end of hyperspecialization

The end of corporate academia

Community driven

Access to frontier compute concentrates

Breath without depth

Incremental work (5 minutes of Claude)

hyperspecialization

grants

 

Academia keeps doing same, disconnects form AI progress and becomes irrelevant

Longer horizons, more ambitous projects

Models capable of autonomos work 

Resource allocation and credit assignment

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

The Future of Physics: Promises and Risks

PROMISES

RISKS

coding is solved

 taste may be the scarce resource

safety concerns

live open science

Carolina Cuesta-Lazaro Flatiron/NYU @ Penn 2026

p(behaviour)

p(behaviour|correct)