How Can AI Help Us Understand the Universe?



Justine Zeghal
Mila, Université de Montréal
Cosmology From Home 2026, Online
Cosmological Inference
Cosmological Inference
Goal: Get the value of the cosmological parameters

with the uncertainty!

Bayes theorem:
To study the nature of Dark Energy
Cosmological Inference


Cosmological Inference


The power spectrum is near gaussian so we have an approximation of the likelihood
Cosmological Inference
The power spectrum is near gaussian so we have an approximation of the likelihood

DES Y3 WL Results (with SBI).
The power spectrum is not a sufficient statistics for non gaussian field
Cosmological Inference






Stage III
Stage IV
Portion of the Virgo cluster, zoom on RSCG 55
Portion of the Virgo cluster, zoom on RSCG 55
Access to new non gaussian small scales. We don't want to lose this new information!
Field-Level Inference
Field-Level Inference
Field-Level Inference
Simulator

Initial conditions
Large Scale Structure

Prediction
Inference
Explicit inference
Needs an explicit simulator to sample the joint posterior through MCMC:
Implicit inference
We use simulations to
learn
Instead of relying on an analytical model to describe the phenomenon, we can simulate it.
Two ways of performing inference from simulations:
Field-Level Inference
Because we work at the map level, considering all cosmological information, we call this inference:
Field-level inference / Pixel-level inference / Full-field inference
Most precise inference!
Explicit inference
Needs an explicit simulator to sample the joint posterior through MCMC:
Implicit inference
We use simulations to
learn
Instead of relying on an analytical model to describe the phenomenon, we can simulate it.
Two ways of performing inference from simulations:


Implicit Inference
Implicit Inference
From the dataset we can learn:
- the marginal likelihood
- the posterior
- the likelihood to evidence ratio
Since the goal of implicit inference is to approximate a distribution, it has greatly benefited from the advent of generative models.
Implicit Inference
Definition:
A generative model is a machine learning model designed to create new data that is similar to its training data .
Different kinds of generative models:
- Generative Adversarial Networks (GANs)
- Variational Autoencoders (VAEs)
- Flow Models (Normalizing Flows, Flow Matching, Diffusion Models, Stochastic Interpolants)
- etc.



Implicit Inference



Implicit Inference



Implicit Inference



Implicit Inference



Implicit Inference



Implicit Inference



Implicit Inference



Change of Variable Formula:
Should be easy to compute
Where is the NN?
Implicit Inference
Where is the NN?
For instance, for affine transformation (RealNVP, Dinh et al. 2017):
NN
It is such a specific function.. Why?
Easy to inverse
Easy to compute the determinant of the Jacobian
Implicit Inference
Credit: François Lanusse
We need a tool to compare distributions:
the Kullback-Leiber Divergence
How to train the NNs?
Implicit Inference
Implicit Inference
We want to minimize the Kullback-Leiber Divergence wrt
Implicit Inference
We want to minimize the Kullback-Leiber Divergence wrt
Implicit Inference
We want to minimize the Kullback-Leiber Divergence wrt
Implicit Inference
Simulations only!
Change of variable formula
We want to minimize the Kullback-Leiber Divergence wrt
Implicit Inference
Implicit Inference
Likelihood approximation (e.g. Papamakarios et al., 2019)
Implicit Inference
Likelihood approximation (e.g. Papamakarios et al., 2019)
Posterior approximation (e.g. Papamakarios et al., 2016)
Implicit Inference
Implicit Inference
https://simulation-based-inference.org/by Kyle Cranmer and Jason Lo
Implicit Inference
E.g. of results on Data

Reminder: Implicit Inference enables approximating the posterior when the likelihood is unknown but implicitly encoded in simulations.
Summary statistics are not always Gaussian distributed.

wavelet phase harmonic
scattering transform coefficients
Computing the covariance matrix of a combination of summary statistics is challenging.
Adding more summary statistics tightens the constraints
Implicit Inference
E.g. of results on Data


NN-based summaries from the maps + Cl
Implicit Inference
E.g. of results on Data


Different NN-based summaries from the maps + Cl
Robust FLI with Implicit Inference
We've seen that implicit inference is usually combined with summary statistics. Why?
- NLE needs to learn
- NLE needs to learn
- NPE needs to learn
Classical NFs are not good in high dimensions
How to build summaries that can extract all the information?

The NF needs to learn the distribution for each AND the complex relation between and
Robust FLI with Implicit Inference
Indeed, we saw in previous slides that the summaries have an impact on the constraining power
Sufficient Statistic:
Mutual information

Robust FLI with Implicit Inference
Sufficient Statistic:
Mutual information

Only a matter of the loss function we use!
Robust FLI with Implicit Inference
Regression Losses
Information-based Losses
→ Build sufficient statistics by definition.
Mean Squared Error (MSE) loss:
→ Approximate the mean of the posterior.
Sufficient Statistic:
Robust FLI with Implicit Inference



Robust FLI with Implicit Inference
Joint Inference with Implicit Inference
Joint Inference with Implicit Inference
Implicit Inference can be extended to the joint posterior of initial conditions. For this, we need generative models that scale to high dimension.
Flow Matching / Stochastic Interpolants / Diffusion Models (Albergo et al. (2025))
We need to learn a continuous transformation solution of the ODE
velocity field
More flexible!
Joint Inference with Implicit Inference
We need to learn a continuous transformation solution of the ODE
Joint Inference with Implicit Inference
Credit: Gagneux et al. 2025
We need to learn a continuous transformation solution of the ODE
Joint Inference with Implicit Inference
We need to learn a continuous transformation solution of the ODE
Too difficult to train under the NLL, as we need to solve the ODE for each simulation:
Joint Inference with Implicit Inference
We need to learn a continuous transformation solution of the ODE
This is the NN!
Lipman et al. (2023)
Joint Inference with Implicit Inference
We need to learn a continuous transformation solution of the ODE
This is the NN!
Lipman et al. (2023)
Joint Inference with Implicit Inference
We need to learn a continuous transformation solution of the ODE
with:

Tong et al. 2023
Joint Inference with Implicit Inference


Flow Matching, Diffusion Models and Stochastic Interpolants are the same generative framework (Albergo et al. (2025)):
For instance:
Joint Inference with Implicit Inference
Normalizing Flow
Stochastic Interpolant


For instance:
Joint Inference with Implicit Inference
Directly learn the joint with Flow Matching
We need an architecture that can work with multimodal data.




Inference Validation
Inference Validation
Underconfident
Biased
Overconfident
ICML Spotlight ✨
Inference Validation
ICML Spotlight ✨
.
Theorem:
Two distributions are equal if their probability measures are the same over all measurable sets.
6 pink samples
5 blue samples
✅
6 pink samples
5 blue samples
✅
Inference Validation
ICML Spotlight ✨
.
Theorem:
Two distributions are equal if their probability measures are the same over all measurable sets.
6 pink samples
5 blue samples
✅
6 pink samples
5 blue samples
✅
6 pink samples
✅
6 pink samples
7 blue samples
✅
Inference Validation
ICML Spotlight ✨
.
Theorem:
Two distributions are equal if their probability measures are the same over all measurable sets.
6 pink samples
5 blue samples
✅
6 pink samples
5 blue samples
✅
6 pink samples
✅
6 pink samples
7 blue samples
✅
✅
8 pink samples
7 blue samples
✅
Inference Validation
ICML Spotlight ✨
Theorem:
Two distributions are equal if their probability measures are the same over all measurable sets.
6 pink samples
5 blue samples
✅
6 pink samples
5 blue samples
✅
6 pink samples
✅
6 pink samples
7 blue samples
✅
✅
8 pink samples
7 blue samples
✅
.
Inference Validation
ICML Spotlight ✨
.
Bayesian approach: what is the probability that the true sample lies inside or outside the region given than n samples from the proposed one are inside?
Theorem:
Lemma:
Inference Validation
ICML Spotlight ✨
- Sample-based
- Can work with few samples
- Works in high dimension
- Does not rely on training a model
- Detect miscalibration even when other scores fail
- It is a scalar value


Which posterior is the best?
MIRA (Mass In Random Area)
A Score for Conditional Distribution Accuracy and Model Comparison
Dealing with Expensive Simulations
Dealing with Expensive Simulations
Simulation-Based Inference methods rely exclusively on simulations
Simulations have to be the most realistic
Very costly


Explicit Inference

Implicit Inference

Compression

Preliminary results!
Dealing with Expensive Simulations
Implicit Inference methods with fewer simulations
Sequential methods
We can sample simulations only where we need!
Prior
Evaluate
Posterior
Observation
Posterior Estimator
Simulator
Neural Posterior Estimation
Sequential
Proposal
Need to re weight for SNPE! See Greenberg et al. (2019)
Dealing with Expensive Simulations
Implicit Inference methods with fewer simulations
Gradient-based methods
?
❌
✅





For instance for NPE:
Brehmer et al. (2018), Zeghal et al. (2022)

Dealing with Expensive Simulations
Implicit Inference methods with fewer simulations
Multi Fidelity methods
For instance:






Dealing with Expensive Simulations
Emulating Simulations
For instance, we can learn the mapping between a cheap and a realistic simulation
Easier than learning the transformation from ICs
Easier than learning the transformation from ICs
Easier than learning the transformation from ICs
→ e.g. log-normal, LPT, PM
| O(ms) runtime | ✅ |
| differentiable | ✅ |
| realistic | ❌ |
Fast simulations

→ e.g. full nbody, hydro

Costly simulations
| O(ms) runtime | ❌ |
| differentiable | ❌ |
| realistic | ✅ |
Trained under MSE
Dealing with Expensive Simulations
Emulating Simulations
minimized by
which is fine if, for instance , is a dirac



when is not a dirac we can use generative models to get probable samples
Dealing with Expensive Simulations
Emulating Simulations
For instance, in Zeghal et al. (2025) we aim to approximate
from unpaired simulations



Dataset 1
Optimal Transport Plan



Dataset 2
is not a dirac so we use Flow Matching
to train with fewer simulations, we employ OT to find the minimal transformation
Dealing with Expensive Simulations
Emulating Simulations


Dealing with Expensive Simulations
Emulating Simulations


LogNormal
Emulated

Challenge simulation
VS
🥳



ML is already reshaping the way we analyze data
Thank you for your attention!
Cosmology From Home 2026
By Justine Zgh
Cosmology From Home 2026
- 22

