Justine Zeghal
Mila, Université de Montréal
Cosmology From Home 2026, Online
Goal: Get the value of the cosmological parameters
with the uncertainty!
Bayes theorem:
To study the nature of Dark Energy
The power spectrum is near gaussian so we have an approximation of the likelihood
The power spectrum is near gaussian so we have an approximation of the likelihood
DES Y3 WL Results (with SBI).
The power spectrum is not a sufficient statistics for non gaussian field
Stage III
Stage IV
Portion of the Virgo cluster, zoom on RSCG 55
Portion of the Virgo cluster, zoom on RSCG 55
Access to new non gaussian small scales. We don't want to lose this new information!
Simulator
Initial conditions
Large Scale Structure
Prediction
Inference
Explicit inference
Needs an explicit simulator to sample the joint posterior through MCMC:
Implicit inference
We use simulations to
learn
Instead of relying on an analytical model to describe the phenomenon, we can simulate it.
Two ways of performing inference from simulations:
Because we work at the map level, considering all cosmological information, we call this inference:
Field-level inference / Pixel-level inference / Full-field inference
Most precise inference!
Explicit inference
Needs an explicit simulator to sample the joint posterior through MCMC:
Implicit inference
We use simulations to
learn
Instead of relying on an analytical model to describe the phenomenon, we can simulate it.
Two ways of performing inference from simulations:
From the dataset we can learn:
Since the goal of implicit inference is to approximate a distribution, it has greatly benefited from the advent of generative models.
Definition:
A generative model is a machine learning model designed to create new data that is similar to its training data .
Different kinds of generative models:
Change of Variable Formula:
Should be easy to compute
Where is the NN?
Where is the NN?
For instance, for affine transformation (RealNVP, Dinh et al. 2017):
NN
It is such a specific function.. Why?
Easy to inverse
Easy to compute the determinant of the Jacobian
Credit: François Lanusse
We need a tool to compare distributions:
the Kullback-Leiber Divergence
How to train the NNs?
We want to minimize the Kullback-Leiber Divergence wrt
We want to minimize the Kullback-Leiber Divergence wrt
We want to minimize the Kullback-Leiber Divergence wrt
Simulations only!
Change of variable formula
We want to minimize the Kullback-Leiber Divergence wrt
Likelihood approximation (e.g. Papamakarios et al., 2019)
Likelihood approximation (e.g. Papamakarios et al., 2019)
Posterior approximation (e.g. Papamakarios et al., 2016)
https://simulation-based-inference.org/by Kyle Cranmer and Jason Lo
E.g. of results on Data
Reminder: Implicit Inference enables approximating the posterior when the likelihood is unknown but implicitly encoded in simulations.
Summary statistics are not always Gaussian distributed.
wavelet phase harmonic
scattering transform coefficients
Computing the covariance matrix of a combination of summary statistics is challenging.
Adding more summary statistics tightens the constraints
E.g. of results on Data
NN-based summaries from the maps + Cl
E.g. of results on Data
Different NN-based summaries from the maps + Cl
We've seen that implicit inference is usually combined with summary statistics. Why?
Classical NFs are not good in high dimensions
How to build summaries that can extract all the information?
The NF needs to learn the distribution for each AND the complex relation between and
Indeed, we saw in previous slides that the summaries have an impact on the constraining power
Sufficient Statistic:
Mutual information
Sufficient Statistic:
Mutual information
Only a matter of the loss function we use!
Regression Losses
Information-based Losses
→ Build sufficient statistics by definition.
Mean Squared Error (MSE) loss:
→ Approximate the mean of the posterior.
Sufficient Statistic:
Implicit Inference can be extended to the joint posterior of initial conditions. For this, we need generative models that scale to high dimension.
Flow Matching / Stochastic Interpolants / Diffusion Models (Albergo et al. (2025))
We need to learn a continuous transformation solution of the ODE
velocity field
More flexible!
We need to learn a continuous transformation solution of the ODE
Credit: Gagneux et al. 2025
We need to learn a continuous transformation solution of the ODE
We need to learn a continuous transformation solution of the ODE
Too difficult to train under the NLL, as we need to solve the ODE for each simulation:
We need to learn a continuous transformation solution of the ODE
This is the NN!
Lipman et al. (2023)
We need to learn a continuous transformation solution of the ODE
This is the NN!
Lipman et al. (2023)
We need to learn a continuous transformation solution of the ODE
with:
Tong et al. 2023
Flow Matching, Diffusion Models and Stochastic Interpolants are the same generative framework (Albergo et al. (2025)):
For instance:
Normalizing Flow
Stochastic Interpolant
For instance:
Directly learn the joint with Flow Matching
We need an architecture that can work with multimodal data.
Underconfident
Biased
Overconfident
ICML Spotlight ✨
ICML Spotlight ✨
.
Theorem:
Two distributions are equal if their probability measures are the same over all measurable sets.
6 pink samples
5 blue samples
✅
6 pink samples
5 blue samples
✅
ICML Spotlight ✨
.
Theorem:
Two distributions are equal if their probability measures are the same over all measurable sets.
6 pink samples
5 blue samples
✅
6 pink samples
5 blue samples
✅
6 pink samples
✅
6 pink samples
7 blue samples
✅
ICML Spotlight ✨
.
Theorem:
Two distributions are equal if their probability measures are the same over all measurable sets.
6 pink samples
5 blue samples
✅
6 pink samples
5 blue samples
✅
6 pink samples
✅
6 pink samples
7 blue samples
✅
✅
8 pink samples
7 blue samples
✅
ICML Spotlight ✨
Theorem:
Two distributions are equal if their probability measures are the same over all measurable sets.
6 pink samples
5 blue samples
✅
6 pink samples
5 blue samples
✅
6 pink samples
✅
6 pink samples
7 blue samples
✅
✅
8 pink samples
7 blue samples
✅
.
ICML Spotlight ✨
.
Bayesian approach: what is the probability that the true sample lies inside or outside the region given than n samples from the proposed one are inside?
Theorem:
Lemma:
ICML Spotlight ✨
Which posterior is the best?
MIRA (Mass In Random Area)
A Score for Conditional Distribution Accuracy and Model Comparison
Simulation-Based Inference methods rely exclusively on simulations
Simulations have to be the most realistic
Very costly
Explicit Inference
Implicit Inference
Compression
Preliminary results!
Implicit Inference methods with fewer simulations
Sequential methods
We can sample simulations only where we need!
Prior
Evaluate
Posterior
Observation
Posterior Estimator
Simulator
Neural Posterior Estimation
Sequential
Proposal
Need to re weight for SNPE! See Greenberg et al. (2019)
Implicit Inference methods with fewer simulations
Gradient-based methods
?
❌
✅
For instance for NPE:
Brehmer et al. (2018), Zeghal et al. (2022)
Implicit Inference methods with fewer simulations
Multi Fidelity methods
For instance:
Emulating Simulations
For instance, we can learn the mapping between a cheap and a realistic simulation
Easier than learning the transformation from ICs
Easier than learning the transformation from ICs
Easier than learning the transformation from ICs
→ e.g. log-normal, LPT, PM
| O(ms) runtime | ✅ |
| differentiable | ✅ |
| realistic | ❌ |
Fast simulations
→ e.g. full nbody, hydro
Costly simulations
| O(ms) runtime | ❌ |
| differentiable | ❌ |
| realistic | ✅ |
Trained under MSE
Emulating Simulations
minimized by
which is fine if, for instance , is a dirac
when is not a dirac we can use generative models to get probable samples
Emulating Simulations
For instance, in Zeghal et al. (2025) we aim to approximate
from unpaired simulations
Dataset 1
Optimal Transport Plan
Dataset 2
is not a dirac so we use Flow Matching
to train with fewer simulations, we employ OT to find the minimal transformation
Emulating Simulations
Emulating Simulations
LogNormal
Emulated
Challenge simulation
VS
ML is already reshaping the way we analyze data