Justine Zeghal
Mila, Université de Montréal
PIML workshop 2026, Heraklion, Crete, GreeceGoal: Get the value of the cosmological parameters
with the uncertainty!
Bayes theorem:
The power spectrum is near gaussian so we have an approximation of the likelihood
The power spectrum is near gaussian so we have an approximation of the likelihood
DES Y3 WL Results (with SBI).
The power spectrum is not a sufficient statistics for non gaussian field
Stage III
Stage IV
Portion of the Virgo cluster, zoom on RSCG 55
Portion of the Virgo cluster, zoom on RSCG 55
Access to new non gaussian small scales. We don't want to lose this new information!
Simulator
Initial conditions
Large Scale Structure
Prediction
Inference
Explicit inference
Needs an explicit simulator to sample the joint posterior through MCMC:
Implicit inference
We use simulations to
learn
Instead of relying on an analytical model to describe the phenomenon, we can simulate it.
Two ways of performing inference from simulations:
Explicit inference
Needs an explicit simulator to sample the joint posterior through MCMC:
Implicit inference
We use simulations to
learn
Instead of relying on an analytical model to describe the phenomenon, we can simulate it.
Two ways of performing inference from simulations:
Because we work at the map level, considering all cosmological information, we call this inference:
Field-level inference / Pixel-level inference / Full-field inference
Most precise inference!
From the dataset we can learn:
We use generative models
Most of the time: Normalizing Flows
Change of Variable Formula:
Should be easy to compute
We need to learn the mapping to approximate the complex distribution
Credit: François Lanusse
We need a tool to compare distributions:
the Kullback-Leiber Divergence
We want to minimize the Kullback-Leiber Divergence wrt
We want to minimize the Kullback-Leiber Divergence wrt
We want to minimize the Kullback-Leiber Divergence wrt
Simulations only!
Change of variable formula
We want to minimize the Kullback-Leiber Divergence wrt
Likelihood approximation (e.g. Papamakarios et al., 2019)
Likelihood approximation (e.g. Papamakarios et al., 2019)
Posterior approximation (e.g. Papamakarios et al., 2016)
Both methods require a lot of costly simulations!
Brehmer et al. (2018), Zeghal et al. (2022)
For instance for NPE:
Neural-based SBI
?
❌
✅
We also want to use the simulator's gradients information:
and we want to learn the marginal posterior but
Minimized by
Neural-based SBI
Lack expressivity!
Brehmer et al. (2018), Zeghal et al. (2022)
Neural-based SBI
Without gradients
With gradients
Some results
Without gradients
With gradients
Brehmer et al. (2018), Zeghal et al. (2022)
Neural-based SBI
Application to Weak Lensing Cosmology
Zeghal et al. (2024)
Neural-based SBI
Application to Weak Lensing Cosmology
Zeghal et al. (2024)
No improvements when using the gradients.
We are using the joint gradients:
Need to check your gradients before. And maybe use variance reduction techniques.
For instance for NPE:
Neural-based SBI
?
❌
✅
Lack expressivity!
Brehmer et al. (2018), Zeghal et al. (2022)
→ e.g. log-normal, LPT, PM
| O(ms) runtime | ✅ |
| differentiable | ✅ |
| realistic | ❌ |
Fast simulations
→ e.g. full nbody, hydro
Costly simulations
| O(ms) runtime | ❌ |
| differentiable | ❌ |
| realistic | ✅ |
We can learn the mapping between a cheap and a realistic simulation
Easier to learn a small correction & requires fewer simulations
One way to do:
minimized by
which is fine if, for instance , is a dirac
We can learn the mapping between a cheap and a realistic simulation
Easier to learn a small correction & requires fewer simulations
One way to do:
minimized by
which is fine if, for instance , is a dirac
We can learn the mapping between a cheap and a realistic simulation
Easier to learn a small correction & requires fewer simulations
One way to do:
minimized by
which is fine if, for instance , is a dirac
when is not a dirac we should use generative models to get probable samples
For instance, in Zeghal et al. (2025) we aim to approximate
from unpaired simulations
Dataset 2
Dataset 1
unlearning dataset 1 would requires more simulations than starting from gaussian noise
Conditional Optimal Transport Flow Macthing (Kerrigan et al. 2024)
Conditional Optimal Transport Flow Macthing (Kerrigan et al. 2024)
We need to learn a continuous transformation solution of the ODE
velocity field
More flexible!
Conditional Optimal Transport Flow Macthing (Kerrigan et al. 2024)
Conditional Optimal Transport Flow Macthing (Kerrigan et al. 2024)
Credit: Gagneux et al. 2025
Conditional Optimal Transport Flow Macthing (Kerrigan et al. 2024)
Lipman et al. (2023)
Conditional Optimal Transport Flow Macthing (Kerrigan et al. 2024)
Lipman et al. (2023)
Conditional Optimal Transport Flow Macthing (Kerrigan et al. 2024)
Lipman et al. (2023)
Conditional Optimal Transport Flow Macthing (Kerrigan et al. 2024)
with:
Tong et al. 2023
Conditional Optimal Transport Flow Macthing (Kerrigan et al. 2024)
with:
Tong et al. 2023
Conditional Optimal Transport Flow Macthing (Kerrigan et al. 2024)
Unlearning dataset 1 would requires more simulations than starting from gaussian noise
OT to find the minimal effort mapping
Conditional Optimal Transport Flow Macthing (Kerrigan et al. 2024)
Unlearning dataset 1 would requires more simulations than starting from gaussian noise
OT to find the minimal effort mapping
Conditional Optimal Transport Flow Macthing (Kerrigan et al. 2024)
Unlearning dataset 1 would requires more simulations than starting from gaussian noise
OT to find the minimal effort mapping
Conditional Optimal Transport Flow Macthing (Kerrigan et al. 2024)
Unlearning dataset 1 would requires more simulations than starting from gaussian noise
OT to find the minimal effort mapping
Conditional Optimal Transport Flow Macthing (Kerrigan et al. 2024)
❌
Conditional Optimal Transport Flow Macthing (Kerrigan et al. 2024)
OT will find the "closest" map!
Conditional Optimal Transport Flow Macthing (Kerrigan et al. 2024)
✅
Zeghal et al. (2025)
LogNormal
Emulated
Challenge simulation
VS
Zeghal et al. (2025)
Implicit inference with high dimensional observations is hard
NF are not good in high dimensions
Goal: to build sufficient statistics as a first step, and then run the NPE or NLE method.
The NF needs to learn the distribution for each AND the complex relation between and
Sufficient Statistic:
Mutual information
Sufficient Statistic:
Mutual information
Only a matter of the loss function we use!
Lanzieri & Zeghal et al. (2025)
Regression Losses
Information-based Losses
→ Build sufficient statistics by definition.
Mean Squared Error (MSE) loss:
→ Approximate the mean of the posterior.
Sufficient Statistic:
Lanzieri & Zeghal et al. (2025)
Biased
Overconfident
Underconfident
Sharief & Zeghal et al. (2026)
.
Theorem:
Two distributions are equal if their probability measures are the same over all measurable sets.
6 pink samples
5 blue samples
✅
6 pink samples
5 blue samples
✅
Sharief & Zeghal et al. (2026)
.
Theorem:
Two distributions are equal if their probability measures are the same over all measurable sets.
6 pink samples
5 blue samples
✅
6 pink samples
5 blue samples
✅
6 pink samples
✅
6 pink samples
7 blue samples
✅
Sharief & Zeghal et al. (2026)
.
Theorem:
Two distributions are equal if their probability measures are the same over all measurable sets.
6 pink samples
5 blue samples
✅
6 pink samples
5 blue samples
✅
6 pink samples
✅
6 pink samples
7 blue samples
✅
✅
8 pink samples
7 blue samples
✅
Sharief & Zeghal et al. (2026)
Theorem:
Two distributions are equal if their probability measures are the same over all measurable sets.
6 pink samples
5 blue samples
✅
6 pink samples
5 blue samples
✅
6 pink samples
✅
6 pink samples
7 blue samples
✅
✅
8 pink samples
7 blue samples
✅
.
Sharief & Zeghal et al. (2026)
.
Bayesian approach: what is the probability that the true sample lies inside or outside the region given than n samples from the proposed one are inside?
Theorem:
Lemma:
Sharief & Zeghal et al. (2026)
Is it working?
Check the paper for other examples!
Sharief & Zeghal et al. (2026)
Benefits
Which posterior is the best?
Sharief & Zeghal et al. (2026)
Omori & Zeghal et al. (2026)
Explicit inference becomes more challenging as the realism (i.e. the complexity) of the simulator increase
In WL: we infer 3D ICs from a 2D field
Before us in WL: only LPT simulators have been used but never validated against implicit
LPT sampling works fine
We pushed the comparison to PM Nbody simulators
PM sampling works fine but it is very challenging
We need better sampling methods
We need explicit inference methods that can perform with fewer simulations
Zeghal et al. (2024)