Large Science Models:

Foundation Models for
Generalizable Insights Into Complex Systems

with Psycho-social Application 

PI: Ishanu Chattopadhyay, PhD

Assistant Professor of Biomedical Informatics & Computer Science

University of Kentucky

DARPA-EA-25-02-05-MAGICS-PA-025

HR0011-26-3-E016

July 2026

Proposed Concept

  • Develop Foundation models of complex systems with
    • hundreds to thousands of evolving variables with apriori unknown cross-talk
    • no governing equations are know a priori
    • reflexivity: system changes if observed
  • Learn intrinsic system geometry from data
  • Derive  equations of motion with variational principles (stationary action on Lagrangian). 
  • Inference under data sparsity
  • Detect data (in)sufficiency, adapt to model drift
  • Support forward simulation and perturbation analysis
  • Digital twins of individuals & groups wrt to opinion dynamics

MAGICS Alignment

Data inference boundaries & limitations

Alignment validation 

Complex phenomena

Adaptation to model obsolence

Psychosocial domain limitations

Precise validation protocols to assess process drift triggering re-calibration/training

Built-in flexibility for changing contexts and non-ergodicity

Scalable to thousands to millions of variables, intrinsic reflexivity

Validate social theories with granular simulations from  digital twins of opinion dynamics and social behavior

Component LSM predictors enforce statistical significance of splits in recursive partitioning, ensuring precise uncertainty quantification

*Hothorn, Torsten, Kurt Hornik, and Achim Zeileis. "Unbiased recursive partitioning: A conditional inference framework." Journal of Computational and Graphical statistics 15, no. 3 (2006): 651-674.

emergent macro-structure

Component predictor (Conditional Inference Tree*)

Example: Influenza A HA protein

Recursive

LSM

forest

LSM Forest

Recursive LSM forest: hyperlinked nodes capturing emergent macro-structures

GSS 2018 dataset

  • Set of conditional inference trees (CIT)
    • Strict statistical guarantees: quantifies inference uncertainty
  • Each tree models exactly one variable as a function of potentially all other variables
  • Non-leaf nodes are "hyperlinked" to other trees

Large Science Models

Computationally tractable LSM tree structure given, as proposed, hundreds to thousands of observable variables.

GSS 2018 dataset

  • Each predictor is inferred independently
  • Can scale up to thousands of variables in Python implementation
  • Further scale-up \(10^6 - 10^8\) needs C/C++ implementation

Full Example  of Hyperlinked Trees

DTAG: Global Digital Twin of Opinions

Recall Bail etal.

“Exposure to opposing views on social media can increase political polarization” by Christopher A. Bail et al., published in PNAS in September 2018 (Vol. 115, No. 37, pp. 9216–9221; DOI: 10.1073/pnas.1804840115)

We find more general possibilities: We can make world-views go more extreme or less extreme based on the line of questions and the persona

Perturbing with opposing views made conservatives more conservative (statistically significant), liberals more liberal (not statistically significant)

LSM

  • Fixed Question Sets Exist That Move Different Persona Towards Polarization/Depolarization
  • We can optimize question sequences to move the same persona in a chosen direction
  • Note: It is relatively easy to polarize than to depolarize

Prospective validation in Human Cohorts

  • Beyond MAGICS Scope
  • Prolific Experiments using non-DARPA External Funding with UKy IRB approval

x

demographic filter
persona filter
P2
P1
fixed question sets
attention questions

Questions