From Supercomputers to Health Departments

Roux Institute + Network Science Institute
Northeastern University

George G. Vega Yon, Ph.D.

The University of Utah

2026-09-14

The research map

---
config:
  htmlLabels: false
---
%% htmlLabels:false is load-bearing, not cosmetic. With HTML labels mermaid
%% sizes each box from getBoundingClientRect(), which reveal's slide-scaling
%% transform multiplies -- every box comes out ~40% too narrow and every label
%% is clipped. SVG labels are measured with getBBox(), which the transform
%% does not touch. (Labels still wrap at 200px either way.) The price is that
%% mermaid then never centres them, so style.scss re-anchors the text; see
%% "Research-map mindmap" there. No themeVariables.fontSize either: the SVG is
%% scaled to the slide, which grows boxes and text together anyway.
%%
%% Mindmaps have no `classDef` -- the palette for `branch` / `today` also
%% lives in style.scss. Two syntax traps: `:::class` must sit on its OWN line
%% under the node it decorates (`node:::x` is a parse error), and a bare node
%% label cannot contain parentheses, hence the quoted `["..."]` form on the
%% leaves. The layout is force-directed: identical on every load at a given
%% window size, but it can settle into a different arrangement on a screen of
%% a different shape. Check it on the room's projector, not just your laptop.
mindmap
  root(("Statistical<br>computing and<br>scientific<br>software"))
    NS(Network science)
    :::branch
      E1["Applied ERGMs"]
      :::today
      E2["Other graphical<br>models"]
      E3["ERGM methods"]
      E4["Social Influence"]
    PH(Public health)
    :::branch
      P1["Measles"]
      :::today
      P2["Agent-Based Models"]
      P3["Mechanistic Machine<br>Learning"]
      P4["Economic Impact"]
    OTH(Other domains)
    :::branch
      O1["Genomics"]
      O2["Quantum<br>Computing"]

Note: I owe much of the diversity in my work to my collaborators: it is not a one-person show!

Dive 1 – Agent-Based Models

Can you build a model of measles spreading in schools in 3 months?

With Randon Gruninger – Joshua Kelley – Damon Toth – Leisha Nolen – Delaney Thornton – Andrew Pavia – Derek Meyer – Andrew Pulsipher

How it came together

Dec 2024

Playing with epiworldRShiny

Jan 2025

Let’s see what we can build for measles

Kickoff (Jan 30th)

Goal: Build a better Shiny app

We agreed in January to undertake a three-month project.

Suddenly…

On February 5th, 2025, public health officials sound the alarm: a measles outbreak has begun in Gaines County, Texas.

Source: Mole (2025)

Source: Mole (2025)

Can you build a model of measles spreading in schools in 3 months weeks?

How it came together (cont.)

We agreed in January to undertake a three-month project… With the legislative session approaching (March 7th), it became a three-week project.

Dec 2024

Playing with epiworldRShiny

Jan 2025

Let’s see what we can build for measles

Kickoff (Jan 30th)

Goal: Build a better Shiny app measles model for schools

Model design

Literature review

Disease progression

Economic impact

First version

Reviewed by the team

Adapted following feedback

Run analyses (March 7th)

Selected schools.

Discussed results with lawmakers.

Shiny App

We took a deep dive to reflect on how it unfolded in Vega Yon, Thornton, et al. (2026)

Results: Shiny App

You can access the app at https://ggvy.cl/shiny/measles

Measles School Letters

or “We sent the report; DHHS used it to inform policy… now what?”

Results: 1,200 letters

A personalized letter for every school in Utah

  • Inputs: enrollment and up-to-date vaccination coverage
  • Output: outbreaks with and without the quarantine process

Original template text released with the letters

Original template text released with the letters

What school nurses receiving letters look like according to Microsoft Copilot

What school nurses receiving letters look like according to Microsoft Copilot

The workflow is free and public on GitHub. Minnesota (MADMC, another InsightNet site) is running it; other sites are evaluating it.

The measles model

  • Robust C++ core. Implemented as an extension of epiworld (Vega Yon 2022). The measles tool is available as an R package on GitHub and CRAN (Vega Yon 2026d).

  • Discrete time. Each day is one time step.

  • Homogeneous mixing. Contact rate-based interaction.

  • Full disease progression. Times are geometrically distributed.

  • Quarantine process. Probability of detection and agents’ willingness to follow it.

Two levers: vaccination coverage and the quarantine process.

The engine: epiworld

  • A fremework for fast, scalable agent-based models.
  • Written in C++: 19K lines of code + 9K lines of tests.
  • Multi/evolving-disease models.
  • 15K downloads on CRAN (epiworldR).
  • Parallel computing via OpenMP out of the box.

It is also available in R (Meyer and Vega Yon 2023), Python (Vega Yon and Banks 2026), Shiny (Meyer and Vega Yon 2025), and WebAssembly (Banks and Vega Yon 2024).

More about its <speed>

Measles – Utah-Arizona border outbreak

With Abbey Marye – Damon Toth – Jake Wagoner – Eric Lofgren – Matthew Samore – Andrew Pavia – Leisha Nolen – Kenneth Komatsu – Kelly Oakeson – Lindsay Keegan

From one school to a community

  • Using the school model as a baseline, we implemented a community-scale version1.
  • Multiple groups; heterogeneous mixing between them; group-specific vaccination; contact-tracing-based quarantine.

  • The R package measles (Vega Yon 2026d) implements these models on top of epiworldR. It is available on GitHub and CRAN and is under review at JOSS.

Utah-Arizona border outbreak

  • Largest outbreak in Utah.
  • 585 confirmed cases associated with that outbreak.
  • A population of about 6,000 people.
  • One of the areas with the lowest vaccination coverage in the state, at about 50% (see <details>).
  • A complex relationship with public health.

The number of confirmed cases does not align with the data (Marye et al. 2026)

Results: Utah-Arizona outbreak

- Utah DHHS estimated outbreak size using omics.

- We used epiworldR+measles, incorporating (a) school vaccination coverage, (b) census age distributions, and (c) POLYMOD/Epistorm contact patterns.

- Performed a sensitivity analysis accounting for parameter uncertainty.

Source Infections1
Confirmed cases 585
ABM estimates ~2,226
Genomic estimates ~2,051

On the bright side… Source: Utah DHHS (link)

Both methods put the number closer to 2,000 than to 600, implying that roughly three in four infections went uncounted.

Where else the model goes

2026 FIFA World Cup

Simulations include up to 1,000,000 agents (on free computing resources) across all 11 US host cities, with each city calibrated to its own age distribution and vaccination coverage (Lessler et al. 2026; Vega Yon 2026a).

Graded quarantine

Should quarantine length depend on exposure – same school vs. same classroom vs. a known direct contact? We built a version of the model to answer exactly that question.

Health district reports

Which schools put the surrounding community at risk? A compartmental version with no quarantine answers this question (Toth et al. 2026). Reports have already been delivered to every Utah health district.

Economic evaluation

EpiCC is a serverless cost calculator for measles and TB outbreaks and interventions (Banks et al. 2026).

Dive 2 – Network Models
or: How I Learned to Stop Worrying and Love the Normalizing Constant

Multilevel Bipartite Exponential-family Random Graph Models (ERGMs) in patient-provider networks

With Chong Zhang – Karim Khader – Nai-Chung Nelson Chang – Lindsay Visnovsky – Alun Thomas – Candace Haroldsen – Kristina Stratford – Matthew Samore

Healthcare-Associated Infections

  • Worldwide, “around 1 in 10 patients is affected by Healthcare-Associated Infections” (HAIs) (Pittet 2026).

  • “[I]n the [US], the CDC reported nearly 687,000 HAIs in acute care facilities, resulting in 72,000 associated deaths” (Odoom and Donkor 2025).

  • Contact networks are a key mechanism for the spread of HAIs (Curtis et al. 2013).

Figure 1 from Curtis et al. (2013): A marked-up architectural CAD drawing fragment of the Medical Intensive Care Unit at the University of Iowa Hospitals and Clinics

Figure 1 from Curtis et al. (2013): A marked-up architectural CAD drawing fragment of the Medical Intensive Care Unit at the University of Iowa Hospitals and Clinics

Our Data: Contact networks

ERGMs in resident–healthcare provider networks (Vega Yon, Zhang, et al. 2026) · Social Networks (in press)

  • Contact networks inside long-term care facilities: healthcare personnel on one side and residents on the other (bipartite).

  • 93 bipartite networks from 47 facility units, ranging in size from 10 to 100 nodes.

  • Healthcare personnel: certified nurse assistants (CNAs), registered nurses (RNs), physical therapists (PTs), occupational therapists (OTs), respiratory therapists (RTs), and other healthcare providers.

  • Residents: dialysis, wound care, bedridden, ventilator, MDRO status.

Figure 1 from Vega Yon, Zhang, et al. (2026) drawn using netplot (Vega Yon 2026e)

The key questions

We wanted to address the following questions regarding the contact networks:

HCP contact network #66: Purple nodes correspond to healthcare personnel, whereas green nodes correspond to residents.

HCP contact network #66: Purple nodes correspond to healthcare personnel, whereas green nodes correspond to residents.
  • Do different types of healthcare personnel differ in their contact patterns?
  • Do residents with different care needs differ in their contact patterns?
  • Can we say anything about cohorting?
  • Do we observe any differences between facilities?

Exponential-family Random Graph Models (ERGMs) provide a natural framework for answering these questions.

ERGMs

ERGMs provide a parametric representation of the distribution of \mathbf{y}\in \mathcal{Y}1.

\begin{aligned} {\mathbb{P}_{}\left\{\mathbf{Y}= \mathbf{y}; \boldsymbol{\theta}\right\} } & = \frac{\text{exp}\left\{\theta^{\text{T}}\textcolor{green}{\mathbf{g}(\mathbf{y})}\right\}}{\kappa\left(\theta, \mathcal{Y}\right)},\quad\mathbf{y} \in\mathcal{Y} \\ \textcolor{red}{\kappa\left(\theta,\mathcal{Y}\right)} & = \sum_{\mathbf{y}\in\mathcal{Y}}\text{exp}\left\{\theta^{\text{T}}\textcolor{green}{\mathbf{g}(\mathbf{y})}\right\} \end{aligned} \tag{1}

  • \textcolor{green}{\mathbf{g}(\mathbf{y})} is the key: it captures the essential features of the network structure.
  • \textcolor{red}{\kappa} (the normalizing constant)2 is intractable for most networks, so estimation requires simulation.
  • Simulation is efficient via the conditional distribution of a dyad given the rest of the network (a logit model!):

{\mathbb{P}_{\boldsymbol{\theta}}\left\{\mathbf{Y}_{ij}{=}1\vphantom{\mathbf{Y}_{-ij}}\;\right|\left.\vphantom{\mathbf{Y}_{ij}{=}1}\mathbf{Y}_{-ij}\right\}} = \frac{1}{1 + \text{exp}\left\{-{\boldsymbol{\theta}}^\mathbf{t}\textcolor{purple}{\delta_{ij}}\right\}}

where \textcolor{purple}{\delta_{ij} = \mathbf{g}(\mathbf{Y}_{ij}^{+}) - \mathbf{g}(\mathbf{Y}_{ij}^{-})}.

ERGMs: Simulating edges + homophily

This model features only edges and nodematch (homophily):

{\mathbb{P}_{}\left\{\mathbf{Y}= \mathbf{y}; \boldsymbol{\theta}\right\} } = \frac{\text{exp}\left\{\theta_{\text{edges}}\textcolor{green}{\mathbf{g}_{\text{edges}}(\mathbf{y})} + \theta_{\text{nodematch}}\textcolor{green}{\mathbf{g}_{\text{nodematch}}(\mathbf{y})}\right\}}{\kappa\left(\theta, \mathcal{Y}\right)}

Yet \textcolor{red}{\kappa} never shows up in the simulation: it cancels in the conditional. That is how we sample from a model we cannot evaluate – and how I learned to stop worrying about it.

Challenges in our data

Projectivity

An estimate for a network of size n does not necessarily apply to a network of size n' \gg n.

SOLUTION
Use size interactions for terms that are sensitive to network size (as in Krivitsky et al., 2011).

Repeated nets

The networks are not IID: some were observed more than once.

SOLUTION
Use a block bootstrap to account for the repeated measurements.

Multicollinearity

Some statistics are highly correlated, which can lead to numerical instability and convergence issues.

SOLUTION
Use an affine reparameterization of some terms.

This is one of the first applications of multilevel bipartite ERGMs.

More about the <terms> used in the model.

The compute

Two of our own contributions – not more hardware – turned this from an overnight job into a coffee break:

Step (MC-MLE) Stock Ours Speedup
Single fit 1 hr
(1 core)
30 min
(affine reparam.)
2×
100-rep bootstrap 12 hrs
(40 cores)
10 min
(Robbins-Monro)
72×
Final model
(200 reps)
– 20 min
(40 cores)
–

The functions are available as part of the publication. We plan to submit a patch to ergm.multi.

Results: HCP-Resident Networks

Results from 200 block-bootstrap replicates of the pooled ERGM fit:

Network size interaction

Strong evidence
sparser, not busier

A unit with twice as many residents does not have twice as many contacts

Shared partners

Strong evidence
spreads out

Staff who share a resident rarely share many other residents

Role mixing

Strong evidence
different roles pair

Two staff members caring for the same resident are usually different kinds of workers

Care need

Strong evidence
sicker, more contact

Residents who are sicker – on a ventilator, bedridden, or in need of wound care

Using these results, we can simulate realistic contact networks for epidemic modeling.

See more details <link>

More on ERGMs methodological advances

R21: Conditional ERGM

How many networks do you need to detect a homophily effect? –
(Vega Yon 2023 in JASA)

Figure 1 adapted from Krivitsky et al. (2023): “Empirical power p_S for detecting a gender homophily effect with an assumed \theta_{\text{homophily}} = 1.1 as a function of number of networks S, where each network has size 8 nodes and 20 edges in the scenario of Vega Yon (2023, sec. 1).”

Figure 1 adapted from Krivitsky et al. (2023): “Empirical power p_S for detecting a gender homophily effect with an assumed \theta_{\text{homophily}} = 1.1 as a function of number of networks S, where each network has size 8 nodes and 20 edges in the scenario of Vega Yon (2023, sec. 1).”

More about this project <here>

Closing

Recap

Dive 1: Agent-Based Models

Fast and scalable agent-based models for measles in schools and communities.

  • What: Online tool for simulating school outbreaks, school-level reports distributed across the state, and a community-scale model for the Utah-Arizona border outbreak.

  • Takeaway: We can build models and tools that are accessible to non-academic audiences and can inform policy.

Dive 2: Network Models

Advances in exponential-family random graph models (ERGMs)

  • What: A new application (with methodological advances) of ERGMs to bipartite networks in healthcare, plus a glimpse into conditional ERGMs (R21).

  • Takeaway: ERGMs provide a flexible framework for characterizing real-world networks. Downstream applications can leverage this, e.g., epidemiological simulation.

One more thing…

Three completely open (and reproducible) book projects:

All three are also available in Spanish!

Where this goes

Statistical
inference

  • Statistical inference in complex systems: theory and tools.
  • ERGMs: the same, plus extensions beyond SNA.

Hybrid
models

  • AI and machine learning: methods and scientific products.
  • Quantum computing: explore opportunities (seed grant at Utah).

Tools and partnerships

  • Translational science: more collaboration with public health.
  • Innovation: explore private-sector and academic partnerships.

Collaborators

Mentees

  • Sima Najafzadehkhoei
  • Anibal Olivera
  • Olivia Banks
  • Derek Meyer (former)
  • Andrew Pulsipher (former)

DHHS

  • Randon Gruninger
  • Matthew Mietchen
  • Leisha D. Nolen
  • Abbey Marye
  • Joshua Kelley (former)

Utah

  • Bernardo Modenesi
  • Lindsay T. Keegan
  • Karim Khader
  • Andrew T. Pavia
  • Matthew Samore
  • Damon Toth
  • Jake Wagoner
  • Delaney Thornton
  • Lindsay D. Visnovsky
  • Kristina Stratford
  • Richard Nelson

CDC

  • Damon Bayer
  • Dylan Morris
  • Samuel Rosenblatt
  • Emily Pollock
  • Dina Mistry

Other institutions

  • Thomas Valente (USC)
  • Kyosuke Tanaka (Aarhus University)
  • Kayla de la Haye (UCLA)
  • Paul Marjoram (USC)
  • Sarah Piombo (Harvard)
  • Kaitlyn Johnson (LSHTM)

Thank you!

From Supercomputers to Health Departments

George G. Vega Yon, Ph.D.

Division of Epidemiology
University of Utah

https://ggvy.cl

---
config:
  htmlLabels: false
---
%% htmlLabels:false is load-bearing, not cosmetic. With HTML labels mermaid
%% sizes each box from getBoundingClientRect(), which reveal's slide-scaling
%% transform multiplies -- every box comes out ~40% too narrow and every label
%% is clipped. SVG labels are measured with getBBox(), which the transform
%% does not touch. (Labels still wrap at 200px either way.) The price is that
%% mermaid then never centres them, so style.scss re-anchors the text; see
%% "Research-map mindmap" there. No themeVariables.fontSize either: the SVG is
%% scaled to the slide, which grows boxes and text together anyway.
%%
%% Mindmaps have no `classDef` -- the palette for `branch` / `today` also
%% lives in style.scss. Two syntax traps: `:::class` must sit on its OWN line
%% under the node it decorates (`node:::x` is a parse error), and a bare node
%% label cannot contain parentheses, hence the quoted `["..."]` form on the
%% leaves. The layout is force-directed: identical on every load at a given
%% window size, but it can settle into a different arrangement on a screen of
%% a different shape. Check it on the room's projector, not just your laptop.
mindmap
  root(("Statistical<br>computing and<br>scientific<br>software"))
    NS(Network science)
    :::branch
      E1["Applied ERGMs"]
      :::today
      E2["Other graphical<br>models"]
      E3["ERGM methods"]
      E4["Social Influence"]
    PH(Public health)
    :::branch
      P1["Measles"]
      :::today
      P2["Agent-Based Models"]
      P3["Mechanistic Machine<br>Learning"]
      P4["Economic Impact"]
    OTH(Other domains)
    :::branch
      O1["Genomics"]
      O2["Quantum<br>Computing"]

Appendix

The rest of the map

Genomics / evolution

Gene function evolution over phylogenies (Vega Yon 2026b; Vega Yon, Thomas, et al. 2021) · integrative genomics in cancer (NCI P01, Co-I)

CDC collaborations

Forecasting tools (Bayer et al. 2024) · leveraging wastewater (Johnson et al. 2024) · Rt estimation (Epidemics, 2026) (Milando et al. 2026)

Clinical / health systems

MDRO-contaminated equipment networks in skilled nursing facilities (Visnovsky et al. 2026)

Social networks & behavior

Smoking (Haye et al. 2019) and e-cigarette (Piombo et al. 2025) diffusion · political discussion network formation (RSOS 2022) (Hâncean et al. 2022) · officer networks and firearm use (JQC 2022) (Ouellet et al. 2022) · imaginary network motifs (Social Networks 2024) (Tanaka and Vega Yon 2024)

Other epi

Mpox transmission heterogeneity (Love et al. 2023) · 2026 World Cup epidemic risk (Lessler et al. 2026) · quantum-enhanced epi modeling (seed grant, Co-PI)

Why not a network

One school. Homogeneous mixing. That was a decision, not a default.

  • We planned a network model first. There were no data to build it from – nothing that would pin down who contacts whom inside a school
  • And detail has a price at the other end: every extra mechanism is another thing to explain, and – we watched this happen – another thing to be questioned about
  • At the community scale, the same reasoning gives you age-structured contact matrices: coarse, defensible, and runnable by the people who need them

Fewer assumptions, stated at a level our partners could argue with.

Results: How fast is epiworld?

Comparing epiworld with other frameworks by running 100 replicates of an SEIHR model in each. More details here (<why is it so fast?>)

Comparing epiworld with other frameworks by running 100 replicates of an SEIHR model in each. More details here (<why is it so fast?>)

epydemic here is Dobson’s network ABM library – not to be confused with epydemix, Epistorm’s compartmental + ABC-calibration package!

Go back to <compute>

Why is epiworld fast?

Active epidemic queue

Only agents who are infected—or neighbors of active cases—run disease-update logic. This cuts unnecessary contact and probability calculations during sparse outbreaks.

Cache-friendly networks

Agents are stored contiguously, and contacts are compact neighbor-ID vectors. Low-degree networks use fast linear scans; high-degree nodes add a hash index.

Allocation-free hot loops

Neighbor views, temporary probability arrays, and reusable event storage avoid repeated heap allocation in daily transmission and transition calculations.

Fast random draws

A lightweight xoshiro generator, efficient bounded-integer sampling, and an optional rare-event Poisson approximation reduce RNG overhead.

Efficient setup and scaling

Sparse networks use geometric-skipping graph generators, while independent simulation replications can run concurrently with OpenMP.

Go back to <compute>

Low vaccination coverage

We estimated overall MMR coverage of ~50% in the region, with close to 3,000 susceptible individuals in the area.

Go <back>

Increased reliability

Exact statistics significantly reduce error rates when estimating ERGMs.

This includes models that, in principle, are not degenerate—that is, models that generate networks that are neither very dense nor very sparse.

The other regime: exact computation

Using the exact approach makes the estimation process orders of magnitude faster.

Compared with other methods, ergmito can be up to 100 times faster at estimating pooled models.

Exponential random graph models for little networks (Social Networks, 2021) (Vega Yon, Slaughter, et al. 2021)

ERGMs as maximum entropy

Maximize the Shannon entropy over distributions on the graph space, subject to matching the expected sufficient statistics:

H(p) = -\sum_{\mathbf{y}\in \mathcal{Y}} p(\mathbf{y})\,\text{log}\left\{p(\mathbf{y})\right\} \qquad \text{s.t.} \qquad {\mathbb{E}\left\{s_{}\left(\mathbf{Y}\right)\right\}} = \bar{\mathbf{s}}

The solution is the exponential family – the ERGM:

p(\mathbf{y}) = \frac{\text{exp}\left\{{\boldsymbol{\theta}}^\mathbf{t}s_{}\left(\mathbf{y}\right)\right\}}{\eta_{}\left(\boldsymbol{\theta}{}\right)}

s_{}\left(\mathbf{y}\right) plays the role of the energy, \eta_{}\left(\boldsymbol{\theta}{}\right) that of the partition function, and \boldsymbol{\theta} that of the inverse temperature.

> Back to ergms (link)

Why we bootstrap

The pooled likelihood treats the 93 networks as independent. They are not: they come from 47 units, most observed more than once.

The procedure

  • Resample the 47 units with replacement – the unit is the block
  • Visits travel with their unit: within-unit dependence is preserved, not modeled
  • Refit the pooled ERGM on each resample – 200 replicates, HPC cluster
  • The spread across replicates provides the standard errors

What it buys

  • No parametric assumption about how repeat visits correlate
  • Absorbs both noise sources: Monte Carlo error and facility heterogeneity
  • Robust where a sandwich-style variance formula would be a guess
  • Embarrassingly parallel – the price is compute

Naive SDs come out 1.5–3× too small – and several naive-significant terms do not survive the bootstrap.

Same model, better coordinates

\begin{aligned} & \theta_1\,\text{edges} \;+\; \theta_2\,\text{edges}\times\text{log}\left\{n\right\} \\[.25em] \longrightarrow\quad & \theta_1^{\ast}\,\text{edges} \;+\; \theta_2\,\text{edges}\times\bigl[\text{log}\left\{n\right\} - \overline{\text{log}\left\{n\right\}}\bigr], \qquad \theta_1^{\ast} = \theta_1 + \theta_2\,\overline{\text{log}\left\{n\right\}} \end{aligned}

The two designs differ by an invertible affine map, so they span the same model: same likelihood surface, same fitted distribution, same fit. Only the coordinates move.

Why bother, then

edges and edges \times\text{log}\left\{n\right\} were nearly collinear: the statistics move together, the information matrix is ill-conditioned, and MC-MLE crawls along a ridge instead of climbing it.

Centering makes them nearly orthogonal. It roughly halved the convergence time and produced far fewer failed fits.

And it fixes the reading

Uncentered, the edges coefficient is the baseline at a one-node facility. Nobody has that facility.

Centered, it is the baseline at an average-size unit and the interaction is the deviation from it – the quantity you report is the quantity you estimated.

Model Terms

Reproduced from Vega Yon, Zhang, et al. (2026). Figures and equations generated using the tabulergm R package (Vega Yon 2026g)

Reproduced from Vega Yon, Zhang, et al. (2026). Figures and equations generated using the tabulergm R package (Vega Yon 2026g)

Go back to the <data challenges>

The countable version

  • >1M cumulative CRAN downloads
  • Network science (rgexf, netplot, netdiffuseR, ergmito) · statistical computing (fmcmc) · HPC (slurmR, Stata parallel)
  • 11 of the 18 CDC-funded R packages on CRAN come from the team I lead at ForeSITE, giving it the largest footprint of any group

ergmitos

Go back to ergms (link)

Bipartite ERGM table

Final pooled bipartite ERGM results, all terms, with naive and bootstrap SDs.

Go back to the key findings (<link>)

Conditional ERGMs

R21 (NLM): “Advancing Exponential-Family Random Graph Model Estimation and Power Analysis for Biomedical and Public Health Research”

\begin{aligned} {\mathbb{P}_{}\left\{\mathbf{Y}= \mathbf{y}; \boldsymbol{\theta}\right\} } = \hphantom{\text{\hspace{3.5cm}}}\\% &= \Prcond{}{\Graph = \graph}{\sufstats{l}{\Graph} = z; \params} \times \Pr{}{\sufstats{l}{\graph} = z; \params}\\ \underbrace{\frac{% \text{exp}\left\{{\boldsymbol{\theta}_{-l}}^\mathbf{t}s_{-l}\left(\mathbf{y}\right)\right\}% }{% \eta_{s_l = z,-l}\left(\boldsymbol{\theta}{}\right) }}_{\text{Conditional ERGM}} \times % \underbrace{\frac{\eta_{s_{k}\left(\mathbf{y}\right)=x}\left(\boldsymbol{\theta}{}\right)}{\eta_{}\left(\boldsymbol{\theta}{}\right)}}_{\text{Marginal ERGM}} \end{aligned}

Go back <here>

Conditional ERGMs (cont.)

How good are conditional ERGM estimates as an initialization point?

Go back <here>

References

Banks, Olivia, and George G. Vega Yon. 2024. Epiworldweb: WebAssembly Port of Epiworld. https://github.com/UofUEpiBio/epiworldweb.
Banks, Olivia, Edward Wang, Richard Nelson, and George G. Vega Yon. 2026. EPICC: Epidemiological Cost Calculator. https://epiforesite.github.io/epicc/.
Bayer, Damon, Dylan H Morris, George G. Vega Yon, Trevor Martin, and Subekshya Bidari. 2024. PyRenew: A Package for Bayesian Renewal Modeling with JAX and NumPyro. https://github.com/CDCgov/PyRenew.
Curtis, Donald E., Christopher S. Hlady, Gaurav Kanade, Sriram V. Pemmaraju, Philip M. Polgreen, and Alberto M. Segre. 2013. “Healthcare Worker Contact Networks and the Prevention of Hospital-Acquired Infections.” PLoS ONE 8 (12): e79906. https://doi.org/10.1371/journal.pone.0079906.
Hâncean, Marian-Gabriel, Matjaž Perc, George G. Vega Yon, Adrian Gheorghiță, and Bianca-Elena Mihăilă. 2022. “The Formation of Political Discussion Networks.” Royal Society Open Science 9 (1): 211609. https://doi.org/10.1098/rsos.211609.
Haye, Kayla de la, Heesung Shin, George G. Vega Yon, and Thomas W. Valente. 2019. “Smoking Diffusion Through Networks of Diverse, Urban American Adolescents over the High School Period.” Journal of Health and Social Behavior 60 (3): 362–76. https://doi.org/10.1177/0022146519870521.
Johnson, Kaitlyn, Dylan Morris, Sam Abbott, et al. 2024. Wwinference: Jointly Infers Infection Dynamics from Wastewater Data and Epidemiological Indicators. https://github.com/cdcgov/ww-inference-model/.
Krivitsky, Pavel N., Pietro Coletti, and Niel Hens. 2023. “Rejoinder to Discussion of ‘A Tale of Two Datasets: Representativeness and Generalisability of Inference for Samples of Networks’.” Journal of the American Statistical Association 118 (544): 2235–38. https://doi.org/10.1080/01621459.2023.2280383.
Lessler, Justin, Claire P. Smith, Praachi Das, et al. 2026. Why Epidemic Risk at the 2026 World Cup May Not Be What You Think. medRxiv, pre-published. https://doi.org/10.64898/2026.05.28.26354384.
Love, Jay, Cormac R. LaPrete, Theresa R. Sheets, et al. 2023. Characterizing Spatiotemporal Variation in Transmission Heterogeneity During the 2022 Mpox Outbreak in the USA. Preprint. Epidemiology. https://doi.org/10.1101/2023.05.10.23289580.
Marye, Abbey, Damon J. A. Toth, George G. Vega Yon, et al. 2026. Hidden Burden of a Measles Outbreak Revealed by Genomic and Transmission Models. medRxiv, pre-published. https://doi.org/10.64898/2026.07.09.26357695.
Meyer, Derek, and George G. Vega Yon. 2023. “epiworldR: Fast Agent-Based Epi Models.” Journal of Open Source Software 8 (90): 5781. https://doi.org/10.21105/joss.05781.
Meyer, Derek, and George G. Vega Yon. 2025. epiworldRShiny: Shiny Interface for epiworldR. https://doi.org/10.32614/CRAN.package.epiworldRShiny.
Milando, Chad W., George G. Vega Yon, Kaitlyn Johnson, et al. 2026. “A Vision for Estimation of the Instantaneous Reproductive Number.” Epidemics 54 (January): 100885. https://doi.org/10.1016/j.epidem.2026.100885.
Mole, Beth. 2025. “Measles Outbreak Erupts in One of Texas’ Least-Vaccinated Counties.” In Ars Technica. https://arstechnica.com/health/2025/02/measles-outbreak-erupts-in-one-of-texas-least-vaccinated-counties/.
Najafzadehkhoei, Sima, George G. Vega Yon, and Bernardo Modenesi. 2026. epiworldRcalibrate: Fast and Effortless Calibration of Agent-Based Models Using Machine Learning. https://doi.org/10.32614/CRAN.package.epiworldRcalibrate.
Najafzadehkhoei, Sima, George G. Vega Yon, Bernardo Modenesi, and Derek S. Meyer. 2025. “Generalized Machine Learning for Fast Calibration of Agent-Based Epidemic Models.” Pre-published September 6. https://doi.org/10.48550/arXiv.2509.07013.
Odoom, Alex, and Eric S. Donkor. 2025. “Prevalence of Healthcare-Acquired Infections Among Adults in Intensive Care Units: A Systematic Review and Meta-Analysis.” Health Science Reports 8 (7): e70939. https://doi.org/10.1002/hsr2.70939.
Ouellet, Marie, Sadaf Hashimi, and George G. Vega Yon. 2022. “Officer Networks and Firearm Behaviors: Assessing the Social Transmission of Weapon-Use.” Journal of Quantitative Criminology 39 (3): 679–703. https://doi.org/10.1007/s10940-022-09546-9.
Piombo, Sarah, George G. Vega Yon, and Thomas W. Valente. 2025. “The Impact of Social Norms on Diffusion Dynamics: A Simulation of e-Cigarette Use Behavior.” Health Education & Behavior 52 (4): 428–38. https://doi.org/10.1177/10901981251327189.
Pittet, Didier. 2026. Key Facts and Figures. World Health Organization. https://www.who.int/campaigns/world-hand-hygiene-day/key-facts-and-figures.
Tanaka, Kyosuke, and George G. Vega Yon. 2024. “Imaginary Network Motifs: Structural Patterns of False Positives and Negatives in Social Networks.” Social Networks 78 (July): 65–80. https://doi.org/10.1016/j.socnet.2023.11.005.
Toth, Damon, Jake Wagoner, Willy Ray, and George G. Vega Yon. 2026. Multigroup.vaccine: Analyze Outbreak Models of Multi-Group Populations with Vaccination. https://epiforesite.github.io/multigroup-vaccine/.
Vega Yon, George G. 2022. Epiworld: A Flexible and General Agent Based Model Engine. https://github.com/UofUEpiBio/epiworld.
Vega Yon, George G. 2023. “Power and Multicollinearity in Small Networks: A Discussion of ‘Tale of Two Datasets: Representativeness and Generalisability of Inference for Samples of Networks’ by Krivitsky, Coletti, and Hens.” Journal of the American Statistical Association 118 (544): 2228–31. https://doi.org/10.1080/01621459.2023.2252041.
Vega Yon, George G. 2026a. Agent-Based Modeling of Measles in US Cities Hosting the 2026 FIFA World Cup. https://epiforesite.github.io/idcup.
Vega Yon, George G. 2026b. Aphylo: Statistical Inference of Annotated Phylogenetic Trees. https://doi.org/10.32614/CRAN.package.aphylo.
Vega Yon, George G. 2026c. Defm: Estimation and Simulation of Multi-Binary Response Models. https://doi.org/10.32614/CRAN.package.defm.
Vega Yon, George G. 2026d. Measles: Measles Epidemiological Models. https://doi.org/10.32614/CRAN.package.measles.
Vega Yon, George G. 2026e. Netplot: Beautiful Graph Drawing. https://doi.org/10.32614/CRAN.package.netplot.
Vega Yon, George G. 2026f. Rgexf: Build, Import and Export GEXF Graph Files. https://doi.org/10.32614/CRAN.package.rgexf.
Vega Yon, George G. 2026g. Tabulergm: Publication-Ready Tables and Summaries for Exponential-Family Random Graph Models. https://doi.org/10.32614/CRAN.package.tabulergm.
Vega Yon, George G., and Olivia Banks. 2026. Epiworldpy: Python Bindings for Epiworld. https://github.com/UofUEpiBio/epiworldpy.
Vega Yon, George G., and Kayla de la Haye. 2025. Ergmito: Exponential Random Graph Models for Small Networks. https://doi.org/10.32614/CRAN.package.ergmito.
Vega Yon, George G., and Paul Marjoram. 2019a. “Fmcmc: A Friendly MCMC Framework.” Journal of Open Source Software 4 (39): 1427. https://doi.org/10.21105/joss.01427.
Vega Yon, George G., and Paul Marjoram. 2019b. “slurmR: A Lightweight Wrapper for HPC with Slurm.” Journal of Open Source Software 4 (42): 1493. https://doi.org/10.21105/joss.01493.
Vega Yon, George G., Mary Jo Pugh, and Thomas W. Valente. 2022. Discrete Exponential-Family Models for Multivariate Binary Outcomes. arXiv:2211.00627. arXiv. https://doi.org/10.48550/arXiv.2211.00627.
Vega Yon, George G., and Brian Quistorff. 2019. “Parallel: A Command for Parallel Computing.” The Stata Journal: Promoting Communications on Statistics and Stata 19 (3): 667–84. https://doi.org/10.1177/1536867X19874242.
Vega Yon, George G., Andrew Slaughter, and Kayla de la Haye. 2021. “Exponential Random Graph Models for Little Networks.” Social Networks 64: 225–38. https://doi.org/10.1016/j.socnet.2020.07.005.
Vega Yon, George G., Duncan C. Thomas, John Morrison, Huaiyu Mi, Paul D. Thomas, and Paul Marjoram. 2021. “Bayesian Parameter Estimation for Automatic Annotation of Gene Functions Using Observational Data and Phylogenetic Trees.” PLOS Computational Biology 17 (2): 1–35. https://doi.org/10.1371/journal.pcbi.1007948.
Vega Yon, George G., Delaney Thornton, Andrew Redd, et al. 2026. Practical Guidelines and Reflections on Building Public Health Software: A Measles Case Study (Preprint). JMIR Public Health; Surveillance. https://doi.org/10.2196/preprints.94240.
Vega Yon, George G., and Thomas Valente. 2026. netdiffuseR: Analysis of Diffusion and Contagion Processes on Networks. https://doi.org/10.32614/CRAN.package.netdiffuseR.
Vega Yon, George G., Chong Zhang, Karim Khader, et al. 2026. “Exponential-Family Random Graph Models in Resident-Healthcare Provider Networks: An Application Using Data from Long-Term Healthcare Facilities in the United States.” Social Networks.
Visnovsky, Lindsay D., Molly Leecaster, George G. Vega Yon, et al. 2026. “P-1104. Multidrug-Resistant Organism (MDRO)-Contaminated Mobile Equipment Networks in Skilled Nursing Facilities.” Open Forum Infectious Diseases 13 (Supplement_1). https://doi.org/10.1093/ofid/ofaf695.1299.