Reference for the conventions used in the YAML term database under
inst/terms/. Follow these standards when adding or editing term
definitions so that the math and figures stay consistent across the
term dictionary.
Files
Each term lives in inst/terms/<term>.<directed|undirected>.yml, where
<term> is the canonical ergm term name as written in a model formula
(e.g., gwesp, b1nodematch). Terms available for both directed and
undirected networks get one file per variant; bipartite terms use
undirected. A file has up to five top-level entries: title and
description (text, both optional), citation (optional, see below),
plot (the network drawing specification), and math (a LaTeX
expression). No parser changes are needed for new terms: files are
looked up by term name.
Aliases
When two ergm term names share an implementation (e.g. dgwesp and
gwesp), the second file can reuse the first with an alias entry
instead of copying it. The target is the file of the same directedness,
so dgwesp.directed.yml below reuses gwesp.directed.yml:
alias: gwespAny other entry in an alias file overrides the target's. plot is
merged field by field (so an alias can, say, recolor one vertex without
repeating the whole drawing); every other entry replaces the target's
entry of the same name:
alias: gwesp
title: Typed geometrically weighted edgewise shared partners
plot:
ecolor: gray
Titles, descriptions, and citations
title is a short label, capitalized like a heading and without a
trailing period (e.g. Uniform homophily). description is one to
three sentences of prose saying what the statistic counts and why a
modeler would include it; write it as a folded block scalar (>-) so
the source stays readable. Both fields are optional: when absent,
tabulergm falls back to the title and description recorded in the
ergm term database, which are accurate but often too long, too
technical, or full of raw LaTeX for a table cell. Prefer writing them
here.
citation records the source(s) that introduced the term. Each entry
carries a key (the marker shown in the table, conventionally
lastnameYEAR) and, where one exists, a machine-readable identifier so
readers can pull the full reference into their own bibliography:
citation:
- key: hunter2007
doi: 10.1016/j.socnet.2006.08.005Cite the paper that introduced the statistic in the ERGM/p* framework
first (for example Wasserman and Pattison 1996 for nodematch); when two
papers introduced a parameterization together, list both. After the
origin, a term may carry at most one well-established substantive
reference for the concept it measures (for example McPherson et al. 2001
on homophily). A term with no identifiable ERGM origin stays uncited
rather than borrowing a loosely related source.
Accepted identifier fields are doi, arxiv, pmid, and url; an
entry may also carry free-text text, which is the right choice when
no stable identifier can be verified. Always resolve an
identifier before committing it (for a DOI, https://doi.org/<id>)
and check that its authors and year match the key; an identifier that
points at the wrong paper is worse than none. Use the arxiv field (not
an arXiv DOI) for preprints, and reuse the same key and identifier for a
reference cited by several terms.
tabulergm_table() renders a (key) marker next to the term's
description and a matching [key] identifier line below the table. Users
can replace any of these fields per table through the override
arguments of tabulergm_table() without editing the YAML.
Math notation
\(y_{ij}\): tie indicator. Undirected statistics sum over \(i<j\); directed statistics sum over \(i \neq j\).
\(x_i\): vertex attribute value of node \(i\); \(x_{ij}\): dyadic covariate value (e.g.,
edgecov).\(\mathbf{1}(\cdot)\): indicator function (
\mathbf{1}in LaTeX).Attribute levels: \(k\) for a single level (factor terms), \((k, l)\) for mixing pairs, and \((p, q)\) for center/leaf values in star-mixing terms.
Bipartite modes: \(B_1\) (first mode) and \(B_2\) (second mode); \(n_{B_1}\) and \(n_{B_2}\) for the mode sizes.
Geometrically weighted terms follow the Hunter (2007) parameterization, e.g. \(\exp(\tau) \sum_i [1 - (1 - e^{-\tau})^i] EP_i(y)\), where the exponent is the summation index and \(EP_i\), \(DP_i\), and \(D_i\) are the edgewise shared partner, dyadwise shared partner, and degree counts; directed and bipartite degree counts carry a superscript naming what is counted, \(D^{\mathrm{in}}_i\), \(D^{\mathrm{out}}_i\), \(D^{B_1}_i\), and \(D^{B_2}_i\) (e.g. the number of first-mode nodes with degree \(i\)).
Degree-based terms write node degree as a sum of ties, e.g. \(\sum_{j \neq i} y_{ij}\) (undirected) or \(\sum_{j \neq i} y_{ji}\) (in-degree), and k-star counts as binomial coefficients, \(\binom{\cdot}{k}\).
Directed shared-partner terms carry the two-path type as a superscript, e.g. \(EP^{\mathrm{OTP}}_i\), because
ergmcounts outgoing two-paths (OTP) by default andgwesp/gwdsptake atypeargument that changes which two-paths are counted.
Always verify a new formula against the ergm term manual (see
ergm::ergmTerm and ergm::search.ergmTerms()) and the source
literature. When the manual is ambiguous, compare the formula
numerically against summary(nw ~ term) on a small test network.
Drawing conventions
The plot entry supports edgelist, vcolor, vshape, vsize,
ecolor, elinetype, and layout (with x and y coordinates).
Edgelists are chains like "0->1->2, 0->3": each consecutive pair is
one edge; a lone node id (e.g. "1->2, 0") adds an isolated node.
Per-vertex vectors follow the node order obtained from the parsed
edgelist (unique node ids, all tail nodes first, then head nodes, then
isolated nodes); render the figure to double-check the mapping.
Vertex color:
blackmarks the focal structure of a term;graymarks non-focal context, both attribute-irrelevant nodes and structurally non-focal nodes (e.g., the shared partners ingwesp/gwdsp);orangemarks nodes whose attribute enters the statistic (matched pairs share orange); mixing terms useorangevs. teal ("#008080", quoted in the YAML) for the two attribute categories, a colorblind-friendly pairing. These colors drive the explanatory notes appended below rendered tables, so use them consistently.Vertex shape:
squaremarks first-mode (\(B_1\)) nodes andcirclemarks second-mode (\(B_2\)) nodes in bipartite drawings. One-mode drawings use circles only.Layout: bipartite drawings place the first mode on the left and the second mode on the right.
Vertex size:
1.0for focal or attribute-relevant nodes,.5for context nodes.Edges: solid black for focal ties;
grayfor context ties (e.g., the two-paths in shared-partner terms); dashed (elinetype: 2) for match/covariate annotations;orangewhen the edge itself carries the covariate (e.g.,edgecov).Directedness comes from the file name; arrows are drawn automatically for
.directed.ymlterms.
Checklist for a new term
Add the YAML file(s) following the standards above, including a
titleanddescription, and acitationwhen the term has an identifiable source.Add tinytest coverage in
inst/tinytest/test_term_db.R(file lookup, math/figure reading, and formula integration).Add the term to the dictionary tables in
README.qmdandvignettes/ergm-with-tabulergm.Rmd; both contain a hidden coverage-check chunk that fails the render if a term is missing.Re-render the README, and run the tests and
devtools::check().
References
Hunter, D. R. (2007). Curved exponential family models for social networks. Social Networks, 29(2), 216–230. doi:10.1016/j.socnet.2006.08.005
Bomiriya, R. P., Bansal, S., & Hunter, D. R. (2014). Modeling homophily in ERGMs for bipartite networks. (No stable identifier verified; the arXiv id previously given here, 1412.1151, belongs to an unrelated paper.)