Closure, connectivity and degree distributions: Exponential random graph (p*) models for directed social networks
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This paper proposes three new triadic-based parameters to represent different versions of triadic closure, and shows that, for some datasets, the path shortening parameter is insufficient for practical modeling, whereas the structural homophily parameters can produce useful models with distinctive interpretations.
Abstract
The new higher order specifications for exponential random graph models introduced by Snijders, Pattison, Robins & Handcock (2006) exhibit dramatic improvements in model fit compared with the commonly used Markov random graphs. Snijders et al briefly presented versions of these new specifications for directed graphs, in particular a directed alternating ktriangle parameter, based on closure of multiple two-paths. In this paper, we present a number of additional higher order parameters for directed graphs. Most importantly, we propose three new triadic-based parameters to represent different versions of triadic closure: cyclic effects; transitivity based on shared choices of partners; and transitivity based on shared popularity. We also introduce corresponding parameters for multiple connectivity effects. We propose some fifty graph features to be investigated in goodness of fit diagnostics for these new parameters. As empirical illustrations, we develop models for two sets of organizational network data, to show that the new parameters help with an optimal representation of the data. The first example is a trust network within a training group, and the second a 'work difficulty' network within a government instrumentality. In the first example we show that our additional parameters are necessary to obtain an acceptable model for the data. The second example is novel in fitting a statistical model, and inferring structural processes, for a negative tie network. Using this second example, we show how the incorporation of additional effects - the number of sources and sinks in the network, and the correlation between the in- and out-degree distributions - can improve representation of the degree distribution. The final model acceptably replicates the negative tie network in terms of: statistics related to twenty different graph configurations; the in- and out-degree distribution, including their correlation; seven different graph clustering coefficients; the triad census; and the geodesic distribution. Model interpretation emphasizes the importance of some nodes receiving high numbers of negative ties.
