ZC_5_16

Computational Social Science: Big Data, Agent-Based Models, and Digital Behavioral Analysis

Verified (Tier 1)
Confidence: 4/5 Section: ZC Updated: April 1, 2026
Source Count: 13 | Weighted Score: 33 | Source Confidence: [4/5] | Primary Tier: 1 | Last Updated: April 1, 2026
Keywords: computational social science, big data, agent-based modeling, social network analysis, digital trace data, natural language processing, machine learning, social simulation, Twitter, social media analytics, causal inference, algorithmic bias
Category Tags: computational-social-science, big-data, social-network-analysis, agent-based-modeling, digital-methods, social-media-research
Cross-References: ZC_5_01 — Modern Social Science Methods · V_4_01 — Computer Science Foundations · ZD_3_05 — Machine Learning & Neural Networks

QUICK SUMMARY

Computational social science (CSS) is the interdisciplinary field that applies computational methods — machine learning, natural language processing, network analysis, agent-based modeling, and large-scale data mining — to study human behavior, social structures, and cultural dynamics. The field was formalized by David Lazer and 14 co-authors in a landmark 2009 Science essay declaring that "the capacity to collect and analyze massive amounts of data has transformed the social sciences." CSS operates on two frontiers: (1) observational analysis of digital trace data — the behavioral exhaust generated by billions of people using social media, mobile phones, web searches, financial transactions, and GPS-enabled devices, providing unprecedented real-time, population-scale records of human activity; and (2) computational simulation — agent-based models (ABMs) and microsimulation that generate emergent social phenomena (segregation, cooperation, market dynamics, epidemic spread) from individual-level behavioral rules. Key achievements include Jon Kleinberg and Duncan Watts's work on social network structure (small-world networks), Matthew Salganik's experimental studies of cultural markets and unpredictability, Sinan Aral's analysis of social contagion and misinformation spread, and the use of NLP-based sentiment analysis to predict election outcomes, stock markets, and public health trends. CSS has also generated significant controversy over research ethics (the Facebook emotional contagion experiment, 2014), algorithmic bias, surveillance implications, and the "replication crisis" in data-driven social research.


1. VERIFIED CLAIMS (Tier 1 — Peer-Reviewed / Established)

1.1 The Founding Manifesto: Lazer et al. (2009)

1.2 Social Network Analysis: Small Worlds and Scale-Free Networks

1.3 Agent-Based Modeling: Emergent Social Phenomena

1.4 Digital Trace Data and Behavioral Prediction


2. CREDIBLE CLAIMS (Tier 2 — Academic / Debated but Supported)

2.1 Social Contagion and Information Spread

2.2 Experimental Approaches: The Salganik Music Lab

2.3 NLP and Text-as-Data in Political Science


3. SPECULATIVE CLAIMS (Tier 3 — Possible but Unverified)

3.1 Predictive Modeling of Social Upheaval

3.2 Digital Twins of Societies


4. DUBIOUS CLAIMS (Tier 4 — No Credible Source / Contradicted by Evidence)

4.1 Social Media Data as Unbiased Population Samples


Counter-Arguments & Criticisms

The Facebook emotional contagion experiment (Kramer, Guillory, and Hancock, 2014) — which manipulated the News Feeds of 689,003 Facebook users without informed consent to test whether exposure to positive or negative content affected users' subsequent posts — generated widespread criticism from scientists, ethicists, and the public. Michelle Meyer (2014) and others argued that the study violated basic research ethics principles, raising fundamental questions about the ethics of large-scale behavioral experiments conducted by technology companies on their users.

Gary King and colleagues (Harvard, 2014) have cautioned about the "big data hubris" — the assumption that large datasets automatically produce better insights. They demonstrated that Google Flu Trends (a prominent early CSS project that predicted flu prevalence from search query data) systematically over-predicted flu activity after 2011, ultimately performing worse than simple autoregressive models. The failure illustrated that correlation-based prediction without causal understanding is fragile and subject to concept drift.


IMAGES

#DescriptionFilenameSourceLicense
1Social network visualization of Twitter retweet clusterstwitter_network_visualization.jpgWikimedia CommonsCC BY-SA 4.0
2Schelling segregation model simulation resultsschelling_segregation_model.jpgWikimedia CommonsPD
3Agent-based model of epidemic spreadabm_epidemic_spread.jpgWikimedia CommonsCC BY-SA 4.0

BIBLIOGRAPHY

  1. Lazer, David, Alex Pentland, Lada Adamic, et al | 2009 | "Computational Social Science" | Science | ∅ | 323.5915::721–723 | ∅ | ∅ | doi:10.1126/science.1167742 | ∅ | ∅ | ∅
  2. Watts, Duncan J.; Steven H | 1998 | "Collective Dynamics of 'Small-World' Networks" | Nature | ∅ | 393.6684::440–442 | Strogatz | ∅ | doi:10.1038/30918 | ∅ | ∅ | ∅
  3. Schelling, Thomas C | 1971 | "Dynamic Models of Segregation" | Journal of Mathematical Sociology | ∅ | 1.2::143–186 | ∅ | ∅ | doi:10.1080/0022250X.1971.9989794 | ∅ | ∅ | ∅
  4. Kosinski, Michal, David Stillwell; Thore Graepel | 2013 | "Private Traits and Attributes Are Predictable from Digital Records of Human Behavior" | Proceedings of the National Academy of Sciences | ∅ | 110.15::5802–5805 | ∅ | ∅ | doi:10.1073/pnas.1218772110 | ∅ | ∅ | ∅
  5. Vosoughi, Soroush, Deb Roy; Sinan Aral | 2018 | "The Spread of True and False News Online" | Science | ∅ | 359.6380::1146–1151 | ∅ | ∅ | doi:10.1126/science.aap9559 | ∅ | ∅ | ∅
  6. Salganik, Matthew J., Peter Sheridan Dodds; Duncan J | 2006 | "Experimental Study of Inequality and Unpredictability in an Artificial Cultural Market" | Science | ∅ | 311.5762::854–856 | Watts | ∅ | doi:10.1126/science.1121066 | ∅ | ∅ | ∅
  7. Grimmer, Justin; Brandon M | 2013 | "Text as Data: The Promise and Pitfalls of Automatic Content Analysis Methods for Political Texts" | Political Analysis | ∅ | 21.3::267–297 | Stewart | ∅ | doi:10.1093/pan/mps028 | ∅ | ∅ | ∅
  8. Barabási, Albert-László; Réka Albert | 1999 | "Emergence of Scaling in Random Networks" | Science | ∅ | 286.5439::509–512 | ∅ | ∅ | doi:10.1126/science.286.5439.509 | ∅ | ∅ | ∅
  9. Epstein, Joshua M.; Robert Axtell | 1996 | ∅ | Growing Artificial Societies: Social Science from the Bottom Up | ∅ | ∅ | Washington, DC: Brookings Institution Press | ∅ | isbn:9780262550253 | ∅ | ∅ | ∅
  10. Salganik, Matthew J. | 2018 | ∅ | Bit by Bit: Social Research in the Digital Age | ∅ | ∅ | Princeton: Princeton University Press | ∅ | isbn:9780691158648 | ∅ | ∅ | ∅
  11. Ruths, Derek; Jürgen Pfeffer | 2014 | "Social Media for Large Studies of Behavior" | Science | ∅ | 346.6213::1063–1064 | ∅ | ∅ | doi:10.1126/science.346.6213.1063 | ∅ | ∅ | ∅
  12. Aral, Sinan | 2020 | ∅ | The Hype Machine: How Social Media Disrupts Our Elections, Our Economy, and Our Health — and How We Must Adapt | ∅ | ∅ | New York: Currency | ∅ | isbn:9780525574514 | ∅ | ∅ | ∅
  13. Lazer, David, Ryan Kennedy, Gary King; Alessandro Vespignani | 2014 | "The Parable of Google Flu: Traps in Big Data Analysis" | Science | ∅ | 343.6176::1203–1205 | ∅ | ∅ | doi:10.1126/science.1248506 | ∅ | ∅ | ∅

CROSS-REFERENCE INDEX

Related DocConnection
ZC_5_01Social science methodological framework
V_4_01Computer science foundations underlying CSS methods
ZD_3_05Machine learning methods used in CSS
ZG_2_05NLP and text analysis methods in CSS
T_3_03Social media's psychological effects studied through CSS methods

Generated from V4 expansion plan. Last Updated: April 1, 2026