← Back to explorer

Capability Provenance in Language Models: A Case Study in Social Reasoning

Type
paper
Venue
Proceedings of the Conference on Language Modeling (COLM 2026); arXiv:2606.19625
Year
2026
Source
web
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

Traces language-model capabilities back to the training text that taught them, using social reasoning as a case study. Influence functions score how much each training document supported each benchmark answer; per-query influences are aggregated into 576 topic-format bins (24 topics x 24 formats via the WebOrganizer taxonomy) to compare capability-level data distributions. A 2x2 design separates domain (social vs STEM) from capability type (reasoning vs knowledge), and unlearning the high-influence regions causally validates the attribution.

Keywords

interpretability · training-data attribution · influence functions · social reasoning · OLMo · unlearning

Topics

interpretability, training-data attribution, influence functions, social reasoning, OLMo

Research notes

  • Discovery: Shared together with https://allenai.org/blog/olmo-capability-tracing (AllenAI blog covering the same Georgia Tech work); treated as one item.
  • Method: Influence-function attribution per benchmark query, aggregated over 576 topic-format corpus bins; contrastive 2x2 design (social/STEM x reasoning/knowledge) on OLMo 3 trained on Dolma 3; causal check by unlearning high-influence regions vs random controls; expansion to theory of mind, moral judgment, social bias, pragmatics probes and a family of open-data models (Marin, DCLM, Common Pile).
  • Key findings: Each capability draws on its own distribution of training text, not one shared pile; Social reasoning is carried by two flavors of text: interactional writing (dialogue, Q&A) and people-centered documentation (guides, manuals); STEM reasoning draws on structured technical text; Capability type differentiates corpus provenance more than domain does (social-minus-STEM contrast ~1.4x wider for reasoning than knowledge); Influence-targeted unlearning damages social reasoning far more than random controls, with weak/absent effects on other benchmarks
  • Limitations: Unlearning validates corpus regions in aggregate, not individually causal documents; which benchmark carries the effect can shift with the training corpus.
  • Year set to 2026 per COLM 2026 venue and arXiv ID 2606.19625; exact date not verified. Fully open artifacts: OLMo 3, Dolma 3, WebOrganizer, OLMo 3 eval suite, OLMES.