This directory contains the Anthropic Insights data we provided to our three external research groups as part of the collaboration detailed in "Enabling independent research on how people use Claude".
Before using this data, we recommend first reading our blog post on this collaboration and the Anthropic Insights paper and blog post. Before drawing conclusions from this data — especially from open-ended clusters — please read "Guidance for Interpreting Open-Ended Anthropic Insights Clusters" in the blog appendix.
The three research groups were:
Each group's study analyzed roughly 250,000 Claude.ai or Claude Code conversations drawn from a fixed window in April-May 2026. We ran the data collection on their behalf, and they conducted their own independent analysis. This dataset contains the exact outputs each partner received for their research. No raw conversations are included here — only aggregated, privacy-preserving cluster data (see "Privacy and review" below).
stanford — Anthropic Insights data for Stanford's study on how humans collaborate with AI.
oxford — Anthropic Insights data for Oxford's study on people's experience while using Claude and how that relates to Claude's behavior.
metr — Anthropic Insights data for METR's study estimating real-world productivity gains from coding agents and how those gains change across model generations.
metr_addendum — Anthropic Insights data for a follow-up run to METR's original run.
See the blog post for more details, and see the blog appendix for our full privacy threat model and the third-party audit details.
Each row is one cluster: a group of conversations that answered one researcher-defined question (a facet) in a similar way, at one level of a cluster hierarchy.
Base columns:
| Column | Meaning |
|---|---|
cluster_id | Unique identifier for the cluster |
facet_id | Which facet (researcher question) this cluster belongs to |
cluster_name | Short, Claude-generated label for the cluster |
cluster_description | Longer Claude-generated summary of what conversations in the cluster involve |
level | Hierarchy level (0 = most granular; level 1 rows are broader parent clusters) |
num_records | Number of conversations in the cluster |
num_orgs | Number of distinct organizations represented in the cluster |
ratio, ratio_95ci_lower, ratio_95ci_upper | The cluster's share of the run's sample, with a 95% confidence interval |
mean_val, sum_val | Cluster-level summary statistics, populated for numeric facets (not all files carry all of these) |
Cross-facet columns make up the bulk of each file: every cluster is cross-tabulated against the study's other facets. For a categorical facet X with value v, the columns X:v_num_records and X:v_ratio give the number and fraction of the cluster's conversations with that value. Empty cells in cross-facet columns indicate that the intersection was either not computed for that facet pair or fell below the privacy threshold.
Data released under CC BY 4.0.
@online{handa2026enablingindependentresearch,
author = {Kunal Handa and Miranda Zhang and Gabriel Nicholas and Miles McCain and Ryan Heller and Saffron Huang and Thomas Millar and Suzanne Wang and Shan Carter and Mo Julapalli and Matt Kearney and Sarah Pollack and Judy Shen and Matthew Jagielski and Shaoyi Zhang and Heather Whitney and Ankur Rathi and Aisling Keenan and David Saunders and Jake Eaton and Sylvie Carr and Jack Clark and Michael Stern and Deep Ganguli},
title = {Enabling independent research on how people use Claude},
date = {2026-08-26},
year = {2026},
url = {https://www.anthropic.com/research/enabling-independent-research},
}
You can submit inquiries to kunal@anthropic.com. We invite researchers to express interest in potential future Societal Impacts external researcher collaborations using this form.
4 commits
This directory contains the Anthropic Insights data we provided to our three external research groups as part of the collaboration detailed in "Enabling independent research on how people use Claude".
Before using this data, we recommend first reading our blog post on this collaboration and the Anthropic Insights paper and blog post. Before drawing conclusions from this data — especially from open-ended clusters — please read "Guidance for Interpreting Open-Ended Anthropic Insights Clusters" in the blog appendix.
The three research groups were:
Each group's study analyzed roughly 250,000 Claude.ai or Claude Code conversations drawn from a fixed window in April-May 2026. We ran the data collection on their behalf, and they conducted their own independent analysis. This dataset contains the exact outputs each partner received for their research. No raw conversations are included here — only aggregated, privacy-preserving cluster data (see "Privacy and review" below).
stanford — Anthropic Insights data for Stanford's study on how humans collaborate with AI.
oxford — Anthropic Insights data for Oxford's study on people's experience while using Claude and how that relates to Claude's behavior.
metr — Anthropic Insights data for METR's study estimating real-world productivity gains from coding agents and how those gains change across model generations.
metr_addendum — Anthropic Insights data for a follow-up run to METR's original run.
See the blog post for more details, and see the blog appendix for our full privacy threat model and the third-party audit details.
Each row is one cluster: a group of conversations that answered one researcher-defined question (a facet) in a similar way, at one level of a cluster hierarchy.
Base columns:
| Column | Meaning |
|---|---|
cluster_id | Unique identifier for the cluster |
facet_id | Which facet (researcher question) this cluster belongs to |
cluster_name | Short, Claude-generated label for the cluster |
cluster_description | Longer Claude-generated summary of what conversations in the cluster involve |
level | Hierarchy level (0 = most granular; level 1 rows are broader parent clusters) |
num_records | Number of conversations in the cluster |
num_orgs | Number of distinct organizations represented in the cluster |
ratio, ratio_95ci_lower, ratio_95ci_upper | The cluster's share of the run's sample, with a 95% confidence interval |
mean_val, sum_val | Cluster-level summary statistics, populated for numeric facets (not all files carry all of these) |
Cross-facet columns make up the bulk of each file: every cluster is cross-tabulated against the study's other facets. For a categorical facet X with value v, the columns X:v_num_records and X:v_ratio give the number and fraction of the cluster's conversations with that value. Empty cells in cross-facet columns indicate that the intersection was either not computed for that facet pair or fell below the privacy threshold.
Data released under CC BY 4.0.
@online{handa2026enablingindependentresearch,
author = {Kunal Handa and Miranda Zhang and Gabriel Nicholas and Miles McCain and Ryan Heller and Saffron Huang and Thomas Millar and Suzanne Wang and Shan Carter and Mo Julapalli and Matt Kearney and Sarah Pollack and Judy Shen and Matthew Jagielski and Shaoyi Zhang and Heather Whitney and Ankur Rathi and Aisling Keenan and David Saunders and Jake Eaton and Sylvie Carr and Jack Clark and Michael Stern and Deep Ganguli},
title = {Enabling independent research on how people use Claude},
date = {2026-08-26},
year = {2026},
url = {https://www.anthropic.com/research/enabling-independent-research},
}
You can submit inquiries to kunal@anthropic.com. We invite researchers to express interest in potential future Societal Impacts external researcher collaborations using this form.
4 commits