CIRCL/vulnerability-attack-techniques

Dataset

1

stars

8

commits

5

linked in READMEs

Aug 11, 2026

updated

attack-techniques
cve
cybersecurity
mitre-attack
security
vulnerability

README

vulnerability-attack-techniques

This dataset maps 1,207 CVEs to MITRE ATT&CK (Enterprise) techniques, joining hand-curated mappings from the MITRE Center for Threat-Informed Defense (CTID) with vulnerability descriptions from CIRCL/vulnerability-scores. It is intended for training and evaluating models that suggest candidate ATT&CK techniques from a vulnerability description: CVSS tells you how bad a vulnerability is, CWE what kind of flaw it is — ATT&CK tells defenders what adversary behavior to expect and detect.

Every label in the techniques column was written by an analyst following the CTID "Mapping ATT&CK to CVE for Impact" methodology, which assigns each CVE up to three kinds of techniques: an exploitation technique (how it is exploited), a primary impact (what exploitation directly yields), and a secondary impact (what the attacker can do next).

This is the gold set of the paper Mapping CVEs to MITRE ATT&CK Techniques: A Curated Gold-Set Classifier and the Limits of LLM-Assisted Label Expansion (arXiv:2607.25572). The classifier trained on it, CIRCL/vulnerability-attack-technique-classification-roberta-base, runs in production on Vulnerability-Lookup.

DOI: 10.57967/hf/9621

Label sources

label_sourcesCVEsOrigin
ctid_cve788attack_to_cve (2021), ATT&CK v9 era
ctid_kev392Mappings Explorer KEV mappings, ATT&CK 16.1
both27

All technique IDs are normalized to enterprise ATT&CK v19.1: techniques revoked since the original mappings are remapped to their successor via the STIX revoked-by relationships (e.g. T1562 Impair Defenses → T1685 Disable or Modify Tools), and Mobile/ICS techniques are dropped (enterprise domain only).

⚠️ techniques vs techniques_derived

The techniques_derived column contains labels from the automatically derived CVE → CWE → CAPEC → ATT&CK chain maintained by CVE2CAPEC. Do not train on this column. Analysis of the chain shows a median fan-out of 4–20 techniques per CVE and top-frequency techniques (e.g. T1574.007 on 53% of 2024 CVEs) that are artifacts of the cross-framework table expansion, not descriptions of real adversary behavior. The column is included as:

  1. a baseline that a trained model must beat;
  2. a comparison column for studying where the deterministic chain diverges from analyst judgment.

Its use as an inference-time candidate prior was measured and rejected (2026-08-06): at the parent-technique level the derived candidate sets cover only 3.3% of the analyst-chosen techniques on the test split, so any re-ranking toward them degrades every ranking metric.

The full source analysis is documented in the VulnTrain documentation.

Fields

FieldTypeDescription
idstringCVE identifier
titlestringVulnerability title
descriptionstringVulnerability description in English (model input)
exploitation_techniqueslist[string]CTID exploitation technique(s)
primary_impactlist[string]CTID primary impact technique(s)
secondary_impactlist[string]CTID secondary impact technique(s)
techniqueslist[string]Union of all curated techniques — the training target
techniques_derivedlist[string]CVE2CAPEC weak labels — not for training
label_sourceslist[string]ctid_cve and/or ctid_kev
attack_versionstringEnterprise ATT&CK version the IDs are normalized to
cvss_vectorstringCVSS vector string, highest available version (empty if none) — v2
cvss_versionstringVersion of cvss_vector: 4.0, 3.1, 3.0 or 2.0 — v2
cweslist[string]CWE assignments, e.g. CWE-502 Deserialization of Untrusted Data — v2
affected_productslist[string]vendor product pairs from the CVE record — v2
cpeslist[string]CPE identifiers — v2
cwes_predictedlist[string]Top-1 prediction of CIRCL/cwe-parent-vulnerability-classification-roberta-base on the description, same display format as cwes — v2.1

Structured metadata columns (v2, added 2026-08-06)

The v2 columns are extracted from the raw CVE records served by Vulnerability-Lookup (CNA container preferred, CISA ADP Vulnrichment filling many gaps — notably 100% CVSS/CWE coverage on the KEV subset); cpes is joined from CIRCL/vulnerability-scores. v1 columns are unchanged (the update is strictly additive: identical rows and splits). Coverage differs by label source — report results stratified by label_sources when using these columns as model inputs:

SubsetCVEscvss_vectorcwesaffected_productscpes
all1,20772.0%84.3%67.4%93.2%
ctid_kev392100%100%79.8%79.1%
ctid_cve78857.1%76.0%62.2%100%
both27100%100%40.7%100%

CVSS versions among the 869 vectors: 677 × v3.1, 173 × v3.0, 18 × v4.0, 1 × v2.0.

Predicted CWE column (v2.1, added 2026-08-08)

cwes_predicted holds the top-1 output of the deployed CIRCL CWE guesser (CIRCL/cwe-parent-vulnerability-classification-roberta-base, parent-level, 303 classes) run on each row's title + description. Coverage is 100% by construction; agreement with the gold cwes column (ancestor level, on the 814 rows whose gold entry carries a parseable CWE id) is 27.3% top-1. The column exists to measure the cascade cost of replacing gold CWE input with a model prediction in downstream CVE→ATT&CK classifiers; it is a model output, not curated ground truth — do not use it as labels. v1/v2 columns are unchanged (strictly additive update).

Label statistics

192 distinct techniques; 66 with at least 5 examples. Most CVEs carry 1–3 techniques. Top techniques: T1190 Exploit Public-Facing Application (348), T1059 Command and Scripting Interpreter (262), T1203 Exploitation for Client Execution (213), T1068 Exploitation for Privilege Escalation (189).

Known limitations

  • Size: ~1,200 CVEs supports a proof-of-concept, not a production model.
  • Selection bias: both label sources over-represent exploited-in-the-wild vulnerabilities (the KEV set by construction).
  • Inherent task ceiling: a CVE description describes a flaw, while ATT&CK describes attacker behavior around it — even human annotators disagree on such mappings. Models trained on this data should suggest candidate techniques for analyst review, not produce authoritative mappings.

Usage

from datasets import load_dataset

dataset = load_dataset("CIRCL/vulnerability-attack-techniques")

for entry in dataset["train"].select(range(3)):
    print(entry["id"], entry["techniques"], "-", entry["description"][:80])

Licensing of upstream sources

The CTID mappings are Apache-2.0. Descriptions come from CIRCL/vulnerability-scores (CC BY 4.0). The techniques_derived column is derived from the GPLv3 CVE2CAPEC project. MITRE ATT&CK® is a registered trademark of The MITRE Corporation; ATT&CK content is used in accordance with the MITRE ATT&CK terms of use.

References

Citation

@misc{bonhomme2026mappingcvesmitreattck,
      title={Mapping CVEs to MITRE ATT&CK Techniques: A Curated Gold-Set Classifier and the Limits of LLM-Assisted Label Expansion},
      author={Cédric Bonhomme and Alexandre Dulaunoy},
      year={2026},
      eprint={2607.25572},
      archivePrefix={arXiv},
      primaryClass={cs.CR},
      url={https://arxiv.org/abs/2607.25572},
}

Acknowledgements

Developed at CIRCL in the context of the AIPITCH project, co-funded by the European Union.

Contributors

CIRCL/vulnerability-attack-techniques

Dataset

1

stars

8

commits

5

linked in READMEs

Aug 11, 2026

updated

attack-techniques
cve
cybersecurity
mitre-attack
security
vulnerability

README

vulnerability-attack-techniques

This dataset maps 1,207 CVEs to MITRE ATT&CK (Enterprise) techniques, joining hand-curated mappings from the MITRE Center for Threat-Informed Defense (CTID) with vulnerability descriptions from CIRCL/vulnerability-scores. It is intended for training and evaluating models that suggest candidate ATT&CK techniques from a vulnerability description: CVSS tells you how bad a vulnerability is, CWE what kind of flaw it is — ATT&CK tells defenders what adversary behavior to expect and detect.

Every label in the techniques column was written by an analyst following the CTID "Mapping ATT&CK to CVE for Impact" methodology, which assigns each CVE up to three kinds of techniques: an exploitation technique (how it is exploited), a primary impact (what exploitation directly yields), and a secondary impact (what the attacker can do next).

This is the gold set of the paper Mapping CVEs to MITRE ATT&CK Techniques: A Curated Gold-Set Classifier and the Limits of LLM-Assisted Label Expansion (arXiv:2607.25572). The classifier trained on it, CIRCL/vulnerability-attack-technique-classification-roberta-base, runs in production on Vulnerability-Lookup.

DOI: 10.57967/hf/9621

Label sources

label_sourcesCVEsOrigin
ctid_cve788attack_to_cve (2021), ATT&CK v9 era
ctid_kev392Mappings Explorer KEV mappings, ATT&CK 16.1
both27

All technique IDs are normalized to enterprise ATT&CK v19.1: techniques revoked since the original mappings are remapped to their successor via the STIX revoked-by relationships (e.g. T1562 Impair Defenses → T1685 Disable or Modify Tools), and Mobile/ICS techniques are dropped (enterprise domain only).

⚠️ techniques vs techniques_derived

The techniques_derived column contains labels from the automatically derived CVE → CWE → CAPEC → ATT&CK chain maintained by CVE2CAPEC. Do not train on this column. Analysis of the chain shows a median fan-out of 4–20 techniques per CVE and top-frequency techniques (e.g. T1574.007 on 53% of 2024 CVEs) that are artifacts of the cross-framework table expansion, not descriptions of real adversary behavior. The column is included as:

  1. a baseline that a trained model must beat;
  2. a comparison column for studying where the deterministic chain diverges from analyst judgment.

Its use as an inference-time candidate prior was measured and rejected (2026-08-06): at the parent-technique level the derived candidate sets cover only 3.3% of the analyst-chosen techniques on the test split, so any re-ranking toward them degrades every ranking metric.

The full source analysis is documented in the VulnTrain documentation.

Fields

FieldTypeDescription
idstringCVE identifier
titlestringVulnerability title
descriptionstringVulnerability description in English (model input)
exploitation_techniqueslist[string]CTID exploitation technique(s)
primary_impactlist[string]CTID primary impact technique(s)
secondary_impactlist[string]CTID secondary impact technique(s)
techniqueslist[string]Union of all curated techniques — the training target
techniques_derivedlist[string]CVE2CAPEC weak labels — not for training
label_sourceslist[string]ctid_cve and/or ctid_kev
attack_versionstringEnterprise ATT&CK version the IDs are normalized to
cvss_vectorstringCVSS vector string, highest available version (empty if none) — v2
cvss_versionstringVersion of cvss_vector: 4.0, 3.1, 3.0 or 2.0 — v2
cweslist[string]CWE assignments, e.g. CWE-502 Deserialization of Untrusted Data — v2
affected_productslist[string]vendor product pairs from the CVE record — v2
cpeslist[string]CPE identifiers — v2
cwes_predictedlist[string]Top-1 prediction of CIRCL/cwe-parent-vulnerability-classification-roberta-base on the description, same display format as cwes — v2.1

Structured metadata columns (v2, added 2026-08-06)

The v2 columns are extracted from the raw CVE records served by Vulnerability-Lookup (CNA container preferred, CISA ADP Vulnrichment filling many gaps — notably 100% CVSS/CWE coverage on the KEV subset); cpes is joined from CIRCL/vulnerability-scores. v1 columns are unchanged (the update is strictly additive: identical rows and splits). Coverage differs by label source — report results stratified by label_sources when using these columns as model inputs:

SubsetCVEscvss_vectorcwesaffected_productscpes
all1,20772.0%84.3%67.4%93.2%
ctid_kev392100%100%79.8%79.1%
ctid_cve78857.1%76.0%62.2%100%
both27100%100%40.7%100%

CVSS versions among the 869 vectors: 677 × v3.1, 173 × v3.0, 18 × v4.0, 1 × v2.0.

Predicted CWE column (v2.1, added 2026-08-08)

cwes_predicted holds the top-1 output of the deployed CIRCL CWE guesser (CIRCL/cwe-parent-vulnerability-classification-roberta-base, parent-level, 303 classes) run on each row's title + description. Coverage is 100% by construction; agreement with the gold cwes column (ancestor level, on the 814 rows whose gold entry carries a parseable CWE id) is 27.3% top-1. The column exists to measure the cascade cost of replacing gold CWE input with a model prediction in downstream CVE→ATT&CK classifiers; it is a model output, not curated ground truth — do not use it as labels. v1/v2 columns are unchanged (strictly additive update).

Label statistics

192 distinct techniques; 66 with at least 5 examples. Most CVEs carry 1–3 techniques. Top techniques: T1190 Exploit Public-Facing Application (348), T1059 Command and Scripting Interpreter (262), T1203 Exploitation for Client Execution (213), T1068 Exploitation for Privilege Escalation (189).

Known limitations

  • Size: ~1,200 CVEs supports a proof-of-concept, not a production model.
  • Selection bias: both label sources over-represent exploited-in-the-wild vulnerabilities (the KEV set by construction).
  • Inherent task ceiling: a CVE description describes a flaw, while ATT&CK describes attacker behavior around it — even human annotators disagree on such mappings. Models trained on this data should suggest candidate techniques for analyst review, not produce authoritative mappings.

Usage

from datasets import load_dataset

dataset = load_dataset("CIRCL/vulnerability-attack-techniques")

for entry in dataset["train"].select(range(3)):
    print(entry["id"], entry["techniques"], "-", entry["description"][:80])

Licensing of upstream sources

The CTID mappings are Apache-2.0. Descriptions come from CIRCL/vulnerability-scores (CC BY 4.0). The techniques_derived column is derived from the GPLv3 CVE2CAPEC project. MITRE ATT&CK® is a registered trademark of The MITRE Corporation; ATT&CK content is used in accordance with the MITRE ATT&CK terms of use.

References

Citation

@misc{bonhomme2026mappingcvesmitreattck,
      title={Mapping CVEs to MITRE ATT&CK Techniques: A Curated Gold-Set Classifier and the Limits of LLM-Assisted Label Expansion},
      author={Cédric Bonhomme and Alexandre Dulaunoy},
      year={2026},
      eprint={2607.25572},
      archivePrefix={arXiv},
      primaryClass={cs.CR},
      url={https://arxiv.org/abs/2607.25572},
}

Acknowledgements

Developed at CIRCL in the context of the AIPITCH project, co-funded by the European Union.

Contributors