
A comprehensive benchmarking framework for evaluating material generation models across multiple metrics including validity, distribution, diversity, novelty, uniqueness, and stability. [NeurIPS AI4Mat 2025 Spotlight]
⚠️ Known issue and ongoing recompute: Previous configs used to launch the benchmark used the MACE model with no checkpoint specified (
mace_mp()), which silently trackedmace-torch's changing default rather than the checkpoint that producedmace_mp_energyinLeMat-Bulk-MLIP-Hull(medium-0b3). Everymace_mp-hull-based stability/SUN/MSUN result computed before this fix mixed two energy scales. The checkpoints are now pinned explicitly.Affected historical results are being recomputed, and the public leaderboard is being updated with corrected numbers. If you're comparing against previously published SUN/MSUN/stability figures, treat them as provisional until this recompute lands — re-run with the current
mainif you need current results now.
# Clone the repository
git clone https://github.com/LeMaterial/lemat-genbench.git
cd lemat-genbench
# Install dependencies
uv sync
# Activate the virtual environment (macOS/Linux)
source .venv/bin/activate
# Set up UMA access (required for stability and distribution benchmarks)
huggingface-cli login
# Run a quick benchmark
uv run scripts/run_benchmarks.py --cifs cif_folder --config comprehensive_multi_mlip_hull --name quick_test
Clone the repository:
git clone https://github.com/LeMaterial/lemat-genbench.git
cd lemat-genbench
Install dependencies:
uv sync
Activate the virtual environment:
# On macOS/Linux:
source .venv/bin/activate
# On Windows:
.venv\Scripts\activate
Set up UMA model access (required for stability and distribution benchmarks):
# Request access to UMA model on HuggingFace
# Visit: https://huggingface.co/facebook/UMA
# Click "Request access" and wait for approval
# Login to HuggingFace CLI
huggingface-cli login
| Family | Description | Cost |
|---|---|---|
validity | Structure validation (charge, distance, plausibility) | Low |
distribution | Distribution similarity (JSD, MMD, Fréchet distance) | Medium |
diversity | Structural diversity (element, space group, site number) | Low |
novelty | Novelty vs. LeMat-Bulk reference dataset | Medium |
uniqueness | Internal uniqueness within generated set | Low |
stability | Thermodynamic stability (formation energy, e_above_hull) | High |
hhi | Supply risk assessment (production/reserve concentration) | Low |
sun | Composite metric (Stability + Uniqueness + Novelty) | High |
⚠️ Note: MMD calculations use a 15K sample from LeMat-Bulk dataset due to computational complexity.
⚠️ Note: Energy above hull calculations may fail for charged species (e.g., Cs+, Br-) as phase diagrams expect neutral compounds.
# Run all benchmark families on CIF files in a directory
uv run scripts/run_benchmarks.py --cifs /path/to/cif/directory --config comprehensive_multi_mlip_hull --name my_benchmark
# Run specific benchmark families
uv run scripts/run_benchmarks.py --cifs structures.txt --config comprehensive_multi_mlip_hull --families validity novelty --name custom_run
# Load structures from CSV file
uv run scripts/run_benchmarks.py --csv my_structures.csv --config comprehensive_multi_mlip_hull --name csv_benchmark
💡 Tip: Use configuration
comprehensive_multi_mlip_hullfor results comparable to the leaderboard.
| Option | Description |
|---|---|
--cifs | Path to directory or file list containing CIF files |
--csv | Path to CSV file containing structures |
--config | Configuration name (default: comprehensive) |
--name | Name for this benchmark run (required) |
--families | Specific benchmark families to run (optional) |
--fingerprint-method | Method: bawl, short-bawl, structure-matcher, pdd |
uv run scripts/run_benchmarks.py --cifs /path/to/cif/directory --config comprehensive_multi_mlip_hull --name my_run
Create a text file with CIF paths:
# my_structures.txt
path/to/structure1.cif
path/to/structure2.cif
path/to/structure3.cif
Then run:
uv run scripts/run_benchmarks.py --cifs my_structures.txt --config comprehensive_multi_mlip_hull --name my_run
uv run scripts/run_benchmarks.py --csv my_structures.csv --config comprehensive_multi_mlip_hull --name my_csv_run
CSV Format Requirements:
structure, LeMatStructs, or cif_stringExample CSV format:
material_id,structure,other_metadata
0,"{""@module"": ""pymatgen.core.structure"", ""@class"": ""Structure"", ""lattice"": {...}, ""sites"": [...]}",metadata1
1,"{""@module"": ""pymatgen.core.structure"", ""@class"": ""Structure"", ""lattice"": {...}, ""sites"": [...]}",metadata2
Note: You can only use one input method at a time (
--cifsOR--csv, not both).
| Method | Description | Speed | Memory |
|---|---|---|---|
bawl | Full BAWL fingerprinting | Fast | Low |
short-bawl | Shortened BAWL fingerprinting (default) | Fast | Low |
structure-matcher | PyMatGen StructureMatcher comparison | Slow | High |
pdd | Packing density descriptor | Medium | Medium |
# Use structure-matcher for better accuracy (slower)
uv run scripts/run_benchmarks.py \
--cifs submissions/test \
--config comprehensive_structure_matcher \
--name test_run \
--fingerprint-method structure-matcher
# Run only validity checks
uv run scripts/run_benchmarks.py --cifs structures/ --config validity --families validity --name validity_only
# Run only stability analysis
uv run scripts/run_benchmarks.py --cifs structures/ --config stability --families stability --name stability_only
# Run validity and novelty (low + medium cost)
uv run scripts/run_benchmarks.py --cifs structures/ --config comprehensive_multi_mlip_hull --families validity novelty --name validity_novelty
# Run diversity, uniqueness, and HHI (all low cost)
uv run scripts/run_benchmarks.py --cifs structures/ --config comprehensive_multi_mlip_hull --families diversity uniqueness hhi --name diversity_analysis
# Run distribution and stability (medium + high cost)
uv run scripts/run_benchmarks.py --cifs structures/ --config comprehensive_multi_mlip_hull --families distribution stability --name distribution_stability
uv run scripts/run_benchmarks.py --cifs structures/ --config comprehensive_multi_mlip_hull --name full_analysis
Results are saved to results/ directory:
{run_name}_{config_name}_{timestamp}.json
{
"run_info": {
"run_name": "my_benchmark",
"config_name": "comprehensive",
"timestamp": "20241204_143022",
"n_structures": 100,
"benchmark_families": ["validity", "distribution", "diversity", "..."]
},
"results": {
"validity": { "..." },
"distribution": { "..." },
"diversity": { "..." }
}
}
# Process a single file
uv run scripts/extract_benchmark_metrics.py results/my_run_comprehensive.json
# Process all JSON files in a directory
uv run scripts/extract_benchmark_metrics.py results_new/ --directory
# Specify custom output directory
uv run scripts/extract_benchmark_metrics.py results/my_run.json --output-dir custom_output/
# Process directory with custom pattern
uv run scripts/extract_benchmark_metrics.py final_results/ --directory --pattern "*comprehensive*.json"
uv run scripts/run_benchmarks.py --cifs notebooks --config validity --name quick_validity
uv run scripts/run_benchmarks.py --cifs my_structures/ --config stability --name stability_analysis
uv run scripts/run_benchmarks.py --cifs structures.txt --config comprehensive_multi_mlip_hull --families validity novelty uniqueness --name custom_analysis
# Quick validation of CSV structures
uv run scripts/run_benchmarks.py --csv my_structures.csv --config validity --name csv_validity
# Full analysis of CSV structures
uv run scripts/run_benchmarks.py --csv generated_structures.csv --config comprehensive_multi_mlip_hull --name csv_full_analysis
# Use SSH-optimized script for large datasets
uv run scripts/run_benchmarks_ssh.py --cifs large_dataset/ --config comprehensive_multi_mlip_hull --name large_run
--families to run only needed benchmarksrun_benchmarks_ssh.py for high-core environments1. UMA Access Denied:
# Ensure you're logged in
huggingface-cli login
# Check access status
huggingface-cli whoami
2. Memory Issues:
# Run fewer families at once
uv run scripts/run_benchmarks.py --cifs structures/ --config validity --families validity --name memory_test
3. Timeout Errors:
4. Private Dataset Access Error:
# Error: 'Entalpic/LeMaterial-Above-Hull-dataset' doesn't exist on the Hub
# Solution: Download datasets locally (one-time setup)
uv run scripts/download_above_hull_datasets.py
5. Structure-Matcher Performance:
src/config/UMA Model: HuggingFace
Wood, Brandon M., et al. "UMA: A Family of Universal Models for Atoms." arXiv preprint arXiv:2506.23971 (2025).
ORB Models: GitHub
Rhodes, Benjamin, et al. "Orb-v3: atomistic simulation at scale." arXiv preprint arXiv:2504.06231 (2025).
MACE Models: GitHub
Batatia, Ilyes, et al. "MACE: Higher order equivariant message passing neural networks for fast and accurate force fields." Advances in Neural Information Processing Systems 35 (2022): 11423-11436.
Distribution Metrics:
Diversity Metrics:
Supply Risk Metrics:
This project is licensed under the Apache License - see the LICENSE file for details.
If you use LeMat-GenBench in your research, please cite:
@article{betala2025lemat,
title = {LeMat-GenBench: A Unified Evaluation Framework for Crystal Generative Models},
author = {Betala, Siddharth and Gleason, Samuel P. and Ramlaoui, Ali and Xu, Andy and
Channing, Georgia and Levy, Daniel and Fourrier, Cl{\'e}mentine and Kazeev, Nikita and
Joshi, Chaitanya K. and Kaba, S{\'e}kou-Oumar and Therrien, F{\'e}lix and
Hernandez-Garcia, Alex and Mercado, Roc{\'\i}o and Krishnan, N. M. Anoop and
Duval, Alexandre},
journal = {arXiv preprint arXiv:2512.04562},
year = {2025}
}
Python
86.9%
Jupyter Notebook
12.9%

A comprehensive benchmarking framework for evaluating material generation models across multiple metrics including validity, distribution, diversity, novelty, uniqueness, and stability. [NeurIPS AI4Mat 2025 Spotlight]
⚠️ Known issue and ongoing recompute: Previous configs used to launch the benchmark used the MACE model with no checkpoint specified (
mace_mp()), which silently trackedmace-torch's changing default rather than the checkpoint that producedmace_mp_energyinLeMat-Bulk-MLIP-Hull(medium-0b3). Everymace_mp-hull-based stability/SUN/MSUN result computed before this fix mixed two energy scales. The checkpoints are now pinned explicitly.Affected historical results are being recomputed, and the public leaderboard is being updated with corrected numbers. If you're comparing against previously published SUN/MSUN/stability figures, treat them as provisional until this recompute lands — re-run with the current
mainif you need current results now.
# Clone the repository
git clone https://github.com/LeMaterial/lemat-genbench.git
cd lemat-genbench
# Install dependencies
uv sync
# Activate the virtual environment (macOS/Linux)
source .venv/bin/activate
# Set up UMA access (required for stability and distribution benchmarks)
huggingface-cli login
# Run a quick benchmark
uv run scripts/run_benchmarks.py --cifs cif_folder --config comprehensive_multi_mlip_hull --name quick_test
Clone the repository:
git clone https://github.com/LeMaterial/lemat-genbench.git
cd lemat-genbench
Install dependencies:
uv sync
Activate the virtual environment:
# On macOS/Linux:
source .venv/bin/activate
# On Windows:
.venv\Scripts\activate
Set up UMA model access (required for stability and distribution benchmarks):
# Request access to UMA model on HuggingFace
# Visit: https://huggingface.co/facebook/UMA
# Click "Request access" and wait for approval
# Login to HuggingFace CLI
huggingface-cli login
| Family | Description | Cost |
|---|---|---|
validity | Structure validation (charge, distance, plausibility) | Low |
distribution | Distribution similarity (JSD, MMD, Fréchet distance) | Medium |
diversity | Structural diversity (element, space group, site number) | Low |
novelty | Novelty vs. LeMat-Bulk reference dataset | Medium |
uniqueness | Internal uniqueness within generated set | Low |
stability | Thermodynamic stability (formation energy, e_above_hull) | High |
hhi | Supply risk assessment (production/reserve concentration) | Low |
sun | Composite metric (Stability + Uniqueness + Novelty) | High |
⚠️ Note: MMD calculations use a 15K sample from LeMat-Bulk dataset due to computational complexity.
⚠️ Note: Energy above hull calculations may fail for charged species (e.g., Cs+, Br-) as phase diagrams expect neutral compounds.
# Run all benchmark families on CIF files in a directory
uv run scripts/run_benchmarks.py --cifs /path/to/cif/directory --config comprehensive_multi_mlip_hull --name my_benchmark
# Run specific benchmark families
uv run scripts/run_benchmarks.py --cifs structures.txt --config comprehensive_multi_mlip_hull --families validity novelty --name custom_run
# Load structures from CSV file
uv run scripts/run_benchmarks.py --csv my_structures.csv --config comprehensive_multi_mlip_hull --name csv_benchmark
💡 Tip: Use configuration
comprehensive_multi_mlip_hullfor results comparable to the leaderboard.
| Option | Description |
|---|---|
--cifs | Path to directory or file list containing CIF files |
--csv | Path to CSV file containing structures |
--config | Configuration name (default: comprehensive) |
--name | Name for this benchmark run (required) |
--families | Specific benchmark families to run (optional) |
--fingerprint-method | Method: bawl, short-bawl, structure-matcher, pdd |
uv run scripts/run_benchmarks.py --cifs /path/to/cif/directory --config comprehensive_multi_mlip_hull --name my_run
Create a text file with CIF paths:
# my_structures.txt
path/to/structure1.cif
path/to/structure2.cif
path/to/structure3.cif
Then run:
uv run scripts/run_benchmarks.py --cifs my_structures.txt --config comprehensive_multi_mlip_hull --name my_run
uv run scripts/run_benchmarks.py --csv my_structures.csv --config comprehensive_multi_mlip_hull --name my_csv_run
CSV Format Requirements:
structure, LeMatStructs, or cif_stringExample CSV format:
material_id,structure,other_metadata
0,"{""@module"": ""pymatgen.core.structure"", ""@class"": ""Structure"", ""lattice"": {...}, ""sites"": [...]}",metadata1
1,"{""@module"": ""pymatgen.core.structure"", ""@class"": ""Structure"", ""lattice"": {...}, ""sites"": [...]}",metadata2
Note: You can only use one input method at a time (
--cifsOR--csv, not both).
| Method | Description | Speed | Memory |
|---|---|---|---|
bawl | Full BAWL fingerprinting | Fast | Low |
short-bawl | Shortened BAWL fingerprinting (default) | Fast | Low |
structure-matcher | PyMatGen StructureMatcher comparison | Slow | High |
pdd | Packing density descriptor | Medium | Medium |
# Use structure-matcher for better accuracy (slower)
uv run scripts/run_benchmarks.py \
--cifs submissions/test \
--config comprehensive_structure_matcher \
--name test_run \
--fingerprint-method structure-matcher
# Run only validity checks
uv run scripts/run_benchmarks.py --cifs structures/ --config validity --families validity --name validity_only
# Run only stability analysis
uv run scripts/run_benchmarks.py --cifs structures/ --config stability --families stability --name stability_only
# Run validity and novelty (low + medium cost)
uv run scripts/run_benchmarks.py --cifs structures/ --config comprehensive_multi_mlip_hull --families validity novelty --name validity_novelty
# Run diversity, uniqueness, and HHI (all low cost)
uv run scripts/run_benchmarks.py --cifs structures/ --config comprehensive_multi_mlip_hull --families diversity uniqueness hhi --name diversity_analysis
# Run distribution and stability (medium + high cost)
uv run scripts/run_benchmarks.py --cifs structures/ --config comprehensive_multi_mlip_hull --families distribution stability --name distribution_stability
uv run scripts/run_benchmarks.py --cifs structures/ --config comprehensive_multi_mlip_hull --name full_analysis
Results are saved to results/ directory:
{run_name}_{config_name}_{timestamp}.json
{
"run_info": {
"run_name": "my_benchmark",
"config_name": "comprehensive",
"timestamp": "20241204_143022",
"n_structures": 100,
"benchmark_families": ["validity", "distribution", "diversity", "..."]
},
"results": {
"validity": { "..." },
"distribution": { "..." },
"diversity": { "..." }
}
}
# Process a single file
uv run scripts/extract_benchmark_metrics.py results/my_run_comprehensive.json
# Process all JSON files in a directory
uv run scripts/extract_benchmark_metrics.py results_new/ --directory
# Specify custom output directory
uv run scripts/extract_benchmark_metrics.py results/my_run.json --output-dir custom_output/
# Process directory with custom pattern
uv run scripts/extract_benchmark_metrics.py final_results/ --directory --pattern "*comprehensive*.json"
uv run scripts/run_benchmarks.py --cifs notebooks --config validity --name quick_validity
uv run scripts/run_benchmarks.py --cifs my_structures/ --config stability --name stability_analysis
uv run scripts/run_benchmarks.py --cifs structures.txt --config comprehensive_multi_mlip_hull --families validity novelty uniqueness --name custom_analysis
# Quick validation of CSV structures
uv run scripts/run_benchmarks.py --csv my_structures.csv --config validity --name csv_validity
# Full analysis of CSV structures
uv run scripts/run_benchmarks.py --csv generated_structures.csv --config comprehensive_multi_mlip_hull --name csv_full_analysis
# Use SSH-optimized script for large datasets
uv run scripts/run_benchmarks_ssh.py --cifs large_dataset/ --config comprehensive_multi_mlip_hull --name large_run
--families to run only needed benchmarksrun_benchmarks_ssh.py for high-core environments1. UMA Access Denied:
# Ensure you're logged in
huggingface-cli login
# Check access status
huggingface-cli whoami
2. Memory Issues:
# Run fewer families at once
uv run scripts/run_benchmarks.py --cifs structures/ --config validity --families validity --name memory_test
3. Timeout Errors:
4. Private Dataset Access Error:
# Error: 'Entalpic/LeMaterial-Above-Hull-dataset' doesn't exist on the Hub
# Solution: Download datasets locally (one-time setup)
uv run scripts/download_above_hull_datasets.py
5. Structure-Matcher Performance:
src/config/UMA Model: HuggingFace
Wood, Brandon M., et al. "UMA: A Family of Universal Models for Atoms." arXiv preprint arXiv:2506.23971 (2025).
ORB Models: GitHub
Rhodes, Benjamin, et al. "Orb-v3: atomistic simulation at scale." arXiv preprint arXiv:2504.06231 (2025).
MACE Models: GitHub
Batatia, Ilyes, et al. "MACE: Higher order equivariant message passing neural networks for fast and accurate force fields." Advances in Neural Information Processing Systems 35 (2022): 11423-11436.
Distribution Metrics:
Diversity Metrics:
Supply Risk Metrics:
This project is licensed under the Apache License - see the LICENSE file for details.
If you use LeMat-GenBench in your research, please cite:
@article{betala2025lemat,
title = {LeMat-GenBench: A Unified Evaluation Framework for Crystal Generative Models},
author = {Betala, Siddharth and Gleason, Samuel P. and Ramlaoui, Ali and Xu, Andy and
Channing, Georgia and Levy, Daniel and Fourrier, Cl{\'e}mentine and Kazeev, Nikita and
Joshi, Chaitanya K. and Kaba, S{\'e}kou-Oumar and Therrien, F{\'e}lix and
Hernandez-Garcia, Alex and Mercado, Roc{\'\i}o and Krishnan, N. M. Anoop and
Duval, Alexandre},
journal = {arXiv preprint arXiv:2512.04562},
year = {2025}
}
Python
86.9%
Jupyter Notebook
12.9%