daviddao/awesome-data-valuation

💱 A curated list of data valuation (DV) to design your next data marketplace

144

53 commits

updated Aug 14, 2026

See the code

README

Awesome Data Valuation

data market problem

💱 A curated list of data valuation (DV) to design your next data marketplace. DV aims to understand the value of a data point for a given machine learning task and is an essential primitive in the design of data marketplaces and explainable AI.

Legend

💻 Code available

🎥 Talk / Slides

Contents

What is your data worth?

Shapley Value & Cooperative Game Theory

Towards Efficient Data Valuation Based on the Shapley ValueRuoxi Jia & David Dao, Boxin Wang, Frances Ann Hubis, Nick Hynes, Nezihe Merve Gürel, Bo Li, Ce Zhang, Dawn Song, Costas J. Spanos2019
Summary Jia et al. (2019) contribute theoretical and practical results for efficient methods for approximating the Shapley value (SV). They show that methods with a sublinear amount of model evaluations are possible and further reductions can be made for sparse SVs. Lastly, they introduce two practical SV estimation methods for ML tasks, one for uniformly stable learning algorithms and one for smooth loss functions.
Bibtex
@inproceedings{jia2019towards,
title={Towards efficient data valuation based on the shapley value},
author={Jia, Ruoxi and Dao, David and Wang, Boxin and Hubis, Frances Ann and Hynes, Nick and G{"u}rel, Nezihe Merve and Li, Bo and Zhang, Ce and Song, Dawn and Spanos, Costas J},
booktitle={The 22nd International Conference on Artificial Intelligence and Statistics},
pages={1167--1176},
year={2019},
organization={PMLR}
}
💻
Data Shapley: Equitable Valuation of Data for Machine LearningAmirata Ghorbani, James Zou2019
Summary Ghorbani & Zou (2019) introduce (data) Shapley value to equitably measure the value of each training point to a supervised learners performance. They further outline several benefits of the Shapley value, e.g. being able to capture outliers or inform what new data to acquire, as well as develop Monte Carlo and gradient-based methods for its efficient estimation.
Bibtex
@inproceedings{ghorbani2019data,
title={Data shapley: Equitable valuation of data for machine learning},
author={Ghorbani, Amirata and Zou, James},
booktitle={International Conference on Machine Learning},
pages={2242--2251},
year={2019},
organization={PMLR}
}
💻
A Distributional Framework for Data ValuationAmirata Ghorbani, Michael P. Kim, James Zou2020
Summary Ghorbani et al. (2020) formulate the Shapley value as a distributional quantity in the context of an underlying data distribution instead of a fixed dataset. They further introduce a novel sampling-based algorithm for the distributional Shapley value with strong approximation guarantees.
Bibtex
@inproceedings{ghorbani2020distributional,
title={A Distributional Framework for Data Valuation},
author={Ghorbani, Amirata, P. Kim, Michael and Zou, James},
booktitle={International Conference on Machine Learning},
year={2020}
}
💻
Asymmetric Shapley values: incorporating causal knowledge into model-agnostic explainabilityChristopher Frye, Colin Rowat, Ilya Feige2020
Summary Frye et al. (2020) incorporate causality into the Shapley value framework. Importantly, their framework can handle any amount of causal knowledge and does not require the complete causal graph underlying the data.
Bibtex
@article{frye2020asymmetric,
title={Asymmetric Shapley values: incorporating causal knowledge into model-agnostic explainability},
author={Frye, Christopher and Rowat, Colin and Feige, Ilya},
journal={Advances in Neural Information Processing Systems},
volume={33},
year={2020}
}
🎥
Collaborative Machine Learning with Incentive-Aware Model RewardsRachael Hwee Ling Sim, Yehong Zhang, Mun Choon Chan, Bryan Kian Hsiang Low2020
Summary Sim et al. (2020) introduce a data valuation method with separate ML models as rewards based on the Shapley value and information gain on model parameters given its data. They further define several conditions for incentives such as Shapley fairness, stability, individual rationality, and group welfare, that are suitable for the freely replicable nature of their model reward scheme.
Bibtex
@inproceedings{sim2020collaborative,
title={Collaborative machine learning with incentive-aware model rewards},
author={Sim, Rachael Hwee Ling and Zhang, Yehong and Chan, Mun Choon and Low, Bryan Kian Hsiang},
booktitle={International Conference on Machine Learning},
pages={8927--8936},
year={2020},
organization={PMLR}
}
Validation free and replication robust volume-based data valuationXinyi Xu, Zhaoxuan Wu, Chuan Sheng Foo, Bryan Kian Hsiang Low2021
Summary Xu et al. (2021) propose using data diversity via robust volume for measuring the value of data. This removes the need for a validation set and allows for guarantees on replication robustness but suffers from the curse of dimensionality and may ignore useful information in the validation set.
Bibtex
@article{xu2021validation,
title={Validation free and replication robust volume-based data valuation},
author={Xu, Xinyi and Wu, Zhaoxuan and Foo, Chuan Sheng and Low, Bryan Kian Hsiang},
journal={Advances in Neural Information Processing Systems},
volume={34},
year={2021}
}
💻
Beta Shapley: a Unified and Noise-reduced Data Valuation Framework for Machine LearningYongchan Kwon, James Zou2021
Summary Kwon & Zou (2022) introduce Beta Shapley, a generalization of Data Shapley by relaxing the efficiency axiom.
Bibtex
@article{kwon2021beta,
title={Beta Shapley: a Unified and Noise-reduced Data Valuation Framework for Machine Learning},
author={Kwon, Yongchan and Zou, James},
journal={arXiv preprint arXiv:2110.14049},
year={2021}
}
Gradient-Driven Rewards to Guarantee Fairness in Collaborative Machine LearningXinyi Xu, Lingjuan Lyu, Xingjun Ma, Chenglin Miao, Chuan Sheng Foo, Bryan Kian Hsiang Low2021
Summary Xu et al. (2021) propose cosine gradient Shapley value to fairly evaluate the expected contribution of each agent's update in the federated learning setting removing the need for an auxiliary validation dataset. They further introduce a novel training-time gradient reward mechanism with a fairness guarantee.
Bibtex
@article{xu2021gradient,
title={Gradient driven rewards to guarantee fairness in collaborative machine learning},
author={Xu, Xinyi and Lyu, Lingjuan and Ma, Xingjun and Miao, Chenglin and Foo, Chuan Sheng and Low, Bryan Kian Hsiang},
journal={Advances in Neural Information Processing Systems},
volume={34},
pages={16104--16117},
year={2021}
}
Improving Cooperative Game Theory-based Data Valuation via Data Utility LearningTianhao Wang, Yu Yang, Ruoxi Jia2022
Summary Wang et al. (2022) propose a general framework to improve effectiveness of sampling-based Shapley value (SV) or Least core (LC) estimation heuristics. They propose learning to predict the performance of a learning algorithm (denoted data utility learning) and using this predictor to estimate learning performance without retraining for cheaper SV and LC estimation.
Bibtex
@article{wang2021improving,
title={Improving cooperative game theory-based data valuation via data utility learning},
author={Wang, Tianhao and Yang, Yu and Jia, Ruoxi},
journal={arXiv preprint arXiv:2107.06336},
year={2021}
}
Data Banzhaf: A Robust Data Valuation Framework for Machine LearningJiachen T. Wang, Ruoxi Jia2023
Summary Wang et al. (2023) propose using the Banzhaf value for data valuation, providing better robustness against noisy performance scores and an efficient estimate using Maximum Sample Reuse (MSR) principle
Bibtex
@InProceedings{pmlr-v206-wang23e, title={Data Banzhaf: A Robust Data Valuation Framework for Machine Learning},
author={Wang, Jiachen T. and Jia, Ruoxi},
booktitle={Proceedings of The 26th International Conference on Artificial Intelligence and Statistics},
pages={6388--6421},
year={2023},
editor={Ruiz, Francisco and Dy, Jennifer and van de Meent, Jan-Willem},
volume={206},
series={Proceedings of Machine Learning Research},
month={25--27 Apr},
publisher={PMLR},
pdf={https://proceedings.mlr.press/v206/wang23e/wang23e.pdf},
url={https://proceedings.mlr.press/v206/wang23e.html}
}
💻
A Multilinear Sampling Algorithm to Estimate Shapley ValuesRamin Okhrati, Aldo Lipani2021
Summary Okhrati and Lipani (2021) propose a new sampling method for Shapley values based on a multilinear extension technique as applied in game theory. It provides more accurate estimations of the Shapley values by reducing the variance of the sampling statistics.
Bibtex
@INPROCEEDINGS{9412511,
title={A Multilinear Sampling Algorithm to Estimate Shapley Values},
author={Okhrati, Ramin and Lipani, Aldo},
booktitle={2020 25th International Conference on Pattern Recognition (ICPR)},
year={2021}
}
💻
If You Like Shapley Then You’ll Love the CoreYan, T., and Procaccia, A. D.2021
Summary Yan and Procaccia (2021) propose an alternative method for credit assignment in data valuation. They use the least core, which can be computed efficiently.
Bibtex
@article{Yan_Procaccia_2021,
title={If You Like Shapley Then You’ll Love the Core},
author={Yan, Tom and Procaccia, Ariel D.},
journal={Proceedings of the AAAI Conference on Artificial Intelligence},
year={2021}
}
CS-Shapley: Class-wise Shapley Values for Data Valuation in ClassificationSchoch, Stephanie, Haifeng Xu, and Yangfeng Ji2022
Summary Schoch et al. (2022) propose a new Shapley value that discriminates between training instances' in-class and out-of-class contributions.
Bibtex
@inproceedings{schoch2022csshapley,
title={{CS}-Shapley: Class-wise Shapley Values for Data Valuation in Classification},
author={Stephanie Schoch and Haifeng Xu and Yangfeng Ji},
booktitle={Advances in Neural Information Processing Systems},
year={2022}
}
💻
Precedence-Constrained Winter Value for Effective Graph Data ValuationHongliang Chi, Wei Jin, Charu Aggarwal, and Yao Ma2022
Summary Introduces a new data valuation concept similar to Data Shapley but based on the Winter value.
Bibtex
@misc{chi2024wintervalue,
title={Precedence-Constrained Winter Value for Effective Graph Data Valuation},
author={Hongliang Chi and Wei Jin and Charu Aggarwal and Yao Ma},
year={2024},
eprint={2402.01943},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2402.01943},
}

Efficient algorithms

Efficient Task-Specific Data Valuation for Nearest Neighbor AlgorithmsRuoxi Jia, David Dao, Boxin Wang, Frances Ann Hubis, Nezihe Merve Gurel, Bo Li, Ce Zhang, Costas J. Spanos, Dawn Song2019
Summary Jia et al. (2019) present algorithms to compute the Shapley value exactly in quasi-linear time and approximations in sublinear time for k-nearest-neighbor models. They empirically evaluate their algorithms at scale and extend them to several other settings.
Bibtex
@article{jia12efficient,
title={Efficient Task-Specific Data Valuation for Nearest Neighbor Algorithms},
author={Jia, Ruoxi and Dao, David and Wang, Boxin and Hubis, Frances Ann and Gurel, Nezihe Merve and Zhang, Bo Li4 Ce and Song, Costas Spanos1 Dawn},
journal={Proceedings of the VLDB Endowment},
volume={12},
number={11}
}
💻
Efficient computation and analysis of distributional Shapley valuesYongchan Kwon, Manuel A. Rivas, James Zou2021
Summary Kwon et al. (2021) develop tractable analytic expressions for the distributional data Shapley value for linear regression, binary classification, and non-parametric density estimation as well as new efficient methods for its estimation.
Bibtex
@inproceedings{kwon2021efficient,
title={Efficient computation and analysis of distributional Shapley values},
author={Kwon, Yongchan and Rivas, Manuel A and Zou, James},
booktitle={International Conference on Artificial Intelligence and Statistics},
pages={793--801},
year={2021},
organization={PMLR}
}
💻
DUPRE: Data Utility Prediction for Efficient Data ValuationKieu Thao Nguyen Pham, Rachael Hwee Ling Sim, Quoc Phong Nguyen, See Kiong Ng, Bryan Kian Hsiang Low2025
Summary Pham et al. (2025) speed up cooperative game-theoretic valuation from the other direction: instead of reducing the number of subsets to evaluate, DUPRE predicts the utility of unseen data subsets with a learned model, skipping costly retraining for each coalition.
Bibtex
@article{pham2025dupre,
title={DUPRE: Data Utility Prediction for Efficient Data Valuation},
author={Pham, Kieu Thao Nguyen and Sim, Rachael Hwee Ling and Nguyen, Quoc Phong and Ng, See Kiong and Low, Bryan Kian Hsiang},
journal={arXiv preprint arXiv:2502.16152},
year={2025}
}

Benchmarks, Criticism & Relaxations

Scalability vs. Utility: Do We Have to Sacrifice One for the Other in Data Importance Quantification?Ruoxi Jia, Fan Wu, Xuehui Sun, Jiacen Xu, David Dao, Bhavya Kailkhura, Ce Zhang, Bo Li, Dawn Song2021
Summary Jia et al. (2021) perform a theoretical analysis on the differences between leave-one-out-based and Shapley value-based methods as well as an empirical study across several ML tasks investigating the two aforementioned methods as well as exact Shapley value-based methods and Shapley over KNN Surrogates.
Bibtex
@misc{jia2021scalability,
title={Scalability vs. Utility: Do We Have to Sacrifice One for the Other in Data Importance Quantification?},
author={Ruoxi Jia and Fan Wu and Xuehui Sun and Jiacen Xu and David Dao and Bhavya Kailkhura and Ce Zhang and Bo Li and Dawn Song},
year={2021},
eprint={1911.07128},
archivePrefix={arXiv},
primaryClass={cs.LG}
}
💻
Shapley values for feature selection: The good, the bad, and the axiomsDaniel Fryer, Inga Strümke, Hien Nguyen2021
Summary Fryer et al. (2021) calls into question the appropriateness of using the Shapley value for feature selection and advise caution against the magical thinking that presenting its abstract general axioms as "favourable and fair" may introduce. They further point out that the four axioms of "efficiency", "null player", "symmetry", and "additivity" do not guarantee that the Shapley value is suited to feature selection and may sometimes even imply the opposite.
Bibtex
@misc{fryer2021shapley,
title={Shapley values for feature selection: The good, the bad, and the axioms},
author={Daniel Fryer and Inga Strümke and Hien Nguyen},
year={2021},
eprint={2102.10936},
archivePrefix={arXiv},
primaryClass={cs.LG}
}
Rethinking Data Shapley for Data Selection Tasks: Misleads and MeritsJiachen T. Wang, Tianji Yang, James Zou, Yongchan Kwon, Ruoxi Jia2024
Summary Wang et al. (ICML 2024 oral) show via a hypothesis-testing framework that Data Shapley's data-selection performance can be no better than random without constraints on the utility function, and identify the class of utility functions (monotonically transformed modular functions) under which it selects optimally.
Bibtex
@inproceedings{pmlr-v235-wang24cg,
title={Rethinking Data Shapley for Data Selection Tasks: Misleads and Merits},
author={Wang, Jiachen T. and Yang, Tianji and Zou, James and Kwon, Yongchan and Jia, Ruoxi},
booktitle={Proceedings of the 41st International Conference on Machine Learning},
pages={52033--52063},
year={2024},
publisher={PMLR}
}
Semivalue-based data valuation is arbitrary and gameableHannah Diehl, Ashia C. Wilson2025
Summary Diehl & Wilson (2025) prove that semivalue-based valuations (incl. Data Shapley, Beta Shapley, Banzhaf) are underdetermined: defensible changes to the utility specification can flip rankings, and strategic actors can exploit this latitude to game outcomes. A foundational challenge to using axiomatic valuations for high-stakes decisions.
Bibtex
@article{diehl2025semivalue,
title={Semivalue-based data valuation is arbitrary and gameable},
author={Diehl, Hannah and Wilson, Ashia C.},
journal={arXiv preprint arXiv:2506.12619},
year={2025}
}
Do Data Valuations Make Good Data Prices?Dongyang Fan, Tyler J. Rotello, Sai Praneeth Karimireddy2025
Summary Fan et al. (2025) revisit data valuation from a market-design perspective and show that popular valuations (LOO, Data Shapley) make poor payments: they fail truthfulness and cost-coverage for heterogeneous data owners. Attribution and pricing are different problems — a mechanism layer is needed on top of valuation.
Bibtex
@article{fan2025data,
title={Do Data Valuations Make Good Data Prices?},
author={Fan, Dongyang and Rotello, Tyler J. and Karimireddy, Sai Praneeth},
journal={arXiv preprint arXiv:2504.05563},
year={2025}
}

Influence functions & LOO

Understanding Black-box Predictions via Influence FunctionsPang Wei Koh, Percy Liang2017
Summary Koh & Liang (2017) introduce the use of influence functions, a technique borrowed from robust statistics, to identify training points most responsible for a model's given prediction without needing to retrain. They further develop a simple and efficient implementation of influence functions that scales to large ML settings.
Bibtex
@inproceedings{koh2017understanding,
title={Understanding black-box predictions via influence functions},
author={Koh, Pang Wei and Liang, Percy},
booktitle={International Conference on Machine Learning},
pages={1885--1894},
year={2017},
organization={PMLR}
}
💻🎥
On the accuracy of influence functions for measuring group effectsPang Wei Koh*, Kai-Siang Ang*, Hubert H. K. Teo*, and Percy Liang2019
Summary Koh et al. (2019) study influence functions to measure effects of large groups of training points instead of individual points. They empirically find a correlation and often underestimation between predicted and actual effects and theoretically show that this need not hold in general, realistic settings.
Bibtex
@article{koh2019accuracy,
title={On the accuracy of influence functions for measuring group effects},
author={Koh, Pang Wei and Ang, Kai-Siang and Teo, Hubert HK and Liang, Percy},
journal={arXiv preprint arXiv:1905.13289},
year={2019}
}
💻🎥
Scaling Up Influence FunctionsSchioppa, Andrea, Polina Zablotskaia, David Vilar, and Artem Sokolov2022
Summary Schioppa et al. (2022) propose a new method to scale the computation of influence functions for large neural networks using the Arnoldi iteration. With this, they achieve successful implementation of influence functions on full-size Transformer models with hundreds of millions of parameters.
Bibtex
@inproceedings{schioppa2022scaling,
title={Scaling Up Influence Functions},
author={Schioppa, Andrea and Zablotskaia, Polina and Vilar, David and Sokolov, Artem},
booktitle={Proceedings of the AAAI Conference on Artificial Intelligence},
year={2022}
}
💻
Studying large language model generalization with influence functionsGrosse, Roger and Bae, Juhan and Anil, Cem and Elhage, Nelson and Tamkin, Alex and Tajdini, Amirhossein and Steiner, Benoit and Li, Dustin and Durmus, Esin and Perez, Ethan and others2023
Summary Grosse et al. (2023) use a method known as EK-FAC to approximate the Hessian of the loss of large language models. They apply this technique to study influence functions on large language models, up to 50 billion parameters.
Bibtex
@article{grosse2023studying,
title={Studying large language model generalization with influence functions},
author={Grosse, Roger and Bae, Juhan and Anil, Cem and Elhage, Nelson and Tamkin, Alex and Tajdini, Amirhossein and Steiner, Benoit and Li, Dustin and Durmus, Esin and Perez, Ethan and others},
journal={arXiv preprint arXiv:2308.03296},
year={2023}
}
Do Influence Functions Work on Large Language Models?Zhe Li, Wei Zhao, Yige Li, Jun Sun2024
Summary Li et al. (2024) systematically evaluate influence functions across multiple LLM tasks and find they consistently perform poorly, tracing the failures to approximation errors in inverse-Hessian estimation, uncertain convergence during fine-tuning, and the definition of influence itself. A caution for influence-based valuation at LLM scale.
Bibtex
@article{li2024influence,
title={Do Influence Functions Work on Large Language Models?},
author={Li, Zhe and Zhao, Wei and Li, Yige and Sun, Jun},
journal={arXiv preprint arXiv:2409.19998},
year={2024}
}

Reinforcement Learning

Data Valuation using Reinforcement LearningJinsung Yoon, Sercan Ö Arık, Tomas Pfister2020
Summary Yoon et al. (2020) propose using reinforcement learning for data valuation to learn data values jointly with the predictor model.
Bibtex
@inproceedings{49189,
title={Data Valuation using Reinforcement Learning},
author={Jinsung Yoon and Sercan Arik and Tomas Pfister},
year={2020}
}
💻🎥

Deep Neural Networks

DAVINZ: Data Valuation using Deep Neural Networks at InitializationZhaoxuan Wu, Yao Shu, Bryan Kian Hsiang Low2022
Summary Wu et al. (2022) introduce a validation-based and training-free method for efficient data valuation with large and complex deep neural networks (DNNs). They derive and exploit a domain-aware generalization bound for DNNs to characterize their performance without training and uses this bound as the scoring function while keeping conventional techniques such as Shapley values as the valuation function.
Bibtex
@inproceedings{wu2022davinz,
title={DAVINZ: Data Valuation using Deep Neural Networks at Initialization},
author={Wu, Zhaoxuan and Shu, Yao and Low, Bryan Kian Hsiang},
booktitle={International Conference on Machine Learning},
pages={24150--24176},
year={2022},
organization={PMLR}
}
🎥
LossVal: Efficient Data Valuation for Neural NetworksTim Wibiral, Mohamed Karim Belaid, Maximilian Rabus, and Ansgar Scherp2024
Summary Wibiral et al. (2024) introduce a efficient method for deep neural networks that learns the importance scores of the data valuation as weights of the loss function during the first training run.
Bibtex
@misc{wibiral2024lossvalefficientdatavaluation,
title={LossVal: Efficient Data Valuation for Neural Networks},
author={Tim Wibiral and Mohamed Karim Belaid and Maximilian Rabus and Ansgar Scherp},
year={2024},
eprint={2412.04158},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2412.04158},
}
💻

Out-of-Bag score

Data-OOB: Out-of-bag Estimate as a Simple and Efficient Data ValueYongchan Kwon, James Zou2023
Summary Kwon et al. (2023) propose using the out-of-bag estimate of a bagging estimator for computationally efficient data valuation.
Bibtex
@inproceedings{DBLP:conf/icml/Kwon023, 
author={Yongchan Kwon and James Zou},
editor={Andreas Krause and Emma Brunskill and Kyunghyun Cho and Barbara Engelhardt and Sivan Sabato and Jonathan Scarlett},
title={Data-OOB: Out-of-bag Estimate as a Simple and Efficient Data Value},
booktitle={International Conference on Machine Learning, {ICML} 2023, 23-29 July 2023, Honolulu, Hawaii, {USA}},
series={Proceedings of Machine Learning Research},
volume={202},
pages={18135--18152},
publisher={{PMLR}},
year={2023},
url={https://proceedings.mlr.press/v202/kwon23e.html},
timestamp={Mon, 28 Aug 2023 17:23:08 +0200},
biburl={https://dblp.org/rec/conf/icml/Kwon023.bib},
bibsource={dblp computer science bibliography, https://dblp.org}
}
💻🎥

Evolutionary Approaches

An evolutionary approach to data valuationNatalia Khuri, Sapan Bhandari, Esteban Murillo Burford, Nathan P. Whitener, and Konghao Zhao2022
Summary Khuri et al. (2022) propose using an evolutionary algorithm for data valuation.
Bibtex
@inproceedings{10.1145/3535508.3545522,
author = {Khuri, Natalia and Bhandari, Sapan and Burford, Esteban Murillo and Whitener, Nathan P. and Zhao, Konghao},
title = {An evolutionary approach to data valuation},
year = {2022},
isbn = {9781450393867},
publisher = {Association for Computing Machinery},
address = {New York, NY, USA},
url = {https://doi.org/10.1145/3535508.3545522},
doi = {10.1145/3535508.3545522}
}

Task Agnostic

Fundamentals of Task-Agnostic Data ValuationMohammad Mohammadi Amiri, Frederic Berdoz, Ramesh Raskar2023
SummaryThis paper addresses the challenge of valuing data without specific task assumptions, focusing on task-agnostic data valuation. It discusses valuing a data seller's dataset from a buyer's perspective without validation requirements. The approach involves estimating statistical differences through diversity and relevance measures without needing the raw data, and designing queries that maintain the seller's blindness to the buyer's raw data. The work is significant for practical scenarios where utility metrics like test accuracy on a validation set are not feasible.
Bibtex
@article{Amiri2023FundamentalsOT,
title={Fundamentals of Task-Agnostic Data Valuation},
author={Mohammad Mohammadi Amiri and Frederic Berdoz and Ramesh Raskar},
journal={Proceedings of the AAAI Conference on Artificial Intelligence},
volume={37},
pages={9226-9234},
year={2023},
doi={10.1609/aaai.v37i8.26106}
}
KAIROS: Scalable Model-Agnostic Data ValuationJiongli Zhu, Parjanya Prajakta Prashant, Alex Cloninger, Babak Salimi2025
Summary Zhu et al. (NeurIPS 2025) assign each example a distributional influence score — its contribution to the MMD between the training distribution and a clean reference set. Closed-form, retraining-free, approximates the exact LOO ranking within O(1/N²) error, and supports O(mN) online updates as new data arrives.
Bibtex
@inproceedings{zhu2025kairos,
title={KAIROS: Scalable Model-Agnostic Data Valuation},
author={Zhu, Jiongli and Prashant, Parjanya Prajakta and Cloninger, Alex and Salimi, Babak},
booktitle={Advances in Neural Information Processing Systems},
year={2025}
}

Trajectory & Usage-Based Valuation (LLM era)

Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund Sundararajan2020
Summary TracIn traces each training example's influence as the loss change it causes along the actual SGD trajectory, via first-order gradient inner products at checkpoints. Per-step credits telescope across training, require no retraining, and form the foundation of trajectory-based valuation.
Bibtex
@article{pruthi2020estimating,
title={Estimating Training Data Influence by Tracing Gradient Descent},
author={Pruthi, Garima and Liu, Frederick and Kale, Satyen and Sundararajan, Mukund},
journal={Advances in Neural Information Processing Systems},
volume={33},
pages={19920--19930},
year={2020}
}
💻
Data Shapley in One Training RunJiachen T. Wang, Prateek Mittal, Dawn Song, Ruoxi Jia2025
Summary Wang et al. (ICLR 2025) introduce In-Run Data Shapley: Shapley-style attribution for the specific model produced by one training run, accumulated from per-iteration first/second-order credits — with negligible overhead over standard training in its most efficient form. Enables data attribution for foundation-model pretraining for the first time, with implications for copyright and pretraining curation.
Bibtex
@inproceedings{wang2025inrun,
title={Data Shapley in One Training Run},
author={Wang, Jiachen T. and Mittal, Prateek and Song, Dawn and Jia, Ruoxi},
booktitle={International Conference on Learning Representations},
year={2025}
}
What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence FunctionsSang Keun Choe, Hwijeen Ahn, Juhan Bae, Kewen Zhao, Minsoo Kang, Youngseog Chung, Adithya Pratapa, Willie Neiswanger, Emma Strubell, Teruko Mitamura, Jeff Schneider, Eduard Hovy, Roger Grosse, Eric Xing2024
Summary Choe et al. (2024) make influence-function valuation practical at LLM scale with LoGra, a low-rank gradient projection strategy yielding orders-of-magnitude improvements in throughput and storage, and ship LogIX, a package that converts existing training code into data-valuation code. Applied to Llama3-8B-Instruct and a 1B-token dataset.
Bibtex
@article{choe2024your,
title={What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence Functions},
author={Choe, Sang Keun and Ahn, Hwijeen and Bae, Juhan and Zhao, Kewen and Kang, Minsoo and Chung, Youngseog and Pratapa, Adithya and Neiswanger, Willie and Strubell, Emma and Mitamura, Teruko and Schneider, Jeff and Hovy, Eduard and Grosse, Roger and Xing, Eric},
journal={arXiv preprint arXiv:2405.13954},
year={2024}
}
💻
Source Attribution in Retrieval-Augmented GenerationIkhtiyor Nematov, Tarik Kalai, Elizaveta Kuzmenko, Gabriele Fugagnoli, Dimitris Sacharidis, Katja Hose, Tomer Sagi2025
Summary Nematov et al. (2025) adapt Shapley-style attribution to the retrieved sources of a RAG answer, where every utility evaluation is a costly LLM call, and study practical approximations that trade fidelity against cost. A step toward per-query, usage-time valuation — pricing data at inference rather than training.
Bibtex
@article{nematov2025source,
title={Source Attribution in Retrieval-Augmented Generation},
author={Nematov, Ikhtiyor and Kalai, Tarik and Kuzmenko, Elizaveta and Fugagnoli, Gabriele and Sacharidis, Dimitris and Hose, Katja and Sagi, Tomer},
journal={arXiv preprint arXiv:2507.04480},
year={2025}
}

Benchmarks

OpenDataVal: a Unified Benchmark for Data ValuationsKevin Jiang, Weixin Liang, James Zou, Yongchan Kwon2023
Summary Jiang et al. (2023) provides a Python library to build and test data evaluators across different datasets, data evaluators, models, and new benchmarks.
Bibtex
@article{jiang2023opendataval,
title={OpenDataVal: a Unified Benchmark for Data Valuation},
author={Kevin Fu Jiang and Weixin Liang and James Zou and Yongchan Kwon},
booktitle={Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track},
year={2023},
url={https://openreview.net/forum?id=eEK99egXeB}
}
💻🎥

Libraries

influenciaeDeel-AI2023
Summary A stable implementation of influence functions in tensorflow.
Bibtex
💻
pyDVLappliedAI Institute2023
Summary A library of stable and efficient implementations of algorithms for computing Shapley values and influence functions in pytorch.
Bibtex
💻
datascopeBojan Karlaš2024
Summary A tool for repairing datasets using Shapley value as a measure of data importance.
Bibtex
  
💻

Surveys

Data Valuation in Machine Learning: “Ingredients”, Strategies, and Open ChallengesRachael Hwee Ling Sim*, Xinyi Xu*, Bryan Kian Hsiang Low2022
Summary Sim et al. (2022) present a technical survey of data valuation and its "ingredients" and properties. The paper outlines common desiderata as well as some open research challenges.
Bibtex
@inproceedings{sim2022data,
title={Data valuation in machine learning:“ingredients”, strategies, and open challenges},
author={Sim, Rachael Hwee Ling and Xu, Xinyi and Low, Bryan Kian Hsiang},
booktitle={Proc. IJCAI},
year={2022}
}
🎥
Training data influence analysis and estimation: a surveyZayd Hammoudeh and Daniel Lowd2022
Summary Hammoudeh and Lowd (2022) present a technical survey of data valuation, including taxonomy and runtime complexity.
Bibtex
@article{hammoudeh2024training,
title={Training data influence analysis and estimation: A survey},
author={Hammoudeh, Zayd and Lowd, Daniel},
journal={Machine Learning},
volume={113},
number={5},
pages={2351--2403},
year={2024},
publisher={Springer}
}

Designing data marketplaces

Data market system designs

A demonstration of sterling: a privacy-preserving data marketplaceNick Hynes, David Dao, David Yan, Raymond Cheng, Dawn Song2018
Bibtex
@article{hynes2018demonstration,
title={A Demonstration of Sterling: A Privacy-Preserving Data Marketplace},
author={Hynes, Nick and Dao, David and Yan, David and Cheng, Raymond and Song, Dawn},
journal={Proceedings of the VLDB Endowment},
volume={11},
number={12},
year={2018}
}
DataBright: Towards a Global Exchange for Decentralized Data Ownership and Trusted ComputationDavid Dao, Dan Alistarh, Claudiu Musat, Ce Zhang2018
Bibtex
@article{dao2018databright,
title={Databright: Towards a global exchange for decentralized data ownership and trusted computation},
author={Dao, David and Alistarh, Dan and Musat, Claudiu and Zhang, Ce},
journal={arXiv preprint arXiv:1802.04780},
year={2018}
}
A Marketplace for Data: An Algorithmic SolutionAnish Agarwal, Munther Dahleh, Tuhin Sarkar2019
Bibtex
@inproceedings{agarwal2019marketplace,
title={A marketplace for data: An algorithmic solution},
author={Agarwal, Anish and Dahleh, Munther and Sarkar, Tuhin},
booktitle={Proceedings of the 2019 ACM Conference on Economics and Computation},
pages={701--726},
year={2019}
}
Computing a Data DividendEric Bax2019
Bibtex
@misc{bax2019computing,
title={Computing a Data Dividend},
author={Eric Bax},
year={2019},
eprint={1905.01805},
archivePrefix={arXiv},
primaryClass={cs.GT}
}
Incentivizing Collaboration in Machine Learning via Synthetic Data RewardsSebastian Shenghong Tay, Xinyi Xu, Chuan Sheng Foo, Bryan Kian Hsiang Low2021
Bibtex
@article{tay2021incentivizing,
title={Incentivizing Collaboration in Machine Learning via Synthetic Data Rewards},
author={Tay, Sebastian Shenghong and Xu, Xinyi and Foo, Chuan Sheng and Low, Bryan Kian Hsiang},
journal={arXiv preprint arXiv:2112.09327},
year={2021}
}

Automatic data compliance

Data Capsule: A New Paradigm for Automatic Compliance with Data Privacy RegulationsLun Wang, Joseph P. Near, Neel Somani, Peng Gao, Andrew Low, David Dao, Dawn Song2019
Bibtex
@misc{wang2019data,
title={Data Capsule: A New Paradigm for Automatic Compliance with Data Privacy Regulations},
author={Lun Wang and Joseph P. Near and Neel Somani and Peng Gao and Andrew Low and David Dao and Dawn Song},
year={2019},
eprint={1909.00077},
archivePrefix={arXiv},
primaryClass={cs.CY}
}
💻

Data valuation applications

A Principled Approach to Data Valuation for Federated LearningTianhao Wang, Johannes Rausch, Ce Zhang, Ruoxi Jia, Dawn Song2020
Bibtex
@misc{wang2020principled,
title={A Principled Approach to Data Valuation for Federated Learning},
author={Tianhao Wang and Johannes Rausch and Ce Zhang and Ruoxi Jia and Dawn Song},
year={2020},
eprint={2009.06192},
archivePrefix={arXiv},
primaryClass={cs.LG}
}
Data valuation for medical imaging using Shapley value and application to a large-scale chest X-ray datasetSiyi Tang, Amirata Ghorbani, Rikiya Yamashita, Sameer Rehman, Jared A Dunnmon, James Zou, Daniel L Rubin2021
Bibtex
@article{tang2021data,
title={Data valuation for medical imaging using Shapley value and application to a large-scale chest X-ray dataset},
author={Tang, Siyi and Ghorbani, Amirata and Yamashita, Rikiya and Rehman, Sameer and Dunnmon, Jared A and Zou, James and Rubin, Daniel L},
journal={Scientific reports},
volume={11},
number={1},
pages={1--9},
year={2021},
publisher={Nature Publishing Group}
}
Efficient and Fair Data Valuation for Horizontal Federated LearningShuyue Wei, Yongxin Tong, Zimu Zhou, Tianshu Song2020
SummaryAvailability of big data is crucial for modern machine learning applications and services. Federated learning is an emerging paradigm to unite different data owners for machine learning on massive data sets without worrying about data privacy. Yet data owners may still be reluctant to contribute unless their data sets are fairly valuated and paid. In this work, the authors adapt Shapley value, a widely used data valuation metric to valuating data providers in federated learning. Prior data valuation schemes for machine learning incur high computation cost because they require training of extra models on all data set combinations. For efficient data valuation, the authors approximately construct all the models necessary for data valuation using the gradients in training a single model, rather than train an exponential number of models from scratch. On this basis, they devise three methods for efficient contribution index estimation. Evaluations show that their methods accurately approximate the contribution index while notably accelerating its calculation.
Bibtex
@inbook{wei2020efficient,
title={Efficient and fair data valuation for horizontal federated learning},
author={Wei, Shuyue and Tong, Yongxin and Zhou, Zimu and Song, Tianshu},
year={2020},
booktitle={Federated Learning: Privacy and Incentive},
pages={139--152},
publisher={Springer}
}
Improving Fairness for Data Valuation in Horizontal Federated LearningZhenan Fan, Huang Fang, Zirui Zhou, Jian Pei, Michael P. Friedlander, Changxin Liu, Yong Zhang2020
SummaryFederated learning is an emerging decentralized machine learning scheme that allows multiple data owners to work collaboratively while ensuring data privacy. This paper focuses on fairness in data valuation within federated learning. The authors propose a new measure called completed federated Shapley value to improve the fairness of federated Shapley value. This approach leverages the concepts and tools from optimization and provides both theoretical analysis and empirical evaluation to verify the improvement in fairness.
Bibtex
@misc{fan2020improving,
title={Improving Fairness for Data Valuation in Horizontal Federated Learning},
author={Zhenan Fan and Huang Fang and Zirui Zhou and Jian Pei and Michael P. Friedlander and Changxin Liu and Yong Zhang},
year={2020},
eprint={2109.09046},
archivePrefix={arXiv},
primaryClass={cs.LG}
}
Data Valuation for Vertical Federated Learning: An Information-Theoretic ApproachXiao Han, Leye Wang, Junjie Wu2021
SummaryFederated learning (FL) is a machine learning paradigm that enables privacy-preserving cross-party data collaboration. This work introduces "FedValue," the first privacy-preserving, task-specific, model-free data valuation method for vertical FL tasks. It incorporates Shapley-CMI, an information-theoretic metric, for assessing data values from a game-theoretic perspective. The paper also proposes a novel server-aided federated computation mechanism and techniques to accelerate Shapley-CMI computation. Extensive experiments demonstrate the effectiveness and efficiency of FedValue.
Bibtex
@misc{han2021datavaluation,
title={Data Valuation for Vertical Federated Learning: An Information-Theoretic Approach},
author={Xiao Han and Leye Wang and Junjie Wu},
year={2021},
eprint={URL or DOI link TBD},
}
Towards More Efficient Data Valuation in Healthcare Federated Learning Using EnsemblingSourav Kumar, A. Lakshminarayanan, Ken Chang, Feri Guretno, Ivan Ho Mien, Jayashree Kalpathy-Cramer, Pavitra Krishnaswamy, Praveer Singh2021
SummaryThis paper addresses the challenge of data valuation in federated learning within healthcare. The authors propose a method called SaFE (Shapley Value for Federated Learning using Ensembling), which is designed to be efficient in settings where the number of contributing institutions is manageable. SaFE approximates the Shapley value using gradients from training a single model and develops methods for efficient contribution index estimation. This approach is particularly relevant in medical imaging where data heterogeneity is common and fast, accurate data valuation is necessary for multi-institutional collaborations.
Bibtex
@article{Kumar2021TowardsME,
title={Towards More Efficient Data Valuation in Healthcare Federated Learning Using Ensembling},
author={Sourav Kumar and A. Lakshminarayanan and Ken Chang and Feri Guretno and Ivan Ho Mien and Jayashree Kalpathy-Cramer and Pavitra Krishnaswamy and Praveer Singh},
journal={ArXiv},
year={2021},
volume={abs/2209.05424}
}
Data Debugging with Shapley Importance over Machine Learning PipelinesBojan Karlaš, David Dao, Matteo Interlandi, Sebastian Schelter, Wentao Wu, Ce Zhang2024
SummaryThis paper focuses on repairing datasets with the goal of improving the quality of end-to-end machine learning pipelines. Data repairs are prioritized by Shapley value. The authors propose methods for efficiently computing the Shapley value for different types of pipelines and empirically demonstrate the effectiveness of this approach.
Bibtex
@inproceedings{
karlas2024data,
title={Data Debugging with Shapley Importance over Machine Learning Pipelines},
author={Bojan Karla{\v{s}} and David Dao and Matteo Interlandi and Sebastian Schelter and Wentao Wu and Ce Zhang},
booktitle={The Twelfth International Conference on Learning Representations},
year={2024},
url={https://openreview.net/forum?id=qxGXjWxabq}
}
💻🎥

Data markets and society

Economics of Data

Nonrivalry and the Economics of DataCharles I. Jones, Christopher Tonetti2019
Bibtex
@article{10.1257/aer.20191330,
Author = {Jones, Charles I. and Tonetti, Christopher},
Title = {Nonrivalry and the Economics of Data},
Journal = {American Economic Review},
Volume = {110},
Number = {9},
Year = {2020},
Month = {September},
Pages = {2819-58},
DOI = {10.1257/aer.20191330},
URL = {https://www.aeaweb.org/articles?id=10.1257/aer.20191330}
}

Data Dignity

Chapter 5: Data as Labor, Radical MarketsEric A. Posner and E Glen Weyl2019
Bibtex
@book{posner2019radical,
title={Radical Markets},
author={Posner, Eric A and Weyl, E Glen},
year={2019},
publisher={Princeton University Press}
}
Should We Treat Data as Labor? Moving beyond "Free"Imanol Arrieta-Ibarra, Leonard Goff, Diego Jiménez-Hernández, Jaron Lanier, E. Glen Weyl2018
Bibtex
@article{10.1257/pandp.20181003,
Author = {Arrieta-Ibarra, Imanol and Goff, Leonard and Jiménez-Hernández, Diego and Lanier, Jaron and Weyl, E. Glen},
Title = {Should We Treat Data as Labor? Moving beyond "Free"},
Journal = {AEA Papers and Proceedings},
Volume = {108},
Year = {2018},
Month = {May},
Pages = {38-42},
DOI = {10.1257/pandp.20181003},
URL = {https://www.aeaweb.org/articles?id=10.1257/pandp.20181003}
}

Strategic adaptation

Performative prediction

Performative PredictionJuan Perdomo, Tijana Zrnic, Celestine Mendler-Dünner, Moritz Hardt2020
Summary Perdomo et al. (2020) introduce the concept of "performative prediction" dealing with predictions that influence the target they aim to predict, e.g. through taking actions based on the predictions, causing a distribution shift. The authors develop a risk minimization framework for performative prediction and introduce the equilibrium notion of performative stability where predictions are calibrated against future outcomes that manifest from acting on the prediction.
Bibtex
@inproceedings{perdomo2020performative,
title={Performative prediction},
author={Perdomo, Juan and Zrnic, Tijana and Mendler-D{"u}nner, Celestine and Hardt, Moritz},
booktitle={International Conference on Machine Learning},
pages={7599--7609},
year={2020},
organization={PMLR}
}
Stochastic Optimization for Performative PredictionCelestine Mendler-Dünner, Juan Perdomo, Tijana Zrnic, Moritz Hardt2020
Summary Mendler-Dünner et al. (2020) look at stochastic optimization for performative prediction and prove convergence rates for greedily deploying models after each stochastic update (which may cause distribution shift affecting convergence to a stability point) or lazily deploying the model after several updates.
Bibtex
@article{mendler2020stochastic,
title={Stochastic optimization for performative prediction},
author={Mendler-D{"u}nner, Celestine and Perdomo, Juan and Zrnic, Tijana and Hardt, Moritz},
journal={Advances in Neural Information Processing Systems},
volume={33},
pages={4929--4939},
year={2020}
}

Strategic classification

Strategic Classification is Causal Modeling in DisguiseJohn Miller, Smitha Milli, Moritz Hardt2020
Summary Miller et al. (2020) argue that strategic classication involves causal modelling and designing incentives for improvement requires solving a non-trivial causal inference problem. The authors provide a distinction between gaming and improvement as well as provide a causal framework for strategic adaptation.
Bibtex
@inproceedings{miller2020strategic,
title={Strategic classification is causal modeling in disguise},
author={Miller, John and Milli, Smitha and Hardt, Moritz},
booktitle={International Conference on Machine Learning},
pages={6917--6926},
year={2020},
organization={PMLR}
}
Alternative Microfoundations for Strategic ClassificationMeena Jagadeesan, Celestine Mendler-Dünner, Moritz Hardt2021
Summary Jagadeesan et al. (2021) show that standard microfoundations in strategic classification, that typically uses individual-level behaviour to deduce aggregate-level responses, can lead to degenerate behaviour in aggregate: discontinuities in the aggregate response, stable points ceasing to exist, and maximizing social burden. The authors introduce a noisy response model inspired by performative prediction that mitigates these limitations for binary classification.
Bibtex
@inproceedings{jagadeesan2021alternative,
title={Alternative microfoundations for strategic classification},
author={Jagadeesan, Meena and Mendler-D{"u}nner, Celestine and Hardt, Moritz},
booktitle={International Conference on Machine Learning},
pages={4687--4697},
year={2021},
organization={PMLR}
}

Data Valuation Researchers

NameInstituteh-index
Costas SpanosUniversity of California, Berkeley61
Jinsung YoonGoogle Cloud AI33
Tomas PfisterGoogle Cloud AI39
Amirata GhorbaniStanford18
James ZouStanford64
Nektaria TryfonaVirginia Tech27
Rachael Hwee Ling SimNational University of Singapore4
Bryan Kian Hsiang LowNational University of Singapore38
Dawn SongUniversity of California, Berkeley142
Zhaoxuan WuNational University of Singapore4
Xinyi XuNational University of Singapore8
Tianhao WangUniversity of Virginia18
José González CabañasUC3M-Santander Big Data Institute7
Ruben Cuevas RuminUniversidad Carlos III de Madrid26
Jiachen T. WangPrinceton University9
Bohong WangTsinghua University6
Yongchan KwonColumbia University10
Siyi TangArtera8
Li XiongEmory University52
Jessica VitakUniversity of Maryland49
Katie Chamberlain KritikosUniversity of Illinois at Urbana-Champaign6
Zhenan FanHuawei Technologies Canada6
Shuyue WeiBeihang University4
Hannah SteinSaarland University3
Wolfgang MaassSaarland University26
Mohammad Mohammadi AmiriRensselaer Polytechnic Institute18
Ramesh RaskarMIT103
Konstantin D. PandlKarlsruhe Institute of Technolgoy6
Ali SunyaevKarlsruhe Institute of Technolgoy43
Ludovico BorattoUniversity of Cagliari25
han xiao70
Junjie WuCenter for High Pressure Science & Technology Advanced Research55
Xiao TianNational University of Singapore1
Kean BirchInstitute for Technoscience & Society40
Callum WardUppsala University10
Praveer SinghUniversity of Colorado School of Medicine19
Anran XuShanghai Jiao Tong University2
Guihai Chen67
Andre EstevaCo-Founder & CEO, Artera23
Prateek MittalPrinceton University55
Hyeontaek OhInstitute for IT Convergence9
Lingjiao ChenStanford13
Xiangyu ChangXi'an Jiaotong University17
Hoang Anh JustVirginia Tech3
David DaoETH13
Mark MazumderHarvard12
Vijay Janapa ReddiHarvard46
Sabri EyubogluStanford6
Wenqian LiNational University of Singapore2
Bojan KarlašHarvard14
ai
data-valuation
market
ml

daviddao/awesome-data-valuation

💱 A curated list of data valuation (DV) to design your next data marketplace

144

53 commits

updated Aug 14, 2026

See the code

README

Awesome Data Valuation

data market problem

💱 A curated list of data valuation (DV) to design your next data marketplace. DV aims to understand the value of a data point for a given machine learning task and is an essential primitive in the design of data marketplaces and explainable AI.

Legend

💻 Code available

🎥 Talk / Slides

Contents

What is your data worth?

Shapley Value & Cooperative Game Theory

Towards Efficient Data Valuation Based on the Shapley ValueRuoxi Jia & David Dao, Boxin Wang, Frances Ann Hubis, Nick Hynes, Nezihe Merve Gürel, Bo Li, Ce Zhang, Dawn Song, Costas J. Spanos2019
Summary Jia et al. (2019) contribute theoretical and practical results for efficient methods for approximating the Shapley value (SV). They show that methods with a sublinear amount of model evaluations are possible and further reductions can be made for sparse SVs. Lastly, they introduce two practical SV estimation methods for ML tasks, one for uniformly stable learning algorithms and one for smooth loss functions.
Bibtex
@inproceedings{jia2019towards,
title={Towards efficient data valuation based on the shapley value},
author={Jia, Ruoxi and Dao, David and Wang, Boxin and Hubis, Frances Ann and Hynes, Nick and G{"u}rel, Nezihe Merve and Li, Bo and Zhang, Ce and Song, Dawn and Spanos, Costas J},
booktitle={The 22nd International Conference on Artificial Intelligence and Statistics},
pages={1167--1176},
year={2019},
organization={PMLR}
}
💻
Data Shapley: Equitable Valuation of Data for Machine LearningAmirata Ghorbani, James Zou2019
Summary Ghorbani & Zou (2019) introduce (data) Shapley value to equitably measure the value of each training point to a supervised learners performance. They further outline several benefits of the Shapley value, e.g. being able to capture outliers or inform what new data to acquire, as well as develop Monte Carlo and gradient-based methods for its efficient estimation.
Bibtex
@inproceedings{ghorbani2019data,
title={Data shapley: Equitable valuation of data for machine learning},
author={Ghorbani, Amirata and Zou, James},
booktitle={International Conference on Machine Learning},
pages={2242--2251},
year={2019},
organization={PMLR}
}
💻
A Distributional Framework for Data ValuationAmirata Ghorbani, Michael P. Kim, James Zou2020
Summary Ghorbani et al. (2020) formulate the Shapley value as a distributional quantity in the context of an underlying data distribution instead of a fixed dataset. They further introduce a novel sampling-based algorithm for the distributional Shapley value with strong approximation guarantees.
Bibtex
@inproceedings{ghorbani2020distributional,
title={A Distributional Framework for Data Valuation},
author={Ghorbani, Amirata, P. Kim, Michael and Zou, James},
booktitle={International Conference on Machine Learning},
year={2020}
}
💻
Asymmetric Shapley values: incorporating causal knowledge into model-agnostic explainabilityChristopher Frye, Colin Rowat, Ilya Feige2020
Summary Frye et al. (2020) incorporate causality into the Shapley value framework. Importantly, their framework can handle any amount of causal knowledge and does not require the complete causal graph underlying the data.
Bibtex
@article{frye2020asymmetric,
title={Asymmetric Shapley values: incorporating causal knowledge into model-agnostic explainability},
author={Frye, Christopher and Rowat, Colin and Feige, Ilya},
journal={Advances in Neural Information Processing Systems},
volume={33},
year={2020}
}
🎥
Collaborative Machine Learning with Incentive-Aware Model RewardsRachael Hwee Ling Sim, Yehong Zhang, Mun Choon Chan, Bryan Kian Hsiang Low2020
Summary Sim et al. (2020) introduce a data valuation method with separate ML models as rewards based on the Shapley value and information gain on model parameters given its data. They further define several conditions for incentives such as Shapley fairness, stability, individual rationality, and group welfare, that are suitable for the freely replicable nature of their model reward scheme.
Bibtex
@inproceedings{sim2020collaborative,
title={Collaborative machine learning with incentive-aware model rewards},
author={Sim, Rachael Hwee Ling and Zhang, Yehong and Chan, Mun Choon and Low, Bryan Kian Hsiang},
booktitle={International Conference on Machine Learning},
pages={8927--8936},
year={2020},
organization={PMLR}
}
Validation free and replication robust volume-based data valuationXinyi Xu, Zhaoxuan Wu, Chuan Sheng Foo, Bryan Kian Hsiang Low2021
Summary Xu et al. (2021) propose using data diversity via robust volume for measuring the value of data. This removes the need for a validation set and allows for guarantees on replication robustness but suffers from the curse of dimensionality and may ignore useful information in the validation set.
Bibtex
@article{xu2021validation,
title={Validation free and replication robust volume-based data valuation},
author={Xu, Xinyi and Wu, Zhaoxuan and Foo, Chuan Sheng and Low, Bryan Kian Hsiang},
journal={Advances in Neural Information Processing Systems},
volume={34},
year={2021}
}
💻
Beta Shapley: a Unified and Noise-reduced Data Valuation Framework for Machine LearningYongchan Kwon, James Zou2021
Summary Kwon & Zou (2022) introduce Beta Shapley, a generalization of Data Shapley by relaxing the efficiency axiom.
Bibtex
@article{kwon2021beta,
title={Beta Shapley: a Unified and Noise-reduced Data Valuation Framework for Machine Learning},
author={Kwon, Yongchan and Zou, James},
journal={arXiv preprint arXiv:2110.14049},
year={2021}
}
Gradient-Driven Rewards to Guarantee Fairness in Collaborative Machine LearningXinyi Xu, Lingjuan Lyu, Xingjun Ma, Chenglin Miao, Chuan Sheng Foo, Bryan Kian Hsiang Low2021
Summary Xu et al. (2021) propose cosine gradient Shapley value to fairly evaluate the expected contribution of each agent's update in the federated learning setting removing the need for an auxiliary validation dataset. They further introduce a novel training-time gradient reward mechanism with a fairness guarantee.
Bibtex
@article{xu2021gradient,
title={Gradient driven rewards to guarantee fairness in collaborative machine learning},
author={Xu, Xinyi and Lyu, Lingjuan and Ma, Xingjun and Miao, Chenglin and Foo, Chuan Sheng and Low, Bryan Kian Hsiang},
journal={Advances in Neural Information Processing Systems},
volume={34},
pages={16104--16117},
year={2021}
}
Improving Cooperative Game Theory-based Data Valuation via Data Utility LearningTianhao Wang, Yu Yang, Ruoxi Jia2022
Summary Wang et al. (2022) propose a general framework to improve effectiveness of sampling-based Shapley value (SV) or Least core (LC) estimation heuristics. They propose learning to predict the performance of a learning algorithm (denoted data utility learning) and using this predictor to estimate learning performance without retraining for cheaper SV and LC estimation.
Bibtex
@article{wang2021improving,
title={Improving cooperative game theory-based data valuation via data utility learning},
author={Wang, Tianhao and Yang, Yu and Jia, Ruoxi},
journal={arXiv preprint arXiv:2107.06336},
year={2021}
}
Data Banzhaf: A Robust Data Valuation Framework for Machine LearningJiachen T. Wang, Ruoxi Jia2023
Summary Wang et al. (2023) propose using the Banzhaf value for data valuation, providing better robustness against noisy performance scores and an efficient estimate using Maximum Sample Reuse (MSR) principle
Bibtex
@InProceedings{pmlr-v206-wang23e, title={Data Banzhaf: A Robust Data Valuation Framework for Machine Learning},
author={Wang, Jiachen T. and Jia, Ruoxi},
booktitle={Proceedings of The 26th International Conference on Artificial Intelligence and Statistics},
pages={6388--6421},
year={2023},
editor={Ruiz, Francisco and Dy, Jennifer and van de Meent, Jan-Willem},
volume={206},
series={Proceedings of Machine Learning Research},
month={25--27 Apr},
publisher={PMLR},
pdf={https://proceedings.mlr.press/v206/wang23e/wang23e.pdf},
url={https://proceedings.mlr.press/v206/wang23e.html}
}
💻
A Multilinear Sampling Algorithm to Estimate Shapley ValuesRamin Okhrati, Aldo Lipani2021
Summary Okhrati and Lipani (2021) propose a new sampling method for Shapley values based on a multilinear extension technique as applied in game theory. It provides more accurate estimations of the Shapley values by reducing the variance of the sampling statistics.
Bibtex
@INPROCEEDINGS{9412511,
title={A Multilinear Sampling Algorithm to Estimate Shapley Values},
author={Okhrati, Ramin and Lipani, Aldo},
booktitle={2020 25th International Conference on Pattern Recognition (ICPR)},
year={2021}
}
💻
If You Like Shapley Then You’ll Love the CoreYan, T., and Procaccia, A. D.2021
Summary Yan and Procaccia (2021) propose an alternative method for credit assignment in data valuation. They use the least core, which can be computed efficiently.
Bibtex
@article{Yan_Procaccia_2021,
title={If You Like Shapley Then You’ll Love the Core},
author={Yan, Tom and Procaccia, Ariel D.},
journal={Proceedings of the AAAI Conference on Artificial Intelligence},
year={2021}
}
CS-Shapley: Class-wise Shapley Values for Data Valuation in ClassificationSchoch, Stephanie, Haifeng Xu, and Yangfeng Ji2022
Summary Schoch et al. (2022) propose a new Shapley value that discriminates between training instances' in-class and out-of-class contributions.
Bibtex
@inproceedings{schoch2022csshapley,
title={{CS}-Shapley: Class-wise Shapley Values for Data Valuation in Classification},
author={Stephanie Schoch and Haifeng Xu and Yangfeng Ji},
booktitle={Advances in Neural Information Processing Systems},
year={2022}
}
💻
Precedence-Constrained Winter Value for Effective Graph Data ValuationHongliang Chi, Wei Jin, Charu Aggarwal, and Yao Ma2022
Summary Introduces a new data valuation concept similar to Data Shapley but based on the Winter value.
Bibtex
@misc{chi2024wintervalue,
title={Precedence-Constrained Winter Value for Effective Graph Data Valuation},
author={Hongliang Chi and Wei Jin and Charu Aggarwal and Yao Ma},
year={2024},
eprint={2402.01943},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2402.01943},
}

Efficient algorithms

Efficient Task-Specific Data Valuation for Nearest Neighbor AlgorithmsRuoxi Jia, David Dao, Boxin Wang, Frances Ann Hubis, Nezihe Merve Gurel, Bo Li, Ce Zhang, Costas J. Spanos, Dawn Song2019
Summary Jia et al. (2019) present algorithms to compute the Shapley value exactly in quasi-linear time and approximations in sublinear time for k-nearest-neighbor models. They empirically evaluate their algorithms at scale and extend them to several other settings.
Bibtex
@article{jia12efficient,
title={Efficient Task-Specific Data Valuation for Nearest Neighbor Algorithms},
author={Jia, Ruoxi and Dao, David and Wang, Boxin and Hubis, Frances Ann and Gurel, Nezihe Merve and Zhang, Bo Li4 Ce and Song, Costas Spanos1 Dawn},
journal={Proceedings of the VLDB Endowment},
volume={12},
number={11}
}
💻
Efficient computation and analysis of distributional Shapley valuesYongchan Kwon, Manuel A. Rivas, James Zou2021
Summary Kwon et al. (2021) develop tractable analytic expressions for the distributional data Shapley value for linear regression, binary classification, and non-parametric density estimation as well as new efficient methods for its estimation.
Bibtex
@inproceedings{kwon2021efficient,
title={Efficient computation and analysis of distributional Shapley values},
author={Kwon, Yongchan and Rivas, Manuel A and Zou, James},
booktitle={International Conference on Artificial Intelligence and Statistics},
pages={793--801},
year={2021},
organization={PMLR}
}
💻
DUPRE: Data Utility Prediction for Efficient Data ValuationKieu Thao Nguyen Pham, Rachael Hwee Ling Sim, Quoc Phong Nguyen, See Kiong Ng, Bryan Kian Hsiang Low2025
Summary Pham et al. (2025) speed up cooperative game-theoretic valuation from the other direction: instead of reducing the number of subsets to evaluate, DUPRE predicts the utility of unseen data subsets with a learned model, skipping costly retraining for each coalition.
Bibtex
@article{pham2025dupre,
title={DUPRE: Data Utility Prediction for Efficient Data Valuation},
author={Pham, Kieu Thao Nguyen and Sim, Rachael Hwee Ling and Nguyen, Quoc Phong and Ng, See Kiong and Low, Bryan Kian Hsiang},
journal={arXiv preprint arXiv:2502.16152},
year={2025}
}

Benchmarks, Criticism & Relaxations

Scalability vs. Utility: Do We Have to Sacrifice One for the Other in Data Importance Quantification?Ruoxi Jia, Fan Wu, Xuehui Sun, Jiacen Xu, David Dao, Bhavya Kailkhura, Ce Zhang, Bo Li, Dawn Song2021
Summary Jia et al. (2021) perform a theoretical analysis on the differences between leave-one-out-based and Shapley value-based methods as well as an empirical study across several ML tasks investigating the two aforementioned methods as well as exact Shapley value-based methods and Shapley over KNN Surrogates.
Bibtex
@misc{jia2021scalability,
title={Scalability vs. Utility: Do We Have to Sacrifice One for the Other in Data Importance Quantification?},
author={Ruoxi Jia and Fan Wu and Xuehui Sun and Jiacen Xu and David Dao and Bhavya Kailkhura and Ce Zhang and Bo Li and Dawn Song},
year={2021},
eprint={1911.07128},
archivePrefix={arXiv},
primaryClass={cs.LG}
}
💻
Shapley values for feature selection: The good, the bad, and the axiomsDaniel Fryer, Inga Strümke, Hien Nguyen2021
Summary Fryer et al. (2021) calls into question the appropriateness of using the Shapley value for feature selection and advise caution against the magical thinking that presenting its abstract general axioms as "favourable and fair" may introduce. They further point out that the four axioms of "efficiency", "null player", "symmetry", and "additivity" do not guarantee that the Shapley value is suited to feature selection and may sometimes even imply the opposite.
Bibtex
@misc{fryer2021shapley,
title={Shapley values for feature selection: The good, the bad, and the axioms},
author={Daniel Fryer and Inga Strümke and Hien Nguyen},
year={2021},
eprint={2102.10936},
archivePrefix={arXiv},
primaryClass={cs.LG}
}
Rethinking Data Shapley for Data Selection Tasks: Misleads and MeritsJiachen T. Wang, Tianji Yang, James Zou, Yongchan Kwon, Ruoxi Jia2024
Summary Wang et al. (ICML 2024 oral) show via a hypothesis-testing framework that Data Shapley's data-selection performance can be no better than random without constraints on the utility function, and identify the class of utility functions (monotonically transformed modular functions) under which it selects optimally.
Bibtex
@inproceedings{pmlr-v235-wang24cg,
title={Rethinking Data Shapley for Data Selection Tasks: Misleads and Merits},
author={Wang, Jiachen T. and Yang, Tianji and Zou, James and Kwon, Yongchan and Jia, Ruoxi},
booktitle={Proceedings of the 41st International Conference on Machine Learning},
pages={52033--52063},
year={2024},
publisher={PMLR}
}
Semivalue-based data valuation is arbitrary and gameableHannah Diehl, Ashia C. Wilson2025
Summary Diehl & Wilson (2025) prove that semivalue-based valuations (incl. Data Shapley, Beta Shapley, Banzhaf) are underdetermined: defensible changes to the utility specification can flip rankings, and strategic actors can exploit this latitude to game outcomes. A foundational challenge to using axiomatic valuations for high-stakes decisions.
Bibtex
@article{diehl2025semivalue,
title={Semivalue-based data valuation is arbitrary and gameable},
author={Diehl, Hannah and Wilson, Ashia C.},
journal={arXiv preprint arXiv:2506.12619},
year={2025}
}
Do Data Valuations Make Good Data Prices?Dongyang Fan, Tyler J. Rotello, Sai Praneeth Karimireddy2025
Summary Fan et al. (2025) revisit data valuation from a market-design perspective and show that popular valuations (LOO, Data Shapley) make poor payments: they fail truthfulness and cost-coverage for heterogeneous data owners. Attribution and pricing are different problems — a mechanism layer is needed on top of valuation.
Bibtex
@article{fan2025data,
title={Do Data Valuations Make Good Data Prices?},
author={Fan, Dongyang and Rotello, Tyler J. and Karimireddy, Sai Praneeth},
journal={arXiv preprint arXiv:2504.05563},
year={2025}
}

Influence functions & LOO

Understanding Black-box Predictions via Influence FunctionsPang Wei Koh, Percy Liang2017
Summary Koh & Liang (2017) introduce the use of influence functions, a technique borrowed from robust statistics, to identify training points most responsible for a model's given prediction without needing to retrain. They further develop a simple and efficient implementation of influence functions that scales to large ML settings.
Bibtex
@inproceedings{koh2017understanding,
title={Understanding black-box predictions via influence functions},
author={Koh, Pang Wei and Liang, Percy},
booktitle={International Conference on Machine Learning},
pages={1885--1894},
year={2017},
organization={PMLR}
}
💻🎥
On the accuracy of influence functions for measuring group effectsPang Wei Koh*, Kai-Siang Ang*, Hubert H. K. Teo*, and Percy Liang2019
Summary Koh et al. (2019) study influence functions to measure effects of large groups of training points instead of individual points. They empirically find a correlation and often underestimation between predicted and actual effects and theoretically show that this need not hold in general, realistic settings.
Bibtex
@article{koh2019accuracy,
title={On the accuracy of influence functions for measuring group effects},
author={Koh, Pang Wei and Ang, Kai-Siang and Teo, Hubert HK and Liang, Percy},
journal={arXiv preprint arXiv:1905.13289},
year={2019}
}
💻🎥
Scaling Up Influence FunctionsSchioppa, Andrea, Polina Zablotskaia, David Vilar, and Artem Sokolov2022
Summary Schioppa et al. (2022) propose a new method to scale the computation of influence functions for large neural networks using the Arnoldi iteration. With this, they achieve successful implementation of influence functions on full-size Transformer models with hundreds of millions of parameters.
Bibtex
@inproceedings{schioppa2022scaling,
title={Scaling Up Influence Functions},
author={Schioppa, Andrea and Zablotskaia, Polina and Vilar, David and Sokolov, Artem},
booktitle={Proceedings of the AAAI Conference on Artificial Intelligence},
year={2022}
}
💻
Studying large language model generalization with influence functionsGrosse, Roger and Bae, Juhan and Anil, Cem and Elhage, Nelson and Tamkin, Alex and Tajdini, Amirhossein and Steiner, Benoit and Li, Dustin and Durmus, Esin and Perez, Ethan and others2023
Summary Grosse et al. (2023) use a method known as EK-FAC to approximate the Hessian of the loss of large language models. They apply this technique to study influence functions on large language models, up to 50 billion parameters.
Bibtex
@article{grosse2023studying,
title={Studying large language model generalization with influence functions},
author={Grosse, Roger and Bae, Juhan and Anil, Cem and Elhage, Nelson and Tamkin, Alex and Tajdini, Amirhossein and Steiner, Benoit and Li, Dustin and Durmus, Esin and Perez, Ethan and others},
journal={arXiv preprint arXiv:2308.03296},
year={2023}
}
Do Influence Functions Work on Large Language Models?Zhe Li, Wei Zhao, Yige Li, Jun Sun2024
Summary Li et al. (2024) systematically evaluate influence functions across multiple LLM tasks and find they consistently perform poorly, tracing the failures to approximation errors in inverse-Hessian estimation, uncertain convergence during fine-tuning, and the definition of influence itself. A caution for influence-based valuation at LLM scale.
Bibtex
@article{li2024influence,
title={Do Influence Functions Work on Large Language Models?},
author={Li, Zhe and Zhao, Wei and Li, Yige and Sun, Jun},
journal={arXiv preprint arXiv:2409.19998},
year={2024}
}

Reinforcement Learning

Data Valuation using Reinforcement LearningJinsung Yoon, Sercan Ö Arık, Tomas Pfister2020
Summary Yoon et al. (2020) propose using reinforcement learning for data valuation to learn data values jointly with the predictor model.
Bibtex
@inproceedings{49189,
title={Data Valuation using Reinforcement Learning},
author={Jinsung Yoon and Sercan Arik and Tomas Pfister},
year={2020}
}
💻🎥

Deep Neural Networks

DAVINZ: Data Valuation using Deep Neural Networks at InitializationZhaoxuan Wu, Yao Shu, Bryan Kian Hsiang Low2022
Summary Wu et al. (2022) introduce a validation-based and training-free method for efficient data valuation with large and complex deep neural networks (DNNs). They derive and exploit a domain-aware generalization bound for DNNs to characterize their performance without training and uses this bound as the scoring function while keeping conventional techniques such as Shapley values as the valuation function.
Bibtex
@inproceedings{wu2022davinz,
title={DAVINZ: Data Valuation using Deep Neural Networks at Initialization},
author={Wu, Zhaoxuan and Shu, Yao and Low, Bryan Kian Hsiang},
booktitle={International Conference on Machine Learning},
pages={24150--24176},
year={2022},
organization={PMLR}
}
🎥
LossVal: Efficient Data Valuation for Neural NetworksTim Wibiral, Mohamed Karim Belaid, Maximilian Rabus, and Ansgar Scherp2024
Summary Wibiral et al. (2024) introduce a efficient method for deep neural networks that learns the importance scores of the data valuation as weights of the loss function during the first training run.
Bibtex
@misc{wibiral2024lossvalefficientdatavaluation,
title={LossVal: Efficient Data Valuation for Neural Networks},
author={Tim Wibiral and Mohamed Karim Belaid and Maximilian Rabus and Ansgar Scherp},
year={2024},
eprint={2412.04158},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2412.04158},
}
💻

Out-of-Bag score

Data-OOB: Out-of-bag Estimate as a Simple and Efficient Data ValueYongchan Kwon, James Zou2023
Summary Kwon et al. (2023) propose using the out-of-bag estimate of a bagging estimator for computationally efficient data valuation.
Bibtex
@inproceedings{DBLP:conf/icml/Kwon023, 
author={Yongchan Kwon and James Zou},
editor={Andreas Krause and Emma Brunskill and Kyunghyun Cho and Barbara Engelhardt and Sivan Sabato and Jonathan Scarlett},
title={Data-OOB: Out-of-bag Estimate as a Simple and Efficient Data Value},
booktitle={International Conference on Machine Learning, {ICML} 2023, 23-29 July 2023, Honolulu, Hawaii, {USA}},
series={Proceedings of Machine Learning Research},
volume={202},
pages={18135--18152},
publisher={{PMLR}},
year={2023},
url={https://proceedings.mlr.press/v202/kwon23e.html},
timestamp={Mon, 28 Aug 2023 17:23:08 +0200},
biburl={https://dblp.org/rec/conf/icml/Kwon023.bib},
bibsource={dblp computer science bibliography, https://dblp.org}
}
💻🎥

Evolutionary Approaches

An evolutionary approach to data valuationNatalia Khuri, Sapan Bhandari, Esteban Murillo Burford, Nathan P. Whitener, and Konghao Zhao2022
Summary Khuri et al. (2022) propose using an evolutionary algorithm for data valuation.
Bibtex
@inproceedings{10.1145/3535508.3545522,
author = {Khuri, Natalia and Bhandari, Sapan and Burford, Esteban Murillo and Whitener, Nathan P. and Zhao, Konghao},
title = {An evolutionary approach to data valuation},
year = {2022},
isbn = {9781450393867},
publisher = {Association for Computing Machinery},
address = {New York, NY, USA},
url = {https://doi.org/10.1145/3535508.3545522},
doi = {10.1145/3535508.3545522}
}

Task Agnostic

Fundamentals of Task-Agnostic Data ValuationMohammad Mohammadi Amiri, Frederic Berdoz, Ramesh Raskar2023
SummaryThis paper addresses the challenge of valuing data without specific task assumptions, focusing on task-agnostic data valuation. It discusses valuing a data seller's dataset from a buyer's perspective without validation requirements. The approach involves estimating statistical differences through diversity and relevance measures without needing the raw data, and designing queries that maintain the seller's blindness to the buyer's raw data. The work is significant for practical scenarios where utility metrics like test accuracy on a validation set are not feasible.
Bibtex
@article{Amiri2023FundamentalsOT,
title={Fundamentals of Task-Agnostic Data Valuation},
author={Mohammad Mohammadi Amiri and Frederic Berdoz and Ramesh Raskar},
journal={Proceedings of the AAAI Conference on Artificial Intelligence},
volume={37},
pages={9226-9234},
year={2023},
doi={10.1609/aaai.v37i8.26106}
}
KAIROS: Scalable Model-Agnostic Data ValuationJiongli Zhu, Parjanya Prajakta Prashant, Alex Cloninger, Babak Salimi2025
Summary Zhu et al. (NeurIPS 2025) assign each example a distributional influence score — its contribution to the MMD between the training distribution and a clean reference set. Closed-form, retraining-free, approximates the exact LOO ranking within O(1/N²) error, and supports O(mN) online updates as new data arrives.
Bibtex
@inproceedings{zhu2025kairos,
title={KAIROS: Scalable Model-Agnostic Data Valuation},
author={Zhu, Jiongli and Prashant, Parjanya Prajakta and Cloninger, Alex and Salimi, Babak},
booktitle={Advances in Neural Information Processing Systems},
year={2025}
}

Trajectory & Usage-Based Valuation (LLM era)

Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund Sundararajan2020
Summary TracIn traces each training example's influence as the loss change it causes along the actual SGD trajectory, via first-order gradient inner products at checkpoints. Per-step credits telescope across training, require no retraining, and form the foundation of trajectory-based valuation.
Bibtex
@article{pruthi2020estimating,
title={Estimating Training Data Influence by Tracing Gradient Descent},
author={Pruthi, Garima and Liu, Frederick and Kale, Satyen and Sundararajan, Mukund},
journal={Advances in Neural Information Processing Systems},
volume={33},
pages={19920--19930},
year={2020}
}
💻
Data Shapley in One Training RunJiachen T. Wang, Prateek Mittal, Dawn Song, Ruoxi Jia2025
Summary Wang et al. (ICLR 2025) introduce In-Run Data Shapley: Shapley-style attribution for the specific model produced by one training run, accumulated from per-iteration first/second-order credits — with negligible overhead over standard training in its most efficient form. Enables data attribution for foundation-model pretraining for the first time, with implications for copyright and pretraining curation.
Bibtex
@inproceedings{wang2025inrun,
title={Data Shapley in One Training Run},
author={Wang, Jiachen T. and Mittal, Prateek and Song, Dawn and Jia, Ruoxi},
booktitle={International Conference on Learning Representations},
year={2025}
}
What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence FunctionsSang Keun Choe, Hwijeen Ahn, Juhan Bae, Kewen Zhao, Minsoo Kang, Youngseog Chung, Adithya Pratapa, Willie Neiswanger, Emma Strubell, Teruko Mitamura, Jeff Schneider, Eduard Hovy, Roger Grosse, Eric Xing2024
Summary Choe et al. (2024) make influence-function valuation practical at LLM scale with LoGra, a low-rank gradient projection strategy yielding orders-of-magnitude improvements in throughput and storage, and ship LogIX, a package that converts existing training code into data-valuation code. Applied to Llama3-8B-Instruct and a 1B-token dataset.
Bibtex
@article{choe2024your,
title={What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence Functions},
author={Choe, Sang Keun and Ahn, Hwijeen and Bae, Juhan and Zhao, Kewen and Kang, Minsoo and Chung, Youngseog and Pratapa, Adithya and Neiswanger, Willie and Strubell, Emma and Mitamura, Teruko and Schneider, Jeff and Hovy, Eduard and Grosse, Roger and Xing, Eric},
journal={arXiv preprint arXiv:2405.13954},
year={2024}
}
💻
Source Attribution in Retrieval-Augmented GenerationIkhtiyor Nematov, Tarik Kalai, Elizaveta Kuzmenko, Gabriele Fugagnoli, Dimitris Sacharidis, Katja Hose, Tomer Sagi2025
Summary Nematov et al. (2025) adapt Shapley-style attribution to the retrieved sources of a RAG answer, where every utility evaluation is a costly LLM call, and study practical approximations that trade fidelity against cost. A step toward per-query, usage-time valuation — pricing data at inference rather than training.
Bibtex
@article{nematov2025source,
title={Source Attribution in Retrieval-Augmented Generation},
author={Nematov, Ikhtiyor and Kalai, Tarik and Kuzmenko, Elizaveta and Fugagnoli, Gabriele and Sacharidis, Dimitris and Hose, Katja and Sagi, Tomer},
journal={arXiv preprint arXiv:2507.04480},
year={2025}
}

Benchmarks

OpenDataVal: a Unified Benchmark for Data ValuationsKevin Jiang, Weixin Liang, James Zou, Yongchan Kwon2023
Summary Jiang et al. (2023) provides a Python library to build and test data evaluators across different datasets, data evaluators, models, and new benchmarks.
Bibtex
@article{jiang2023opendataval,
title={OpenDataVal: a Unified Benchmark for Data Valuation},
author={Kevin Fu Jiang and Weixin Liang and James Zou and Yongchan Kwon},
booktitle={Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track},
year={2023},
url={https://openreview.net/forum?id=eEK99egXeB}
}
💻🎥

Libraries

influenciaeDeel-AI2023
Summary A stable implementation of influence functions in tensorflow.
Bibtex
💻
pyDVLappliedAI Institute2023
Summary A library of stable and efficient implementations of algorithms for computing Shapley values and influence functions in pytorch.
Bibtex
💻
datascopeBojan Karlaš2024
Summary A tool for repairing datasets using Shapley value as a measure of data importance.
Bibtex
  
💻

Surveys

Data Valuation in Machine Learning: “Ingredients”, Strategies, and Open ChallengesRachael Hwee Ling Sim*, Xinyi Xu*, Bryan Kian Hsiang Low2022
Summary Sim et al. (2022) present a technical survey of data valuation and its "ingredients" and properties. The paper outlines common desiderata as well as some open research challenges.
Bibtex
@inproceedings{sim2022data,
title={Data valuation in machine learning:“ingredients”, strategies, and open challenges},
author={Sim, Rachael Hwee Ling and Xu, Xinyi and Low, Bryan Kian Hsiang},
booktitle={Proc. IJCAI},
year={2022}
}
🎥
Training data influence analysis and estimation: a surveyZayd Hammoudeh and Daniel Lowd2022
Summary Hammoudeh and Lowd (2022) present a technical survey of data valuation, including taxonomy and runtime complexity.
Bibtex
@article{hammoudeh2024training,
title={Training data influence analysis and estimation: A survey},
author={Hammoudeh, Zayd and Lowd, Daniel},
journal={Machine Learning},
volume={113},
number={5},
pages={2351--2403},
year={2024},
publisher={Springer}
}

Designing data marketplaces

Data market system designs

A demonstration of sterling: a privacy-preserving data marketplaceNick Hynes, David Dao, David Yan, Raymond Cheng, Dawn Song2018
Bibtex
@article{hynes2018demonstration,
title={A Demonstration of Sterling: A Privacy-Preserving Data Marketplace},
author={Hynes, Nick and Dao, David and Yan, David and Cheng, Raymond and Song, Dawn},
journal={Proceedings of the VLDB Endowment},
volume={11},
number={12},
year={2018}
}
DataBright: Towards a Global Exchange for Decentralized Data Ownership and Trusted ComputationDavid Dao, Dan Alistarh, Claudiu Musat, Ce Zhang2018
Bibtex
@article{dao2018databright,
title={Databright: Towards a global exchange for decentralized data ownership and trusted computation},
author={Dao, David and Alistarh, Dan and Musat, Claudiu and Zhang, Ce},
journal={arXiv preprint arXiv:1802.04780},
year={2018}
}
A Marketplace for Data: An Algorithmic SolutionAnish Agarwal, Munther Dahleh, Tuhin Sarkar2019
Bibtex
@inproceedings{agarwal2019marketplace,
title={A marketplace for data: An algorithmic solution},
author={Agarwal, Anish and Dahleh, Munther and Sarkar, Tuhin},
booktitle={Proceedings of the 2019 ACM Conference on Economics and Computation},
pages={701--726},
year={2019}
}
Computing a Data DividendEric Bax2019
Bibtex
@misc{bax2019computing,
title={Computing a Data Dividend},
author={Eric Bax},
year={2019},
eprint={1905.01805},
archivePrefix={arXiv},
primaryClass={cs.GT}
}
Incentivizing Collaboration in Machine Learning via Synthetic Data RewardsSebastian Shenghong Tay, Xinyi Xu, Chuan Sheng Foo, Bryan Kian Hsiang Low2021
Bibtex
@article{tay2021incentivizing,
title={Incentivizing Collaboration in Machine Learning via Synthetic Data Rewards},
author={Tay, Sebastian Shenghong and Xu, Xinyi and Foo, Chuan Sheng and Low, Bryan Kian Hsiang},
journal={arXiv preprint arXiv:2112.09327},
year={2021}
}

Automatic data compliance

Data Capsule: A New Paradigm for Automatic Compliance with Data Privacy RegulationsLun Wang, Joseph P. Near, Neel Somani, Peng Gao, Andrew Low, David Dao, Dawn Song2019
Bibtex
@misc{wang2019data,
title={Data Capsule: A New Paradigm for Automatic Compliance with Data Privacy Regulations},
author={Lun Wang and Joseph P. Near and Neel Somani and Peng Gao and Andrew Low and David Dao and Dawn Song},
year={2019},
eprint={1909.00077},
archivePrefix={arXiv},
primaryClass={cs.CY}
}
💻

Data valuation applications

A Principled Approach to Data Valuation for Federated LearningTianhao Wang, Johannes Rausch, Ce Zhang, Ruoxi Jia, Dawn Song2020
Bibtex
@misc{wang2020principled,
title={A Principled Approach to Data Valuation for Federated Learning},
author={Tianhao Wang and Johannes Rausch and Ce Zhang and Ruoxi Jia and Dawn Song},
year={2020},
eprint={2009.06192},
archivePrefix={arXiv},
primaryClass={cs.LG}
}
Data valuation for medical imaging using Shapley value and application to a large-scale chest X-ray datasetSiyi Tang, Amirata Ghorbani, Rikiya Yamashita, Sameer Rehman, Jared A Dunnmon, James Zou, Daniel L Rubin2021
Bibtex
@article{tang2021data,
title={Data valuation for medical imaging using Shapley value and application to a large-scale chest X-ray dataset},
author={Tang, Siyi and Ghorbani, Amirata and Yamashita, Rikiya and Rehman, Sameer and Dunnmon, Jared A and Zou, James and Rubin, Daniel L},
journal={Scientific reports},
volume={11},
number={1},
pages={1--9},
year={2021},
publisher={Nature Publishing Group}
}
Efficient and Fair Data Valuation for Horizontal Federated LearningShuyue Wei, Yongxin Tong, Zimu Zhou, Tianshu Song2020
SummaryAvailability of big data is crucial for modern machine learning applications and services. Federated learning is an emerging paradigm to unite different data owners for machine learning on massive data sets without worrying about data privacy. Yet data owners may still be reluctant to contribute unless their data sets are fairly valuated and paid. In this work, the authors adapt Shapley value, a widely used data valuation metric to valuating data providers in federated learning. Prior data valuation schemes for machine learning incur high computation cost because they require training of extra models on all data set combinations. For efficient data valuation, the authors approximately construct all the models necessary for data valuation using the gradients in training a single model, rather than train an exponential number of models from scratch. On this basis, they devise three methods for efficient contribution index estimation. Evaluations show that their methods accurately approximate the contribution index while notably accelerating its calculation.
Bibtex
@inbook{wei2020efficient,
title={Efficient and fair data valuation for horizontal federated learning},
author={Wei, Shuyue and Tong, Yongxin and Zhou, Zimu and Song, Tianshu},
year={2020},
booktitle={Federated Learning: Privacy and Incentive},
pages={139--152},
publisher={Springer}
}
Improving Fairness for Data Valuation in Horizontal Federated LearningZhenan Fan, Huang Fang, Zirui Zhou, Jian Pei, Michael P. Friedlander, Changxin Liu, Yong Zhang2020
SummaryFederated learning is an emerging decentralized machine learning scheme that allows multiple data owners to work collaboratively while ensuring data privacy. This paper focuses on fairness in data valuation within federated learning. The authors propose a new measure called completed federated Shapley value to improve the fairness of federated Shapley value. This approach leverages the concepts and tools from optimization and provides both theoretical analysis and empirical evaluation to verify the improvement in fairness.
Bibtex
@misc{fan2020improving,
title={Improving Fairness for Data Valuation in Horizontal Federated Learning},
author={Zhenan Fan and Huang Fang and Zirui Zhou and Jian Pei and Michael P. Friedlander and Changxin Liu and Yong Zhang},
year={2020},
eprint={2109.09046},
archivePrefix={arXiv},
primaryClass={cs.LG}
}
Data Valuation for Vertical Federated Learning: An Information-Theoretic ApproachXiao Han, Leye Wang, Junjie Wu2021
SummaryFederated learning (FL) is a machine learning paradigm that enables privacy-preserving cross-party data collaboration. This work introduces "FedValue," the first privacy-preserving, task-specific, model-free data valuation method for vertical FL tasks. It incorporates Shapley-CMI, an information-theoretic metric, for assessing data values from a game-theoretic perspective. The paper also proposes a novel server-aided federated computation mechanism and techniques to accelerate Shapley-CMI computation. Extensive experiments demonstrate the effectiveness and efficiency of FedValue.
Bibtex
@misc{han2021datavaluation,
title={Data Valuation for Vertical Federated Learning: An Information-Theoretic Approach},
author={Xiao Han and Leye Wang and Junjie Wu},
year={2021},
eprint={URL or DOI link TBD},
}
Towards More Efficient Data Valuation in Healthcare Federated Learning Using EnsemblingSourav Kumar, A. Lakshminarayanan, Ken Chang, Feri Guretno, Ivan Ho Mien, Jayashree Kalpathy-Cramer, Pavitra Krishnaswamy, Praveer Singh2021
SummaryThis paper addresses the challenge of data valuation in federated learning within healthcare. The authors propose a method called SaFE (Shapley Value for Federated Learning using Ensembling), which is designed to be efficient in settings where the number of contributing institutions is manageable. SaFE approximates the Shapley value using gradients from training a single model and develops methods for efficient contribution index estimation. This approach is particularly relevant in medical imaging where data heterogeneity is common and fast, accurate data valuation is necessary for multi-institutional collaborations.
Bibtex
@article{Kumar2021TowardsME,
title={Towards More Efficient Data Valuation in Healthcare Federated Learning Using Ensembling},
author={Sourav Kumar and A. Lakshminarayanan and Ken Chang and Feri Guretno and Ivan Ho Mien and Jayashree Kalpathy-Cramer and Pavitra Krishnaswamy and Praveer Singh},
journal={ArXiv},
year={2021},
volume={abs/2209.05424}
}
Data Debugging with Shapley Importance over Machine Learning PipelinesBojan Karlaš, David Dao, Matteo Interlandi, Sebastian Schelter, Wentao Wu, Ce Zhang2024
SummaryThis paper focuses on repairing datasets with the goal of improving the quality of end-to-end machine learning pipelines. Data repairs are prioritized by Shapley value. The authors propose methods for efficiently computing the Shapley value for different types of pipelines and empirically demonstrate the effectiveness of this approach.
Bibtex
@inproceedings{
karlas2024data,
title={Data Debugging with Shapley Importance over Machine Learning Pipelines},
author={Bojan Karla{\v{s}} and David Dao and Matteo Interlandi and Sebastian Schelter and Wentao Wu and Ce Zhang},
booktitle={The Twelfth International Conference on Learning Representations},
year={2024},
url={https://openreview.net/forum?id=qxGXjWxabq}
}
💻🎥

Data markets and society

Economics of Data

Nonrivalry and the Economics of DataCharles I. Jones, Christopher Tonetti2019
Bibtex
@article{10.1257/aer.20191330,
Author = {Jones, Charles I. and Tonetti, Christopher},
Title = {Nonrivalry and the Economics of Data},
Journal = {American Economic Review},
Volume = {110},
Number = {9},
Year = {2020},
Month = {September},
Pages = {2819-58},
DOI = {10.1257/aer.20191330},
URL = {https://www.aeaweb.org/articles?id=10.1257/aer.20191330}
}

Data Dignity

Chapter 5: Data as Labor, Radical MarketsEric A. Posner and E Glen Weyl2019
Bibtex
@book{posner2019radical,
title={Radical Markets},
author={Posner, Eric A and Weyl, E Glen},
year={2019},
publisher={Princeton University Press}
}
Should We Treat Data as Labor? Moving beyond "Free"Imanol Arrieta-Ibarra, Leonard Goff, Diego Jiménez-Hernández, Jaron Lanier, E. Glen Weyl2018
Bibtex
@article{10.1257/pandp.20181003,
Author = {Arrieta-Ibarra, Imanol and Goff, Leonard and Jiménez-Hernández, Diego and Lanier, Jaron and Weyl, E. Glen},
Title = {Should We Treat Data as Labor? Moving beyond "Free"},
Journal = {AEA Papers and Proceedings},
Volume = {108},
Year = {2018},
Month = {May},
Pages = {38-42},
DOI = {10.1257/pandp.20181003},
URL = {https://www.aeaweb.org/articles?id=10.1257/pandp.20181003}
}

Strategic adaptation

Performative prediction

Performative PredictionJuan Perdomo, Tijana Zrnic, Celestine Mendler-Dünner, Moritz Hardt2020
Summary Perdomo et al. (2020) introduce the concept of "performative prediction" dealing with predictions that influence the target they aim to predict, e.g. through taking actions based on the predictions, causing a distribution shift. The authors develop a risk minimization framework for performative prediction and introduce the equilibrium notion of performative stability where predictions are calibrated against future outcomes that manifest from acting on the prediction.
Bibtex
@inproceedings{perdomo2020performative,
title={Performative prediction},
author={Perdomo, Juan and Zrnic, Tijana and Mendler-D{"u}nner, Celestine and Hardt, Moritz},
booktitle={International Conference on Machine Learning},
pages={7599--7609},
year={2020},
organization={PMLR}
}
Stochastic Optimization for Performative PredictionCelestine Mendler-Dünner, Juan Perdomo, Tijana Zrnic, Moritz Hardt2020
Summary Mendler-Dünner et al. (2020) look at stochastic optimization for performative prediction and prove convergence rates for greedily deploying models after each stochastic update (which may cause distribution shift affecting convergence to a stability point) or lazily deploying the model after several updates.
Bibtex
@article{mendler2020stochastic,
title={Stochastic optimization for performative prediction},
author={Mendler-D{"u}nner, Celestine and Perdomo, Juan and Zrnic, Tijana and Hardt, Moritz},
journal={Advances in Neural Information Processing Systems},
volume={33},
pages={4929--4939},
year={2020}
}

Strategic classification

Strategic Classification is Causal Modeling in DisguiseJohn Miller, Smitha Milli, Moritz Hardt2020
Summary Miller et al. (2020) argue that strategic classication involves causal modelling and designing incentives for improvement requires solving a non-trivial causal inference problem. The authors provide a distinction between gaming and improvement as well as provide a causal framework for strategic adaptation.
Bibtex
@inproceedings{miller2020strategic,
title={Strategic classification is causal modeling in disguise},
author={Miller, John and Milli, Smitha and Hardt, Moritz},
booktitle={International Conference on Machine Learning},
pages={6917--6926},
year={2020},
organization={PMLR}
}
Alternative Microfoundations for Strategic ClassificationMeena Jagadeesan, Celestine Mendler-Dünner, Moritz Hardt2021
Summary Jagadeesan et al. (2021) show that standard microfoundations in strategic classification, that typically uses individual-level behaviour to deduce aggregate-level responses, can lead to degenerate behaviour in aggregate: discontinuities in the aggregate response, stable points ceasing to exist, and maximizing social burden. The authors introduce a noisy response model inspired by performative prediction that mitigates these limitations for binary classification.
Bibtex
@inproceedings{jagadeesan2021alternative,
title={Alternative microfoundations for strategic classification},
author={Jagadeesan, Meena and Mendler-D{"u}nner, Celestine and Hardt, Moritz},
booktitle={International Conference on Machine Learning},
pages={4687--4697},
year={2021},
organization={PMLR}
}

Data Valuation Researchers

NameInstituteh-index
Costas SpanosUniversity of California, Berkeley61
Jinsung YoonGoogle Cloud AI33
Tomas PfisterGoogle Cloud AI39
Amirata GhorbaniStanford18
James ZouStanford64
Nektaria TryfonaVirginia Tech27
Rachael Hwee Ling SimNational University of Singapore4
Bryan Kian Hsiang LowNational University of Singapore38
Dawn SongUniversity of California, Berkeley142
Zhaoxuan WuNational University of Singapore4
Xinyi XuNational University of Singapore8
Tianhao WangUniversity of Virginia18
José González CabañasUC3M-Santander Big Data Institute7
Ruben Cuevas RuminUniversidad Carlos III de Madrid26
Jiachen T. WangPrinceton University9
Bohong WangTsinghua University6
Yongchan KwonColumbia University10
Siyi TangArtera8
Li XiongEmory University52
Jessica VitakUniversity of Maryland49
Katie Chamberlain KritikosUniversity of Illinois at Urbana-Champaign6
Zhenan FanHuawei Technologies Canada6
Shuyue WeiBeihang University4
Hannah SteinSaarland University3
Wolfgang MaassSaarland University26
Mohammad Mohammadi AmiriRensselaer Polytechnic Institute18
Ramesh RaskarMIT103
Konstantin D. PandlKarlsruhe Institute of Technolgoy6
Ali SunyaevKarlsruhe Institute of Technolgoy43
Ludovico BorattoUniversity of Cagliari25
han xiao70
Junjie WuCenter for High Pressure Science & Technology Advanced Research55
Xiao TianNational University of Singapore1
Kean BirchInstitute for Technoscience & Society40
Callum WardUppsala University10
Praveer SinghUniversity of Colorado School of Medicine19
Anran XuShanghai Jiao Tong University2
Guihai Chen67
Andre EstevaCo-Founder & CEO, Artera23
Prateek MittalPrinceton University55
Hyeontaek OhInstitute for IT Convergence9
Lingjiao ChenStanford13
Xiangyu ChangXi'an Jiaotong University17
Hoang Anh JustVirginia Tech3
David DaoETH13
Mark MazumderHarvard12
Vijay Janapa ReddiHarvard46
Sabri EyubogluStanford6
Wenqian LiNational University of Singapore2
Bojan KarlašHarvard14
ai
data-valuation
market
ml