A curated collection of papers and related projects on using LLMs for privacy.
39
14 commits
updated Oct 8, 2025
This is a collection of papers (and projects) related to the use of LLMs for privacy that I curate for my research. The focus is on privacy applications powered by modern, decoder-only LLMs (GPT, Llama, etc.), not on training data privacy (e.g., membership inference attacks, data extraction), and not on pre-GPT LLMs like BERT.
All of these papers are from 2023 onward. I try to include paper links from the official publication venues whenever possible, falling back to OpenReview and then to arXiv. I also try to include the earliest date when a paper was first submitted to arXiv. I categorize the papers based on their goals/domains (e.g., detection vs anonymization) and further add the following tags:
CO (constructive) if they are building something new or on top of some other existing workEV (evaluation) if they are evaluating/benchmarking (usually there's a novel accompanying dataset)PO (position) if they are primarily presenting a position or opinionCI if they involve the Contextual Integrity theory.Please feel free to open an issue or a pull request if you would like to share an interesting paper that uses LLMs for privacy. I will try to read and update the list whenever I have the time.
Papers that focus on detecting/assessing privacy leakages/risks or privacy policy violations:
EVEVEV, CO, CIEVEVEVCO, CIEVEVCO, CIEVCO, CIEVEV(While NER is not necessarily about sensitive data detection, it's very closely related)
COCOCOPapers that study the use of LLMs for data anonymization, de-identification, sanitization, authorship obfuscation, etc. (further sub-categorized by the target evaluation domain):
EV, CIEVPO, CIEV, CIEVEV, CICOCOCOCOCO, CICOCOCOCOCOCOCOCOCOCOCOCOCICO, CICO, CIEV, CIEV, CICO, CICOCOCOCOCOCOEV14 commits
A curated collection of papers and related projects on using LLMs for privacy.
39
14 commits
updated Oct 8, 2025
This is a collection of papers (and projects) related to the use of LLMs for privacy that I curate for my research. The focus is on privacy applications powered by modern, decoder-only LLMs (GPT, Llama, etc.), not on training data privacy (e.g., membership inference attacks, data extraction), and not on pre-GPT LLMs like BERT.
All of these papers are from 2023 onward. I try to include paper links from the official publication venues whenever possible, falling back to OpenReview and then to arXiv. I also try to include the earliest date when a paper was first submitted to arXiv. I categorize the papers based on their goals/domains (e.g., detection vs anonymization) and further add the following tags:
CO (constructive) if they are building something new or on top of some other existing workEV (evaluation) if they are evaluating/benchmarking (usually there's a novel accompanying dataset)PO (position) if they are primarily presenting a position or opinionCI if they involve the Contextual Integrity theory.Please feel free to open an issue or a pull request if you would like to share an interesting paper that uses LLMs for privacy. I will try to read and update the list whenever I have the time.
Papers that focus on detecting/assessing privacy leakages/risks or privacy policy violations:
EVEVEV, CO, CIEVEVEVCO, CIEVEVCO, CIEVCO, CIEVEV(While NER is not necessarily about sensitive data detection, it's very closely related)
COCOCOPapers that study the use of LLMs for data anonymization, de-identification, sanitization, authorship obfuscation, etc. (further sub-categorized by the target evaluation domain):
EV, CIEVPO, CIEV, CIEVEV, CICOCOCOCOCO, CICOCOCOCOCOCOCOCOCOCOCOCOCICO, CICO, CIEV, CIEV, CICO, CICOCOCOCOCOCOEV14 commits