shaoshuo-ss/Awesome-LLM-Fingerprinting

Paper list of LLM fingerprinting, based on our paper titled "SoK: Large Language Model Copyright Auditing via Fingerprinting".

33

8 commits

updated Aug 28, 2025

See the code

README

Awesome-LLM-Fingerprinting

This repository provides a collection of papers about LLM fingerprinting and is based on the paper entitled "SoK: Large Language Model Copyright Auditing via Fingerprinting".

Definition of LLM Fingerprinting

LLM fingerprinting is a passive approach that does not require modifying the model. Instead, it non-intrusively extracts a set of inherent yet distinctive characteristics that collectively serve as the model's "fingerprint". The core idea is that any model derived from the source model $M_o$ will preserve a statistically significant portion of this fingerprint.

In this paper and repository, we follow a classic narrow definition of LLM watermarking and fingerprinting. The primary difference between watermarking and fingerprinting is whether the method requires modifying the model's parameters and LLM fingerprinting is a non-intrusive method. Some existing papers embed "fingerprints" into LLMs by modifying their parameters. In this paper and repository, we classify these methods as LLM watermarking.

The primary advantage of LLM fingerprinting lies in its non-intrusive nature. Since it does not alter the model, it introduces no performance degradation. The computational overhead is typically confined, which is often less intensive than the fine-tuning required for watermarking. Consequently, fingerprinting offers a more flexible and widely applicable paradigm for copyright auditing, especially for models that are already in the public domain.

LLM Fingerprinting Benchmark

In this SoK, we also provide a comprehensive benchmark named LeaFBench for evaluating LLM fingerprinting methods.

Citation

If you find this repository useful, please cite our paper:

@article{shao2025sok,
    title={SoK: Large Language Model Copyright Auditing via Fingerprinting},
    author={Shao, Shuo and Li, Yiming and He, Yu and Yao, Hongwei and Yang, Wenyuan and Tao, Dacheng and Qin, Zhan},
    journal={arXiv preprint arXiv:2508.19843},
    year={2025}
}

List of Papers

TitleConference/JournalYearTypeSubtypeQuery DataRelied FeaturesFingerprint ComparisonCode
A DNN Fingerprint for Non-Repudiable Model Ownership Identification and Piracy DetectionTIFS2022White-boxStaticNALower Layer WeightsAdjusted Cosine Similarity/
HuRef: HUman-REadable Fingerprint for Large Language ModelsNeurIPS2024White-boxStaticNAParameters' DirectionCosine SimilarityLink
EasyDetector: Using Linear Probe to Detect the Provenance of Large Language ModelsTrustCom2024White-boxForward-passExisting DatasetIntermediate FeaturesLinear Probe/
Intrinsic Fingerprint of LLMs: Continue Training is NOT All You Need to Steal A Model!arXiv2025White-boxStaticNAParameters' StatisticsCorrelation Coefficient/
Matrix-driven instant review: Confident detection and reconstruction of LLM plagiarism on PCarXiv2025White-boxStaticNAModel Weight MatricesMatrix Transformations based on Large Deviation Theory/
REEF: Representation Encoding Fingerprints for Large Language ModelsICLR2025White-boxForward-passExisting DatasetIntermediate FeaturesCentered Kernel AlignmentLink
Gradient-Based Model Fingerprinting for LLM Similarity Detection and Family ClassificationarXiv2025White-boxBackward-passUnknownGradientEuclidean Distance/
A Fingerprint for Large Language ModelsarXiv2024Black-boxUntargetedRandom QueriesOutput LogitsEuclidean Distance & Dimension DifferenceLink
Hide and Seek: Fingerprinting Large Language Models with Evolutionary LearningarXiv2024Black-boxUntargetedLLM-generated PromptsOutput ContentDetective LLMLink
TRAP: Targeted Random Adversarial Prompt Honeypot for Black-Box IdentificationACL Findings2024Black-boxTargetedA Base Prompt Combined with an Optimized SuffixOutput ContentExact MatchLink
ProFLingo: A Fingerprinting-based Intellectual Property Protection Scheme for Large Language ModelsIEEE CNS2024Black-boxTargetedA Prompt with an Optimized PrefixOutput ContentTarget Response RateLink
LLMmap: Fingerprinting For Large Language ModelsUSENIX Security2025Black-boxUntargetedManually-crafted PromptsOutput ContentCosine SimilarityLink
Model Equality Testing: Which Model is this API Serving?ICLR2025Black-boxUntargetedManually-crafted PromptsHamming Distance KernelMaximum Mean DiscrepancyLink
Your Large Language Models are Leaving FingerprintsCOLING Workshop2025Black-boxUntargetedExisting DatasetN-gramML Classifier/
Invisible Traces: Using Hybrid Fingerprinting to identify underlying LLMs in GenAI AppsarXiv2025Black-boxUntargetedCrafted Strategic Prompts & Generic PromptsOutput ContentTransformer-based Classifier & LLM/
CoTSRF: Utilize Chain of Thought as Stealthy and Robust Fingerprint of Large Language ModelsarXiv2025Black-boxUntargetedReasoning Questions with CoT PromptsOutput ContentKL Divergence/
DuFFin: A Dual-Level Fingerprinting Framework for LLMs IP ProtectionarXiv2025Black-boxUntargetedA Combination of Trigger PromptsOutput ContentCosine Similarity & Hamming DistanceLink
Auditing Black-Box LLM APIs with a Rank-Based Uniformity TestarXiv2025Black-boxUntargetedNatural PromptsThe Log-rank of Response TokensCramér-von Mises Statistical Test/
Attacks and defenses against LLM fingerprintingarXiv2025Black-boxUntargetedAn RL-optimized Set of QueriesOutput ContentTransformer-based Classifier/
RAP-SM: Robust Adversarial Prompt via Shadow Models for Copyright Verification of Large Language ModelsarXiv2025Black-boxTargetedA Prompt Combined with an Optimized SuffixOutput ContentExact Match/
RoFL: Robust Fingerprinting of Language ModelsarXiv2025Black-boxTargetedOptimized Unlikely Token SequencesOutput ContentExact MatchLink
LLM-FIN: Large Language Models Fingerprinting Attack on Edge DevicesISQED2024Side-channel//Memory Usage PatternML-based/
LLMs Have Rhythm: Fingerprinting Large Language Models Using Inter-Token Times and Network Traffic AnalysisOJCOMS2025Side-channel//Inter-token TimesDL-based/

shaoshuo-ss/Awesome-LLM-Fingerprinting

Paper list of LLM fingerprinting, based on our paper titled "SoK: Large Language Model Copyright Auditing via Fingerprinting".

33

8 commits

updated Aug 28, 2025

See the code

README

Awesome-LLM-Fingerprinting

This repository provides a collection of papers about LLM fingerprinting and is based on the paper entitled "SoK: Large Language Model Copyright Auditing via Fingerprinting".

Definition of LLM Fingerprinting

LLM fingerprinting is a passive approach that does not require modifying the model. Instead, it non-intrusively extracts a set of inherent yet distinctive characteristics that collectively serve as the model's "fingerprint". The core idea is that any model derived from the source model $M_o$ will preserve a statistically significant portion of this fingerprint.

In this paper and repository, we follow a classic narrow definition of LLM watermarking and fingerprinting. The primary difference between watermarking and fingerprinting is whether the method requires modifying the model's parameters and LLM fingerprinting is a non-intrusive method. Some existing papers embed "fingerprints" into LLMs by modifying their parameters. In this paper and repository, we classify these methods as LLM watermarking.

The primary advantage of LLM fingerprinting lies in its non-intrusive nature. Since it does not alter the model, it introduces no performance degradation. The computational overhead is typically confined, which is often less intensive than the fine-tuning required for watermarking. Consequently, fingerprinting offers a more flexible and widely applicable paradigm for copyright auditing, especially for models that are already in the public domain.

LLM Fingerprinting Benchmark

In this SoK, we also provide a comprehensive benchmark named LeaFBench for evaluating LLM fingerprinting methods.

Citation

If you find this repository useful, please cite our paper:

@article{shao2025sok,
    title={SoK: Large Language Model Copyright Auditing via Fingerprinting},
    author={Shao, Shuo and Li, Yiming and He, Yu and Yao, Hongwei and Yang, Wenyuan and Tao, Dacheng and Qin, Zhan},
    journal={arXiv preprint arXiv:2508.19843},
    year={2025}
}

List of Papers

TitleConference/JournalYearTypeSubtypeQuery DataRelied FeaturesFingerprint ComparisonCode
A DNN Fingerprint for Non-Repudiable Model Ownership Identification and Piracy DetectionTIFS2022White-boxStaticNALower Layer WeightsAdjusted Cosine Similarity/
HuRef: HUman-REadable Fingerprint for Large Language ModelsNeurIPS2024White-boxStaticNAParameters' DirectionCosine SimilarityLink
EasyDetector: Using Linear Probe to Detect the Provenance of Large Language ModelsTrustCom2024White-boxForward-passExisting DatasetIntermediate FeaturesLinear Probe/
Intrinsic Fingerprint of LLMs: Continue Training is NOT All You Need to Steal A Model!arXiv2025White-boxStaticNAParameters' StatisticsCorrelation Coefficient/
Matrix-driven instant review: Confident detection and reconstruction of LLM plagiarism on PCarXiv2025White-boxStaticNAModel Weight MatricesMatrix Transformations based on Large Deviation Theory/
REEF: Representation Encoding Fingerprints for Large Language ModelsICLR2025White-boxForward-passExisting DatasetIntermediate FeaturesCentered Kernel AlignmentLink
Gradient-Based Model Fingerprinting for LLM Similarity Detection and Family ClassificationarXiv2025White-boxBackward-passUnknownGradientEuclidean Distance/
A Fingerprint for Large Language ModelsarXiv2024Black-boxUntargetedRandom QueriesOutput LogitsEuclidean Distance & Dimension DifferenceLink
Hide and Seek: Fingerprinting Large Language Models with Evolutionary LearningarXiv2024Black-boxUntargetedLLM-generated PromptsOutput ContentDetective LLMLink
TRAP: Targeted Random Adversarial Prompt Honeypot for Black-Box IdentificationACL Findings2024Black-boxTargetedA Base Prompt Combined with an Optimized SuffixOutput ContentExact MatchLink
ProFLingo: A Fingerprinting-based Intellectual Property Protection Scheme for Large Language ModelsIEEE CNS2024Black-boxTargetedA Prompt with an Optimized PrefixOutput ContentTarget Response RateLink
LLMmap: Fingerprinting For Large Language ModelsUSENIX Security2025Black-boxUntargetedManually-crafted PromptsOutput ContentCosine SimilarityLink
Model Equality Testing: Which Model is this API Serving?ICLR2025Black-boxUntargetedManually-crafted PromptsHamming Distance KernelMaximum Mean DiscrepancyLink
Your Large Language Models are Leaving FingerprintsCOLING Workshop2025Black-boxUntargetedExisting DatasetN-gramML Classifier/
Invisible Traces: Using Hybrid Fingerprinting to identify underlying LLMs in GenAI AppsarXiv2025Black-boxUntargetedCrafted Strategic Prompts & Generic PromptsOutput ContentTransformer-based Classifier & LLM/
CoTSRF: Utilize Chain of Thought as Stealthy and Robust Fingerprint of Large Language ModelsarXiv2025Black-boxUntargetedReasoning Questions with CoT PromptsOutput ContentKL Divergence/
DuFFin: A Dual-Level Fingerprinting Framework for LLMs IP ProtectionarXiv2025Black-boxUntargetedA Combination of Trigger PromptsOutput ContentCosine Similarity & Hamming DistanceLink
Auditing Black-Box LLM APIs with a Rank-Based Uniformity TestarXiv2025Black-boxUntargetedNatural PromptsThe Log-rank of Response TokensCramér-von Mises Statistical Test/
Attacks and defenses against LLM fingerprintingarXiv2025Black-boxUntargetedAn RL-optimized Set of QueriesOutput ContentTransformer-based Classifier/
RAP-SM: Robust Adversarial Prompt via Shadow Models for Copyright Verification of Large Language ModelsarXiv2025Black-boxTargetedA Prompt Combined with an Optimized SuffixOutput ContentExact Match/
RoFL: Robust Fingerprinting of Language ModelsarXiv2025Black-boxTargetedOptimized Unlikely Token SequencesOutput ContentExact MatchLink
LLM-FIN: Large Language Models Fingerprinting Attack on Edge DevicesISQED2024Side-channel//Memory Usage PatternML-based/
LLMs Have Rhythm: Fingerprinting Large Language Models Using Inter-Token Times and Network Traffic AnalysisOJCOMS2025Side-channel//Inter-token TimesDL-based/