Code and data for the USENIX 2025 paper "We Have a Package for You! A Comprehensive Analysis of Package Hallucinations by Code Generating LLMs"
Python
34
11 commits
updated Aug 14, 2026
A Comprehensive Analysis of Package Hallucinations by Code-Generating LLMs
This repository contains the code, data, and instructions for reproducing the experiments and results from our paper:
We Have a Package for You! A Comprehensive Analysis of Package Hallucinations by Code Generating LLMs
Joseph Spracklen, Raveen Wijewickrama, A H M Nazmus Sakib, Anindya Maiti, Bimal Viswanath, Murtuza Jadliwala
In Proceedings of the USENIX Security Symposium, 2025.
π Paper PDF
Package hallucinations occur when an LLM generates code that references a non-existent package (e.g., via pip install xyz or npm install xyz where xyz does not exist).
This creates a software supply-chain risk: adversaries can upload a malicious package using that hallucinated name.
This repo provides:
.
βββ run_test.py # Runs a full hallucination detection experiment
βββ Models/ # Place tested models here (one default model included)
βββ Data/ # Prompt datasets & per-language resources
β βββ Python/
β β βββ LLM_AT.json
β β βββ LLM_LY.json
β β βββ SO_AT.json
β β βββ SO_LY.json
β βββ JavaScript/
β βββ LLM_AT.json
β βββ LLM_LY.json
β βββ SO_AT.json
β βββ SO_LY.json
βββ Tests/ # Output directory for experiment results (starts empty)
βββ Mitigation/ # Mitigation experiments
β βββ run_model_RAG.py
β βββ run_model_SD.py
β βββ run_model_combined.py
β βββ Data/ # RAG DB + build data
β βββ Fine_tuned/ # Fine-tuned & quantized models used in mitigation testing
β βββ RAG_setup.py # Builds the vector DB from Mitigation/Data
βββ Plots/ # Code and data to reproduce paper figures
βββ environment.yml # Conda environment
βββ requirements.txt # (Optional) pip dependencies
βββ README.md
git clone https://github.com/Spracks/PackageHallucination.git
cd PackageHallucination
Using Conda (recommended):
conda env create -f environment.yml
conda activate pkg-hallucination
Or using pip:
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txt
The environment listed is bloated, really just need PyTorch + transformers and the associated dependencies.
Run a full hallucination detection experiment for one model:
python run_test.py DeepSeek_1B --language Python
# or
python run_test.py DeepSeek_1B --language JavaScript
What it does
Data/<LANG>/Tests/Notes
Models/Included prompt datasets (per language):
LLM_AT.json β LLM-generated prompts based on all-time most popular packagesLLM_LY.json β LLM-generated prompts based on last-year most popular packagesSO_AT.json β Top Stack Overflow questions (all-time)SO_LY.json β Top Stack Overflow questions (last-year)Each language directory also contains the master list of valid package names used for detection.
We do not publish the master list of hallucinated package names or per-prompt detailed results (see Security & Ethics). Verified researchers can request access.
Reproduce main tables/figures by re-running experiments and then building plots:
# 1) Run experiments (example)
python run_test.py DeepSeek_1B --language Python
python run_test.py CodeLlama_7B --language Python
# ... (repeat for desired model/language combinations)
# 2) Build figures
cd Plots
python reproduce_figures.py
Where applicable, figure scripts read from Tests/ to regenerate the paper plots.
We provide three mitigation strategies:
RAG (Retrieval-Augmented Generation)
Augments prompts with retrieved package-context from a vector DB built from package descriptions.
# Build RAG DB (once)
python Mitigation/RAG_setup.py
# Run RAG experiment
python Mitigation/run_model_RAG.py DeepSeek_1B --language Python
Self-Detection / Self-Refinement
The model checks its own suggested package list; if invalid, regenerate with constraints.
python Mitigation/run_model_SD.py CodeLlama_7B --language Python
Fine-Tuning
Fine-tune on valid (non-hallucinated) package recommendations derived from the pipeline.
# Use the fine-tuned checkpoints under Mitigation/Fine_tuned/
python Mitigation/run_model_combined.py DeepSeek_1B --language Python
See paper for comparative results; fine-tuning produced the largest hallucination reduction.
environment.yml and keep decoding parameters consistent unless you are explicitly testing RQ2-style variations.environment.yml.Models/ and the name matches your CLI arg.If you use this repository, please cite:
@inproceedings{spracklen2025packagehallucination,
title = {We Have a Package for You! A Comprehensive Analysis of Package Hallucinations by Code Generating LLMs},
author = {Joseph Spracklen and Raveen Wijewickrama and A H M Nazmus Sakib and Anindya Maiti and Bimal Viswanath and Murtuza Jadliwala},
booktitle = {USENIX Security Symposium},
year = {2025}
}
This project is licensed under the MIT License.
Questions or collaboration:
Code and data for the USENIX 2025 paper "We Have a Package for You! A Comprehensive Analysis of Package Hallucinations by Code Generating LLMs"
Python
34
11 commits
updated Aug 14, 2026
A Comprehensive Analysis of Package Hallucinations by Code-Generating LLMs
This repository contains the code, data, and instructions for reproducing the experiments and results from our paper:
We Have a Package for You! A Comprehensive Analysis of Package Hallucinations by Code Generating LLMs
Joseph Spracklen, Raveen Wijewickrama, A H M Nazmus Sakib, Anindya Maiti, Bimal Viswanath, Murtuza Jadliwala
In Proceedings of the USENIX Security Symposium, 2025.
π Paper PDF
Package hallucinations occur when an LLM generates code that references a non-existent package (e.g., via pip install xyz or npm install xyz where xyz does not exist).
This creates a software supply-chain risk: adversaries can upload a malicious package using that hallucinated name.
This repo provides:
.
βββ run_test.py # Runs a full hallucination detection experiment
βββ Models/ # Place tested models here (one default model included)
βββ Data/ # Prompt datasets & per-language resources
β βββ Python/
β β βββ LLM_AT.json
β β βββ LLM_LY.json
β β βββ SO_AT.json
β β βββ SO_LY.json
β βββ JavaScript/
β βββ LLM_AT.json
β βββ LLM_LY.json
β βββ SO_AT.json
β βββ SO_LY.json
βββ Tests/ # Output directory for experiment results (starts empty)
βββ Mitigation/ # Mitigation experiments
β βββ run_model_RAG.py
β βββ run_model_SD.py
β βββ run_model_combined.py
β βββ Data/ # RAG DB + build data
β βββ Fine_tuned/ # Fine-tuned & quantized models used in mitigation testing
β βββ RAG_setup.py # Builds the vector DB from Mitigation/Data
βββ Plots/ # Code and data to reproduce paper figures
βββ environment.yml # Conda environment
βββ requirements.txt # (Optional) pip dependencies
βββ README.md
git clone https://github.com/Spracks/PackageHallucination.git
cd PackageHallucination
Using Conda (recommended):
conda env create -f environment.yml
conda activate pkg-hallucination
Or using pip:
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txt
The environment listed is bloated, really just need PyTorch + transformers and the associated dependencies.
Run a full hallucination detection experiment for one model:
python run_test.py DeepSeek_1B --language Python
# or
python run_test.py DeepSeek_1B --language JavaScript
What it does
Data/<LANG>/Tests/Notes
Models/Included prompt datasets (per language):
LLM_AT.json β LLM-generated prompts based on all-time most popular packagesLLM_LY.json β LLM-generated prompts based on last-year most popular packagesSO_AT.json β Top Stack Overflow questions (all-time)SO_LY.json β Top Stack Overflow questions (last-year)Each language directory also contains the master list of valid package names used for detection.
We do not publish the master list of hallucinated package names or per-prompt detailed results (see Security & Ethics). Verified researchers can request access.
Reproduce main tables/figures by re-running experiments and then building plots:
# 1) Run experiments (example)
python run_test.py DeepSeek_1B --language Python
python run_test.py CodeLlama_7B --language Python
# ... (repeat for desired model/language combinations)
# 2) Build figures
cd Plots
python reproduce_figures.py
Where applicable, figure scripts read from Tests/ to regenerate the paper plots.
We provide three mitigation strategies:
RAG (Retrieval-Augmented Generation)
Augments prompts with retrieved package-context from a vector DB built from package descriptions.
# Build RAG DB (once)
python Mitigation/RAG_setup.py
# Run RAG experiment
python Mitigation/run_model_RAG.py DeepSeek_1B --language Python
Self-Detection / Self-Refinement
The model checks its own suggested package list; if invalid, regenerate with constraints.
python Mitigation/run_model_SD.py CodeLlama_7B --language Python
Fine-Tuning
Fine-tune on valid (non-hallucinated) package recommendations derived from the pipeline.
# Use the fine-tuned checkpoints under Mitigation/Fine_tuned/
python Mitigation/run_model_combined.py DeepSeek_1B --language Python
See paper for comparative results; fine-tuning produced the largest hallucination reduction.
environment.yml and keep decoding parameters consistent unless you are explicitly testing RQ2-style variations.environment.yml.Models/ and the name matches your CLI arg.If you use this repository, please cite:
@inproceedings{spracklen2025packagehallucination,
title = {We Have a Package for You! A Comprehensive Analysis of Package Hallucinations by Code Generating LLMs},
author = {Joseph Spracklen and Raveen Wijewickrama and A H M Nazmus Sakib and Anindya Maiti and Bimal Viswanath and Murtuza Jadliwala},
booktitle = {USENIX Security Symposium},
year = {2025}
}
This project is licensed under the MIT License.
Questions or collaboration: