This repository contains the official implementation and experimental artifacts for “Prompt Sensitivity or Evaluation Artifact? A Task-Aware Analysis for Large Language Models.”
Prompt sensitivity is commonly interpreted as an intrinsic robustness weakness of large language models. This work asks a different question: how much of the measured sensitivity is caused by the model, and how much is introduced by the evaluation method?
We introduce Evaluation-Attributable Sensitivity (EAS) and Signed EAS, instance-level diagnostics that compare sensitivity measured using task-specific heuristic metrics with sensitivity measured by a semantic LLM judge.
The current implementation, final analysis code, and updated outputs are located in
study/. Other top-level folders are retained as earlier experimental artifacts.
Each benchmark item is evaluated under a controlled (2^3) factorial prompt design with three binary structural factors:
This produces 8 prompt variants per item. Model responses are scored using both a task-specific heuristic and a held-out LLM judge. Their disagreement is then analysed through EAS, Signed EAS, a four-zone taxonomy, and structural-factor effects.
Evaluation-Attributable Sensitivity (EAS)
Quantifies the magnitude of disagreement between heuristic-based and judge-based prompt-sensitivity estimates.
Signed EAS
Identifies the direction of disagreement:
Four-zone diagnostic taxonomy
Classifies each instance as Artifact, Underdetected, Genuine, or Stable.
Cross-family evaluation
Evaluates 9 instruction-tuned models from 5 model families on 3 task formats.
| Component | Configuration |
|---|---|
| Datasets | ARC-Challenge, BoolQ, SQuAD |
| Items | 200 per dataset |
| Prompt variants | 8 per item |
| Prompt factors | Role, format, answer prefix |
| Evaluated models | 9 models from Llama, Qwen, Mistral, Gemma, and OpenAI |
| Decoding | Greedy decoding, temperature 0, 64-token budget |
| Semantic judge | Held-out Claude Haiku 4.5 |
| Total response-judge evaluations | 43,200 |
| Confidence intervals | 1,000 bootstrap resamples |
For an item $i$, let the heuristic and judge scores across the eight prompt templates be $h_{i,t}$ and $j_{i,t}$, respectively.
$$ \mathrm{SensH}i = \sigma_t\left(h{i,t}\right) $$
$$ \mathrm{SensJ}i = \sigma_t\left(j{i,t}\right) $$
$$ \mathrm{EAS}_i = \left| \mathrm{SensH}_i - \mathrm{SensJ}_i \right| $$
$$ \mathrm{SignedEAS}_i = \mathrm{SensH}_i - \mathrm{SensJ}_i $$
The heuristic evaluation is task-aware:
The semantic evaluation uses a held-out LLM judge that returns a binary correctness verdict.
The current implementation uses the 75th-percentile thresholds of SensH and SensJ to separate low- and high-sensitivity instances.
| Zone | Heuristic sensitivity | Judge sensitivity | Interpretation |
|---|---|---|---|
| Stable | Low | Low | Both evaluations indicate stable behaviour |
| Artifact | High | Low | The heuristic introduces or inflates apparent sensitivity |
| Underdetected | Low | High | The heuristic misses sensitivity detected semantically |
| Genuine | High | High | Both evaluations indicate real prompt sensitivity |
For each prompt factor $k$, the code computes its main effect as the difference between the mean score when that factor is enabled and disabled:
$$ \Delta_k^e = \mathbb{E}[e \mid k = 1] - \mathbb{E}[e \mid k = 0] $$
where $e$ is either the heuristic or judge score. Comparing $\Delta_k^h$ and $\Delta_k^j$ distinguishes semantic improvements from changes that merely make answers easier for a heuristic to parse.
.
├── README.md
├── requirements.txt
├── figures/
│ └── eas_workflow.png
└── study/
├── templates.py # 2^3 structural prompt design
├── prepare.py # Load datasets and create prompts.jsonl
├── infer.py # Ollama, Hugging Face, OpenAI, and Anthropic inference
├── judge.py # Held-out LLM-as-Judge evaluation
├── metrics.py # SensH, SensJ, EAS, Signed EAS, zones, and factor effects
├── plot.py # Publication-quality result figures
├── run.sh # Convenience script for a single model
├── run_all.sh # Lightweight multi-model example
└── output/
├── <model>/ # Per-model responses, judgements, and metrics
└── combined_final/
├── metrics_instance.csv
├── metrics_dataset.csv
└── figures/
git clone https://github.com/sayumimuthu/llm-prompt-sensitivity-evaluation-artifact.git
cd llm-prompt-sensitivity-evaluation-artifact
python -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
pip install -r requirements.txt
Install the packages required by the inference backend you plan to use:
# Hugging Face local models
pip install torch transformers accelerate
# OpenAI API
pip install openai
# Anthropic API
pip install anthropic
study/infer.py and study/judge.py read environment variables from study/.env or att1/.env.
# Add only the variables required by your selected backends.
HF_TOKEN=your_huggingface_token
OPENAI_API_KEY=your_openai_api_key
ANTHROPIC_API_KEY=your_anthropic_api_key
OLLAMA_BASE_URL=ollama_installed_server_port [Ex:http://localhost:11434]
Do not commit .env files or API keys.
python study/prepare.py \
--n 200 \
--seed 42 \
--out-dir study/output
This creates study/output/prompts.jsonl with:
[ 200\ \text{items} \times 3\ \text{datasets} \times 8\ \text{templates} = 4,800\ \text{prompts}. ]
Example using Ollama:
python study/infer.py \
--in-file study/output/prompts.jsonl \
--out-file study/output/llama3.1-8b/responses.jsonl \
--backend ollama \
--model llama3.1:8b \
--temperature 0 \
--max-tokens 64
Other supported backends are hf_local, openai, and anthropic.
python study/judge.py \
--in-file study/output/llama3.1-8b/responses.jsonl \
--out-file study/output/llama3.1-8b/judged.jsonl \
--judge-backend anthropic \
--judge-model claude-haiku-4-5
The published experiment uses a held-out judge to avoid self-evaluation bias.
For one model:
python study/metrics.py \
--in-files study/output/llama3.1-8b/judged.jsonl \
--out-instance study/output/llama3.1-8b/metrics_instance.csv \
--out-dataset study/output/llama3.1-8b/metrics_dataset.csv
For a combined analysis, provide all judged files after --in-files.
python study/plot.py \
--in-instance study/output/combined_final/metrics_instance.csv \
--in-dataset study/output/combined_final/metrics_dataset.csv \
--out-dir study/output/combined_final/figures
The script generates:
fig1_trizone.pngfig2_scatter.pngfig3_ablation.pngfig4_eas_by_task.pngfig5_eas_ci.pngfig6_signed_eas.pngThe final aggregate outputs used for the repository visualisations are available in:
study/output/combined_final/
metrics_instance.csv contains one row per item-model pair.metrics_dataset.csv contains dataset-level aggregates and bootstrap confidence intervals.figures/ contains the generated result plots.The individual model directories contain the corresponding raw responses, judge verdicts, and intermediate outputs.
University of Peradeniya, Sri Lanka.
This research was funded by the University Research Council (URC), University of Peradeniya, under Grant No. 32.
Python
94.2%
Shell
5.8%
This repository contains the official implementation and experimental artifacts for “Prompt Sensitivity or Evaluation Artifact? A Task-Aware Analysis for Large Language Models.”
Prompt sensitivity is commonly interpreted as an intrinsic robustness weakness of large language models. This work asks a different question: how much of the measured sensitivity is caused by the model, and how much is introduced by the evaluation method?
We introduce Evaluation-Attributable Sensitivity (EAS) and Signed EAS, instance-level diagnostics that compare sensitivity measured using task-specific heuristic metrics with sensitivity measured by a semantic LLM judge.
The current implementation, final analysis code, and updated outputs are located in
study/. Other top-level folders are retained as earlier experimental artifacts.
Each benchmark item is evaluated under a controlled (2^3) factorial prompt design with three binary structural factors:
This produces 8 prompt variants per item. Model responses are scored using both a task-specific heuristic and a held-out LLM judge. Their disagreement is then analysed through EAS, Signed EAS, a four-zone taxonomy, and structural-factor effects.
Evaluation-Attributable Sensitivity (EAS)
Quantifies the magnitude of disagreement between heuristic-based and judge-based prompt-sensitivity estimates.
Signed EAS
Identifies the direction of disagreement:
Four-zone diagnostic taxonomy
Classifies each instance as Artifact, Underdetected, Genuine, or Stable.
Cross-family evaluation
Evaluates 9 instruction-tuned models from 5 model families on 3 task formats.
| Component | Configuration |
|---|---|
| Datasets | ARC-Challenge, BoolQ, SQuAD |
| Items | 200 per dataset |
| Prompt variants | 8 per item |
| Prompt factors | Role, format, answer prefix |
| Evaluated models | 9 models from Llama, Qwen, Mistral, Gemma, and OpenAI |
| Decoding | Greedy decoding, temperature 0, 64-token budget |
| Semantic judge | Held-out Claude Haiku 4.5 |
| Total response-judge evaluations | 43,200 |
| Confidence intervals | 1,000 bootstrap resamples |
For an item $i$, let the heuristic and judge scores across the eight prompt templates be $h_{i,t}$ and $j_{i,t}$, respectively.
$$ \mathrm{SensH}i = \sigma_t\left(h{i,t}\right) $$
$$ \mathrm{SensJ}i = \sigma_t\left(j{i,t}\right) $$
$$ \mathrm{EAS}_i = \left| \mathrm{SensH}_i - \mathrm{SensJ}_i \right| $$
$$ \mathrm{SignedEAS}_i = \mathrm{SensH}_i - \mathrm{SensJ}_i $$
The heuristic evaluation is task-aware:
The semantic evaluation uses a held-out LLM judge that returns a binary correctness verdict.
The current implementation uses the 75th-percentile thresholds of SensH and SensJ to separate low- and high-sensitivity instances.
| Zone | Heuristic sensitivity | Judge sensitivity | Interpretation |
|---|---|---|---|
| Stable | Low | Low | Both evaluations indicate stable behaviour |
| Artifact | High | Low | The heuristic introduces or inflates apparent sensitivity |
| Underdetected | Low | High | The heuristic misses sensitivity detected semantically |
| Genuine | High | High | Both evaluations indicate real prompt sensitivity |
For each prompt factor $k$, the code computes its main effect as the difference between the mean score when that factor is enabled and disabled:
$$ \Delta_k^e = \mathbb{E}[e \mid k = 1] - \mathbb{E}[e \mid k = 0] $$
where $e$ is either the heuristic or judge score. Comparing $\Delta_k^h$ and $\Delta_k^j$ distinguishes semantic improvements from changes that merely make answers easier for a heuristic to parse.
.
├── README.md
├── requirements.txt
├── figures/
│ └── eas_workflow.png
└── study/
├── templates.py # 2^3 structural prompt design
├── prepare.py # Load datasets and create prompts.jsonl
├── infer.py # Ollama, Hugging Face, OpenAI, and Anthropic inference
├── judge.py # Held-out LLM-as-Judge evaluation
├── metrics.py # SensH, SensJ, EAS, Signed EAS, zones, and factor effects
├── plot.py # Publication-quality result figures
├── run.sh # Convenience script for a single model
├── run_all.sh # Lightweight multi-model example
└── output/
├── <model>/ # Per-model responses, judgements, and metrics
└── combined_final/
├── metrics_instance.csv
├── metrics_dataset.csv
└── figures/
git clone https://github.com/sayumimuthu/llm-prompt-sensitivity-evaluation-artifact.git
cd llm-prompt-sensitivity-evaluation-artifact
python -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
pip install -r requirements.txt
Install the packages required by the inference backend you plan to use:
# Hugging Face local models
pip install torch transformers accelerate
# OpenAI API
pip install openai
# Anthropic API
pip install anthropic
study/infer.py and study/judge.py read environment variables from study/.env or att1/.env.
# Add only the variables required by your selected backends.
HF_TOKEN=your_huggingface_token
OPENAI_API_KEY=your_openai_api_key
ANTHROPIC_API_KEY=your_anthropic_api_key
OLLAMA_BASE_URL=ollama_installed_server_port [Ex:http://localhost:11434]
Do not commit .env files or API keys.
python study/prepare.py \
--n 200 \
--seed 42 \
--out-dir study/output
This creates study/output/prompts.jsonl with:
[ 200\ \text{items} \times 3\ \text{datasets} \times 8\ \text{templates} = 4,800\ \text{prompts}. ]
Example using Ollama:
python study/infer.py \
--in-file study/output/prompts.jsonl \
--out-file study/output/llama3.1-8b/responses.jsonl \
--backend ollama \
--model llama3.1:8b \
--temperature 0 \
--max-tokens 64
Other supported backends are hf_local, openai, and anthropic.
python study/judge.py \
--in-file study/output/llama3.1-8b/responses.jsonl \
--out-file study/output/llama3.1-8b/judged.jsonl \
--judge-backend anthropic \
--judge-model claude-haiku-4-5
The published experiment uses a held-out judge to avoid self-evaluation bias.
For one model:
python study/metrics.py \
--in-files study/output/llama3.1-8b/judged.jsonl \
--out-instance study/output/llama3.1-8b/metrics_instance.csv \
--out-dataset study/output/llama3.1-8b/metrics_dataset.csv
For a combined analysis, provide all judged files after --in-files.
python study/plot.py \
--in-instance study/output/combined_final/metrics_instance.csv \
--in-dataset study/output/combined_final/metrics_dataset.csv \
--out-dir study/output/combined_final/figures
The script generates:
fig1_trizone.pngfig2_scatter.pngfig3_ablation.pngfig4_eas_by_task.pngfig5_eas_ci.pngfig6_signed_eas.pngThe final aggregate outputs used for the repository visualisations are available in:
study/output/combined_final/
metrics_instance.csv contains one row per item-model pair.metrics_dataset.csv contains dataset-level aggregates and bootstrap confidence intervals.figures/ contains the generated result plots.The individual model directories contain the corresponding raw responses, judge verdicts, and intermediate outputs.
University of Peradeniya, Sri Lanka.
This research was funded by the University Research Council (URC), University of Peradeniya, under Grant No. 32.
Python
94.2%
Shell
5.8%