This repository contains the code and data for the paper :
| Label | Name | Description |
|---|---|---|
| LV | Lexical Vagueness | Ambiguous or imprecise word choices that obscure the intended behavior |
| SF | Style/Formatting | Changes to surface form (e.g., spacing, phrasing) without semantic change |
| US | Underspecification | Missing constraints, edge cases, or output requirements |
| CLEAN | Original | Unmodified benchmark prompt |
Mutations are applied to three benchmarks: HumanEval, MBPP, and LiveCodeBench.
Python 3.10+ and CUDA-capable GPU(s) are required for local model experiments.
pip install -r requirements.txtCopy .env.example to .env and fill in your keys:
cp .env.example .env
# then edit .env:
# OPENAI_API_KEY=sk-...
# ANTHROPIC_API_KEY=sk-ant-...All API-based scripts load .env automatically at startup (load_env.py), so keys don't need to be exported into the shell. .env is gitignored — never commit it.
.
├── datasets/ # Base benchmark datasets
│ ├── humanEval/ # HumanEval original + 3 mutation variants
│ ├── mbpp/ # MBPP original + 3 mutation variants
│ ├── livecodebench/ # LiveCodeBench original + 3 mutation
│
├── mutations/ # Mutated datasets with test cases appended
│ # (used as input to inference scripts)
│
├── data_split.py # Shared problem-disjoint train/val split
├── load_env.py # .env loader used by all API-based scripts
│
├── inference_results/ # All classifier results live here
│ ├── qwen_linear_classifier_outputs/ # Frozen-encoder + linear head results (seed42/, seed123456/)
│ ├── qwen_lora_classifier_outputs/ # LoRA classifier results (seed42/, seed123456/)
│ ├── qwen_full_classifier_outputs/ # Full fine-tune classifier results (seed42/, seed123456/)
│ └── api_classifier_results/ # Prompting-based classifier results
│ ├── zero_shot/
│ │ ├── claude/ # Claude zero-shot classifier results
│ │ └── gpt/ # GPT-5-mini zero-shot classifier results
│ └── few_shot/
│ ├── claude/ # Claude few-shot classifier results
│ └── gpt/ # GPT-5-mini few-shot classifier results
├── figures/ # Generated plots and visualizations
│
└── judge_output/ # Mutation-judge outputs