Skip to content

Latest commit

 

History

29 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Prompt Defects in Code Generation Benchmarks

This repository contains the code and data for the paper :

Natural Language Defects in Code Generation Prompts: Empirical Impact and Lightweight Detection

Mutation Types

Label Name Description
LV Lexical Vagueness Ambiguous or imprecise word choices that obscure the intended behavior
SF Style/Formatting Changes to surface form (e.g., spacing, phrasing) without semantic change
US Underspecification Missing constraints, edge cases, or output requirements
CLEAN Original Unmodified benchmark prompt

Mutations are applied to three benchmarks: HumanEval, MBPP, and LiveCodeBench.


Setup

Requirements

Python 3.10+ and CUDA-capable GPU(s) are required for local model experiments.

pip install -r requirements.txt

API Keys

Copy .env.example to .env and fill in your keys:

cp .env.example .env
# then edit .env:
#   OPENAI_API_KEY=sk-...
#   ANTHROPIC_API_KEY=sk-ant-...

All API-based scripts load .env automatically at startup (load_env.py), so keys don't need to be exported into the shell. .env is gitignored — never commit it.


Repository Structure

.
├── datasets/                  # Base benchmark datasets
│   ├── humanEval/             # HumanEval original + 3 mutation variants
│   ├── mbpp/                  # MBPP original + 3 mutation variants
│   ├── livecodebench/         # LiveCodeBench original + 3 mutation 
│
├── mutations/                 # Mutated datasets with test cases appended
│                              # (used as input to inference scripts)
│
├── data_split.py               # Shared problem-disjoint train/val split 
├── load_env.py                 # .env loader used by all API-based scripts
│
├── inference_results/                    # All classifier results live here
│   ├── qwen_linear_classifier_outputs/   # Frozen-encoder + linear head results (seed42/, seed123456/)
│   ├── qwen_lora_classifier_outputs/     # LoRA classifier results (seed42/, seed123456/)
│   ├── qwen_full_classifier_outputs/     # Full fine-tune classifier results (seed42/, seed123456/)
│   └── api_classifier_results/           # Prompting-based classifier results
│       ├── zero_shot/
│       │   ├── claude/                    # Claude zero-shot classifier results
│       │   └── gpt/                       # GPT-5-mini zero-shot classifier results
│       └── few_shot/
│           ├── claude/                    # Claude few-shot classifier results
│           └── gpt/                       # GPT-5-mini few-shot classifier results
├── figures/                    # Generated plots and visualizations
│
└── judge_output/               # Mutation-judge outputs

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages