This project implements an adversarial testing framework for evaluating the robustness of large language models (LLMs). It utilizes methods like Monte Carlo Tree Search (MCTS) with Double Progressive Widening (DPW) and adversarial suffix generation to test and optimize prompts against predefined ethical and moderation constraints.
- Adversarial Testing: Evaluate LLM responses against adversarially crafted prompts.
- Monte Carlo Tree Search (MCTS): Includes a DPW solver for efficient exploration of adversarial space.
- Support for Multiple Models: Works with models such as GPT, Vicuna, and LLaMa.
- Moderation API: Integrates OpenAI's Moderation API for content filtering.
- Visualization: Includes tools for loss tracking and visualization.
KOV.py: The main script managing the adversarial testing process and framework setup.WhiteBox.py: Implements the white-box optimization methods, including token sampling and adversarial suffix generation.DPW_Solver.py: Contains the DPW-based MCTS solver for optimizing adversarial search.utils.py: Utility functions for model loading, prompt generation, moderation checks, and loss visualization.
- Python 3.8 or higher
- PyTorch
- Transformers
- OpenAI API (requires
OPENAI_API_KEYenvironment variable) - Additional dependencies listed in the
requirements.txtfile.
-
Clone the repository:
git clone [https://github.com/your-username/jailbreaking-starter.git](https://github.com/phantom2810/Jailbreak-Simulator-Adversarial-Testing-for-LLMs.git)
-
Install dependencies:
pip install -r requirements.txt
-
Set up your OpenAI API key:
export OPENAI_API_KEY=your-api-key
- Configure the input parameters and models in
KOV.py. - Run the main script:
python KOV.py
The framework evaluates adversarial prompts and generates scores for LLM responses:
- Moderation category scores
- Adversarial success rates
- Loss and perplexity metrics
KOV.py: Main script for orchestrating the adversarial testing process.WhiteBox.py: Implements the white-box Markov Decision Process (MDP) for adversarial prompt generation.DPW_Solver.py: Provides MCTS logic with progressive widening for state-action exploration.utils.py: Helper functions for prompt processing, scoring, and response generation.
Contributions are welcome! If you encounter issues or have ideas for improvement, feel free to open an issue or submit a pull request.