A Dual-Agent AI Research System for Automated Prompt Refinement using the RISE Framework
PromptFlow is a research-driven, dual-agent AI system designed to evaluate whether automated prompt refinement improves the quality of responses generated by Large Language Models (LLMs).
Instead of directly sending a user's raw query to an LLM, PromptFlow first refines the query into a structured prompt using the RISE framework (Role, Instruction, Steps, Expectation). The refined prompt is then passed to a generator LLM. By running an A/B test pipeline, the project directly measures the impact of prompt engineering on output quality.
Users interacting with LLMs often submit queries that are:
- Incomplete
- Ambiguous
- Unstructured
- Missing critical constraints
These raw inputs significantly degrade the quality of LLM responses. PromptFlow solves this by automatically transforming raw inputs into optimized, structured prompts without altering the user's original intent.
PromptFlow uses a lightweight, dual-agent architecture optimized for local inference.
- Model:
Gemma 3 1B Instruction - Purpose: Automatically structures raw user queries into the RISE format.
- Fine-Tuning: Supervised Fine-Tuning (QLoRA) on a custom dataset of ~2,100+ RISE prompt samples.
- Input: Raw user query.
- Task: Understand intent, preserve constraints, infer missing structure, apply RISE.
- Model:
Gemma 4 E2B Instruction GGUF - Purpose: Generates the final output response. No additional fine-tuning.
- Execution: Runs efficiently via
llama.cppfor local CPU inference.
- Model:
Nemotron Ultra 3 - Purpose: An automated evaluator model that compares the direct response versus the refined response across six specific metrics.
Every prompt refined by Agent 1 strictly adheres to the RISE structure:
- Role: Assigns an expert persona relevant to the query.
- Instruction: Provides a clear, unambiguous task description.
- Steps: Outlines a logical workflow or methodology for solving the problem.
- Expectation: Defines the exact formatting and output requirements.
The core contribution of this project is investigating the hypothesis: Does automatic prompt refinement improve LLM response quality?
For every user query, the system evaluates two parallel generation tracks:
- 🔴 Left Panel (Control): Raw Query →
Agent 2→ Baseline Response - 🟢 Right Panel (Variable): Raw Query →
Agent 1→ Refined RISE Prompt →Agent 2→ Optimized Response
The Judge Model quantitatively scores both outputs using the following six criteria:
- Relevance
- Clarity
- Completeness
- Actionability
- Structure
- Depth
- Backend: FastAPI, Python
- AI / Inference: Hugging Face Transformers,
llama.cpp, GGUF models - Fine-Tuning: PEFT (QLoRA)
- Frontend: React
- Hardware Profile: Designed for 100% Local Inference (CPU/Consumer GPU)
Unlike traditional prompt engineering tools that act merely as wrappers, PromptFlow:
- Operates completely offline with open-source local LLMs.
- Preserves constraints while intelligently filling contextual gaps.
- Quantitatively validates outcomes using an automated judge.
Expected Outcome: The refined prompt pipeline should consistently produce responses that are significantly more relevant, complete, structured, actionable, deeper, and easier to understand than responses generated from raw user input.
(Placeholder for future deployment instructions)
# 1. Clone the repository
git clone (https://github.com/Devansh-Mankad/promptflow.git)
cd promptflow
# 2. Setup Python Virtual Environment
python -m venv venv
source venv/bin/activate # On Windows: venv\Scripts\activate
# 3. Install dependencies
pip install -r requirements.txt
# 4. Start the FastAPI Backend
uvicorn main:app --reload
# 5. Start the React Frontend
cd frontend
npm install
npm startDevansh Mankad
Computer Engineering Student
If you found this project useful, consider giving it a ⭐ Star on GitHub.
This project is licensed under the MIT License.