An end-to-end AI-assisted test automation pipeline that reads business requirements and OpenAPI specs, designs test scenarios, generates Gherkin, writes Java/Maven Cucumber tests, executes them, analyzes failures, and reports coverage.
The project is Windows-first, but most Python and Node commands also work on Linux/macOS with small shell changes.
Blue.and.White.Minimalist.Geometric.Modern.Thesis.Defense.Presentation.mp4
This repository automates API and end-to-end test creation for microservices.
It combines:
- ๐ค LangGraph multi-agent workflow for test design and generation
- ๐ง LLM-powered agents for scenarios, Gherkin, validation, test writing, failure analysis, and coverage analysis
- ๐ Local RAG with Chroma to retrieve real-world BDD examples from local datasets
- ๐งช Generated Java/Maven Cucumber tests
- ๐ JaCoCo coverage analysis
- ๐ ๏ธ Self-healing retry loops for failing generated tests
- ๐ Optional React/Vite dashboard to edit configuration and launch runs
- ๐ Metrics and plotting tools for evaluating pipeline behavior
The current sample domain is a leave-management system with:
- ๐
authservice on port9000 - ๐
leaveservice on port9001
The architecture is service-aware, so you can add more services in config/services_matrix.yaml.
Business Requirements + User Story + OpenAPI Specs
โ
โผ
LangGraph Multi-Agent Workflow
โ
โโโโโโโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโโโ
โผ โผ โผ
Scenario Design Gherkin Generation Gherkin Validation
โ โ โ
โโโโโโโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโโโ
Java Test Writer
โ
โผ
Maven Test Executor
โ
โโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโ
โผ โผ
Failure Analyst Coverage Analyst
โ โ
โโโโโโโโโ retry loops โโโโโโโโโ
โ
โผ
Generated Reports + Tests
The pipeline is implemented in graph/workflow.py and passes a shared state object from graph/state.py through the agents.
| Step | Agent | File | Responsibility |
|---|---|---|---|
| 1๏ธโฃ | Scenario Designer | agents/scenario_designer.py |
Builds meaningful happy-path, error, edge, security, and integration scenarios |
| 2๏ธโฃ | Gherkin Generator | agents/gherkin_generator.py |
Converts scenarios into .feature files |
| 3๏ธโฃ | Gherkin Validator | agents/gherkin_validator.py |
Validates syntax, completeness, and scenario quality |
| 4๏ธโฃ | Test Writer | agents/test_writer.py |
Generates Java step definitions, runners, Maven config, and test project files |
| 5๏ธโฃ | Test Executor | agents/test_executor.py |
Runs Maven/Cucumber tests and captures execution results |
| 6๏ธโฃ | Failure Analyst | agents/failure_analyst.py |
Diagnoses failing tests and recommends repairs |
| 7๏ธโฃ | Coverage Analyst | agents/coverage_analyst.py |
Reads JaCoCo reports, checks thresholds, and suggests coverage improvements |
The workflow includes bounded retry loops:
- ๐ Gherkin validation failures can route back to generation
- ๐ Test failures can route through failure analysis and back to test writing
- ๐ Low coverage can route back to scenario design
.
โโโ agents/ # LangGraph agent implementations
โโโ graph/ # Workflow graph and shared state models
โโโ config/ # Runtime settings and service matrix
โโโ rag/ # Local RAG ingestion and retrieval
โโโ tools/ # Swagger parsing, service registry, metrics, plotting, RAG helpers
โโโ examples/ # User stories and sample OpenAPI specs
โโโ data/raw/ # Local datasets used by RAG
โโโ output/ # Generated features, tests, reports, and coverage artifacts
โโโ src/ # React/Vite dashboard source
โโโ dist/ # Built dashboard assets
โโโ chroma_db/ # Persisted Chroma vector database
โโโ run_pipeline.py # Main service-aware pipeline runner
โโโ run_pipeline_windows.py # Windows UTF-8 wrapper
โโโ main.py # CLI for RAG and demo commands
โโโ app_server.py # Optional dashboard backend
โโโ business_requirements.yaml # Business rules, scenarios, and coverage targets
โโโ requirements.txt # Python dependencies
โโโ package.json # Node/Vite/gherkin-lint dependencies
โโโ README.md # This documentation
Install these before running the full pipeline:
-
๐ Python 3.11+
Python 3.12 is used successfully in this workspace. -
โ Java 17
-
๐ฆ Maven 3.9+
-
๐ข Node.js + npm
Used for the UI andgherkin-lint. -
๐ Target microservices running or reachable
The generated API tests need live services unless you only generate files. -
๐ LLM provider credentials
The sample.env.exampleis configured for Groq, with optional Hugging Face support.
PowerShell:
python -m venv .venv
.\.venv\Scripts\Activate.ps1
pip install -r requirements.txtnpm installCopy the example file:
Copy-Item .env.example .envThen fill in your real values:
# LLM provider config
LLM_PROVIDER=groq
GROQ_API_KEY=your_key_here
# Optional Hugging Face token
HUGGINGFACEHUB_API_TOKEN=
# RAG config
RAG_ENABLE=true
RAG_PERSIST_DIR=chroma_db
RAG_COLLECTION=tier3_rag
RAG_EMBEDDING_MODEL=sentence-transformers/all-MiniLM-L6-v2.env. It can contain API keys, JWTs, credentials, and service secrets.
Services are declared in:
config/services_matrix.yaml
Each service can define:
- โ whether it is enabled
- ๐ base URL and port
- ๐ local Swagger/OpenAPI file path
- ๐ remote Swagger/OpenAPI URL
- ๐งช Java package, runner class, step class, and Maven location
- ๐๏ธ database metadata
- ๐งฉ dependencies on other services
- ๐งฌ test data used by generated scenarios
Example service shape:
services:
auth:
enabled: true
port: 9000
base_url: http://localhost:9000
swagger_spec: ''
swagger_url: http://localhost:9000/v3/api-docs
role: authentication
dependencies: []Useful inspection commands:
python run_pipeline.py --list
python run_pipeline.py --orderpython run_pipeline_windows.pypython run_pipeline.pypython run_pipeline.py --services auth
python run_pipeline.py --services leave
python run_pipeline.py --services auth,leaveThe runner will:
- ๐ Read
examples/comprehensive_user_story.md - ๐ Load rules and coverage targets from
business_requirements.yaml - ๐งญ Load service definitions from
config/services_matrix.yaml - ๐ก Fetch OpenAPI specs from each configured
swagger_urlorswagger_spec - ๐ค Execute the LangGraph workflow
- ๐งช Generate and run Java/Maven tests
- ๐ Analyze JaCoCo coverage
- ๐ฆ Write results into
output/
| File | Purpose |
|---|---|
examples/comprehensive_user_story.md |
Main user story used by the pipeline |
business_requirements.yaml |
Business rules, critical endpoints, test scenario categories, priorities, and coverage targets |
config/services_matrix.yaml |
Service registry and execution configuration |
.env |
Local secrets and runtime environment variables |
data/raw/GivenWhenThen.json |
Main BDD dataset used by RAG |
Generated artifacts are written under output/.
| Path | Contains |
|---|---|
output/features/ |
Generated .feature files |
output/tests/ |
Generated Java/Maven test project |
output/reports/ |
Workflow summaries, execution summaries, and coverage JSON/YAML |
output/jacoco/ |
JaCoCo execution data and HTML reports when produced |
Useful generated documentation:
- ๐
output/tests/START_HERE.md - ๐
output/tests/README.md - ๐งช
output/tests/TEST_EXECUTION_GUIDE.md - ๐
output/tests/TEST_GENERATION_SUMMARY.md - โ
output/tests/VERIFICATION_REPORT.md
RAG is optional, but it improves scenario quality by giving agents examples from real-world BDD test corpora.
- ๐
data/raw/GivenWhenThen.json - ๐๏ธ Optional E2EGit SQLite/CSV corpus
- ๐ Optional user story text corpora
python main.py ingestCustom example:
python main.py ingest --givenwhenthen-json data/raw/GivenWhenThen.json --persist-dir chroma_db --collection tier3_ragpython main.py query "JWT authentication and leave approval"
python main.py query "invalid leave request boundary case" --k 3$env:RAG_ENABLE="true"
$env:RAG_PERSIST_DIR="chroma_db"
$env:RAG_COLLECTION="tier3_rag"
$env:RAG_EMBEDDING_MODEL="sentence-transformers/all-MiniLM-L6-v2"For more detail, see RAG_STRUCTURE.md.
Coverage targets are configured in business_requirements.yaml:
COVERAGE_TARGETS:
LINE_COVERAGE: 50.0
BRANCH_COVERAGE: 25.0
METHOD_COVERAGE: 70.0You can override them at runtime:
$env:MIN_LINE_COVERAGE="60"
$env:MIN_BRANCH_COVERAGE="40"
$env:MIN_METHOD_COVERAGE="60"
python run_pipeline_windows.pyHelpful flags:
| Variable | Effect |
|---|---|
SKIP_TEST_EXECUTION=1 |
Generate tests without running Maven |
FAIL_ON_COVERAGE_QG=1 |
Fail when coverage gates are not met |
ALLOW_COVERAGE_QG_FAILURE=1 |
Continue even if quality gates fail |
ENABLE_COVERAGE_IMPROVEMENT=false |
Disable the coverage improvement loop |
MAX_HEALING_ATTEMPTS=3 |
Control failure-healing retries |
MAX_GHERKIN_VALIDATION_RETRIES=2 |
Control Gherkin regeneration retries |
MAX_COVERAGE_IMPROVEMENT_ATTEMPTS=1 |
Control coverage improvement retries |
By default, generated Maven tests can produce coverage for the generated test project. For coverage of the actual running backend services, start services with the JaCoCo Java agent and then run the pipeline.
Relevant helpers:
restart_services_with_jacoco.pyrestart_services_with_jacoco.ps1restart-services-with-jacoco.ps1start_backend_with_jacoco.batcollect_jacoco_coverage.pyrun_real_coverage.ps1
High-level flow:
- โ Start backend microservices with the JaCoCo agent
- ๐งช Run generated tests against live services
- ๐ฅ Dump
.execcoverage files - ๐ Generate JaCoCo HTML reports
More notes are available in REAL_BACKEND_COVERAGE_SOLUTION.md.
The project includes a React/Vite dashboard backed by app_server.py.
The dashboard can:
- ๐งญ Load and edit
config/services_matrix.yaml - ๐ Edit the main user story
- ๐ Edit business requirements
โถ๏ธ Start pipeline runs- ๐ก Show run status and logs
- ๐ Link to latest output artifacts
npm run devnpm run build
python app_server.pyThen open:
http://127.0.0.1:8000
Backend API routes:
| Route | Method | Purpose |
|---|---|---|
/api/state |
GET |
Load current UI/project state |
/api/save |
POST |
Save edited config/story/requirements |
/api/run |
POST |
Start a pipeline run |
/api/run-status |
GET |
Poll run status |
/files/{path} |
GET |
Serve safe project artifacts |
python run_pipeline.py
python run_pipeline.py --list
python run_pipeline.py --order
python run_pipeline.py --services auth,leavepython main.py ingest
python main.py query "API validation error scenario"
python main.py demo-gherkinnpm run dev
npm run build
python app_server.pypowershell -ExecutionPolicy Bypass -File tools/run_metrics_and_plots.ps1
python tools/eval_metrics.py
python tools/plot_metrics.pyThe tools/ folder includes scripts for inspecting generated outputs and comparing pipeline performance:
- ๐
tools/eval_metrics.py - ๐
tools/plot_metrics.py - ๐ค
tools/run_llm_benchmark.py - ๐งฎ
tools/recompute_benchmark_metrics.py - ๐ผ๏ธ
tools/render_agent_benchmark_png.py - ๐งญ
tools/render_pipeline_agents_png.py - ๐
tools/plot_loop_vs_no_loop.py - ๐
tools/analyze_cucumber_failures.py
These are useful for research-style evaluation, debugging, and comparing LLM behavior.
The generated test project lives in:
output/tests/
Typical commands:
cd output/tests
mvn clean test
mvn test
mvn jacoco:reportOn Windows, generated helper scripts may also be available:
.\run_tests.batGenerated test docs usually include:
- ๐ฏ quick start
- ๐งช execution guide
- ๐ file index
- โ verification report
- ๐ generation summary
- ๐ซ Do not commit
.env - ๐ซ Do not commit real JWT tokens, passwords, database credentials, or API keys
- ๐ Review
config/services_matrix.yamlbefore sharing because it can contain local credentials or tokens - ๐งผ Treat generated tests as code that must be reviewed before production use
- ๐ Prefer environment variables for secrets instead of hard-coding them into YAML or scripts
Create it from the example:
Copy-Item .env.example .envUse:
python run_pipeline_windows.pyCheck:
- the service is running
- the port matches
config/services_matrix.yaml swagger_urlreturns JSON- firewalls or proxies are not blocking localhost
Either set:
swagger_spec: path/to/openapi.jsonor:
swagger_url: http://localhost:9000/v3/api-docsRun ingestion:
python main.py ingestThen confirm chroma_db/ exists.
Check Java and Maven:
java -version
mvn -versionIf your Maven path is custom, update:
config/services_matrix.yaml
Refresh the JWT/login test data and verify:
- login endpoint
- credentials
- token expiry
- authorization headers generated by the test writer
Lower thresholds temporarily or allow failure while debugging:
$env:ALLOW_COVERAGE_QG_FAILURE="1"
python run_pipeline_windows.py| File | Why it matters |
|---|---|
run_pipeline.py |
Main pipeline entry point |
graph/workflow.py |
LangGraph node order and retry routing |
graph/state.py |
Shared workflow state passed between agents |
agents/scenario_designer.py |
Scenario planning logic |
agents/test_writer.py |
Java/Maven test generation logic |
agents/test_executor.py |
Maven execution logic |
agents/coverage_analyst.py |
Coverage parsing and quality gates |
tools/service_registry.py |
Reads and normalizes configured services |
tools/swagger_parser.py |
OpenAPI parsing helpers |
tools/rag_scenario_retriever.py |
RAG examples injected into prompts |
main.py |
RAG CLI commands |
app_server.py |
Optional dashboard backend |
For a clean run:
- โ Start your backend services
- โ Verify Swagger endpoints in the browser
- โ Activate Python venv
- โ Install Python and Node dependencies
- โ
Fill
.env - โ
Review
config/services_matrix.yaml - โ
Optionally run
python main.py ingest - โ
Run
python run_pipeline_windows.py - โ
Inspect
output/tests/START_HERE.md - โ Review generated tests before using them in CI
The included business requirements describe:
- ๐ authentication and user management
- ๐ข department and role management
- ๐ leave request creation
- โ approval and rejection workflows
- ๐ก๏ธ JWT and role-based access rules
- ๐ date validation, overlap detection, and leave balance constraints
- ๐ integration between the auth and leave services
The project is not limited to that domain. Add or edit services and requirements to generate tests for other APIs.
This project is a complete AI-assisted QA pipeline:
- ๐ง It understands requirements
- ๐ก It reads OpenAPI specs
- ๐งช It generates executable tests
- ๐ It retries and improves when things fail
- ๐ It checks coverage
- ๐ It can be controlled from a local dashboard
- ๐ It can use local RAG to ground generation in real test examples
Use it as a test-generation accelerator, a research prototype, or a foundation for automated API QA across multiple microservices.