Skip to content
ย 
ย 

Latest commit

ย 

History

30 Commits

Folders and files

NameName
Last commit message
Last commit date
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 

Repository files navigation

๐Ÿš€ Multi-Agent Test Automation + Local RAG

An end-to-end AI-assisted test automation pipeline that reads business requirements and OpenAPI specs, designs test scenarios, generates Gherkin, writes Java/Maven Cucumber tests, executes them, analyzes failures, and reports coverage.

The project is Windows-first, but most Python and Node commands also work on Linux/macOS with small shell changes.

Blue.and.White.Minimalist.Geometric.Modern.Thesis.Defense.Presentation.mp4

๐Ÿ“Œ What This Project Does

This repository automates API and end-to-end test creation for microservices.

It combines:

  • ๐Ÿค– LangGraph multi-agent workflow for test design and generation
  • ๐Ÿง  LLM-powered agents for scenarios, Gherkin, validation, test writing, failure analysis, and coverage analysis
  • ๐Ÿ”Ž Local RAG with Chroma to retrieve real-world BDD examples from local datasets
  • ๐Ÿงช Generated Java/Maven Cucumber tests
  • ๐Ÿ“Š JaCoCo coverage analysis
  • ๐Ÿ› ๏ธ Self-healing retry loops for failing generated tests
  • ๐ŸŒ Optional React/Vite dashboard to edit configuration and launch runs
  • ๐Ÿ“ˆ Metrics and plotting tools for evaluating pipeline behavior

The current sample domain is a leave-management system with:

  • ๐Ÿ” auth service on port 9000
  • ๐Ÿ“ leave service on port 9001

The architecture is service-aware, so you can add more services in config/services_matrix.yaml.


๐Ÿงญ High-Level Architecture

Business Requirements + User Story + OpenAPI Specs
                         โ”‚
                         โ–ผ
              LangGraph Multi-Agent Workflow
                         โ”‚
     โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
     โ–ผ                   โ–ผ                   โ–ผ
Scenario Design    Gherkin Generation   Gherkin Validation
     โ”‚                   โ”‚                   โ”‚
     โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                  Java Test Writer
                         โ”‚
                         โ–ผ
                  Maven Test Executor
                         โ”‚
          โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
          โ–ผ                             โ–ผ
   Failure Analyst                Coverage Analyst
          โ”‚                             โ”‚
          โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ retry loops โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                         โ”‚
                         โ–ผ
                Generated Reports + Tests

๐Ÿงฉ Agent Workflow

The pipeline is implemented in graph/workflow.py and passes a shared state object from graph/state.py through the agents.

Step Agent File Responsibility
1๏ธโƒฃ Scenario Designer agents/scenario_designer.py Builds meaningful happy-path, error, edge, security, and integration scenarios
2๏ธโƒฃ Gherkin Generator agents/gherkin_generator.py Converts scenarios into .feature files
3๏ธโƒฃ Gherkin Validator agents/gherkin_validator.py Validates syntax, completeness, and scenario quality
4๏ธโƒฃ Test Writer agents/test_writer.py Generates Java step definitions, runners, Maven config, and test project files
5๏ธโƒฃ Test Executor agents/test_executor.py Runs Maven/Cucumber tests and captures execution results
6๏ธโƒฃ Failure Analyst agents/failure_analyst.py Diagnoses failing tests and recommends repairs
7๏ธโƒฃ Coverage Analyst agents/coverage_analyst.py Reads JaCoCo reports, checks thresholds, and suggests coverage improvements

The workflow includes bounded retry loops:

  • ๐Ÿ” Gherkin validation failures can route back to generation
  • ๐Ÿ” Test failures can route through failure analysis and back to test writing
  • ๐Ÿ” Low coverage can route back to scenario design

๐Ÿ“ Project Structure

.
โ”œโ”€โ”€ agents/                         # LangGraph agent implementations
โ”œโ”€โ”€ graph/                          # Workflow graph and shared state models
โ”œโ”€โ”€ config/                         # Runtime settings and service matrix
โ”œโ”€โ”€ rag/                            # Local RAG ingestion and retrieval
โ”œโ”€โ”€ tools/                          # Swagger parsing, service registry, metrics, plotting, RAG helpers
โ”œโ”€โ”€ examples/                       # User stories and sample OpenAPI specs
โ”œโ”€โ”€ data/raw/                       # Local datasets used by RAG
โ”œโ”€โ”€ output/                         # Generated features, tests, reports, and coverage artifacts
โ”œโ”€โ”€ src/                            # React/Vite dashboard source
โ”œโ”€โ”€ dist/                           # Built dashboard assets
โ”œโ”€โ”€ chroma_db/                      # Persisted Chroma vector database
โ”œโ”€โ”€ run_pipeline.py                 # Main service-aware pipeline runner
โ”œโ”€โ”€ run_pipeline_windows.py         # Windows UTF-8 wrapper
โ”œโ”€โ”€ main.py                         # CLI for RAG and demo commands
โ”œโ”€โ”€ app_server.py                   # Optional dashboard backend
โ”œโ”€โ”€ business_requirements.yaml      # Business rules, scenarios, and coverage targets
โ”œโ”€โ”€ requirements.txt                # Python dependencies
โ”œโ”€โ”€ package.json                    # Node/Vite/gherkin-lint dependencies
โ””โ”€โ”€ README.md                       # This documentation

โœ… Requirements

Install these before running the full pipeline:

  • ๐Ÿ Python 3.11+
    Python 3.12 is used successfully in this workspace.

  • โ˜• Java 17

  • ๐Ÿ“ฆ Maven 3.9+

  • ๐ŸŸข Node.js + npm
    Used for the UI and gherkin-lint.

  • ๐Ÿ”Œ Target microservices running or reachable
    The generated API tests need live services unless you only generate files.

  • ๐Ÿ”‘ LLM provider credentials
    The sample .env.example is configured for Groq, with optional Hugging Face support.


โš™๏ธ Installation

1๏ธโƒฃ Create and activate a Python virtual environment

PowerShell:

python -m venv .venv
.\.venv\Scripts\Activate.ps1
pip install -r requirements.txt

2๏ธโƒฃ Install Node dependencies

npm install

3๏ธโƒฃ Create your .env

Copy the example file:

Copy-Item .env.example .env

Then fill in your real values:

# LLM provider config
LLM_PROVIDER=groq
GROQ_API_KEY=your_key_here

# Optional Hugging Face token
HUGGINGFACEHUB_API_TOKEN=

# RAG config
RAG_ENABLE=true
RAG_PERSIST_DIR=chroma_db
RAG_COLLECTION=tier3_rag
RAG_EMBEDDING_MODEL=sentence-transformers/all-MiniLM-L6-v2

โš ๏ธ Never commit .env. It can contain API keys, JWTs, credentials, and service secrets.


๐Ÿ”ง Service Configuration

Services are declared in:

config/services_matrix.yaml

Each service can define:

  • โœ… whether it is enabled
  • ๐ŸŒ base URL and port
  • ๐Ÿ“„ local Swagger/OpenAPI file path
  • ๐Ÿ”— remote Swagger/OpenAPI URL
  • ๐Ÿงช Java package, runner class, step class, and Maven location
  • ๐Ÿ—„๏ธ database metadata
  • ๐Ÿงฉ dependencies on other services
  • ๐Ÿงฌ test data used by generated scenarios

Example service shape:

services:
  auth:
    enabled: true
    port: 9000
    base_url: http://localhost:9000
    swagger_spec: ''
    swagger_url: http://localhost:9000/v3/api-docs
    role: authentication
    dependencies: []

Useful inspection commands:

python run_pipeline.py --list
python run_pipeline.py --order

๐Ÿงช Running the Pipeline

Recommended on Windows

python run_pipeline_windows.py

Standard runner

python run_pipeline.py

Run only selected services

python run_pipeline.py --services auth
python run_pipeline.py --services leave
python run_pipeline.py --services auth,leave

The runner will:

  1. ๐Ÿ“– Read examples/comprehensive_user_story.md
  2. ๐Ÿ“‹ Load rules and coverage targets from business_requirements.yaml
  3. ๐Ÿงญ Load service definitions from config/services_matrix.yaml
  4. ๐Ÿ“ก Fetch OpenAPI specs from each configured swagger_url or swagger_spec
  5. ๐Ÿค– Execute the LangGraph workflow
  6. ๐Ÿงช Generate and run Java/Maven tests
  7. ๐Ÿ“Š Analyze JaCoCo coverage
  8. ๐Ÿ“ฆ Write results into output/

๐Ÿ“ Important Input Files

File Purpose
examples/comprehensive_user_story.md Main user story used by the pipeline
business_requirements.yaml Business rules, critical endpoints, test scenario categories, priorities, and coverage targets
config/services_matrix.yaml Service registry and execution configuration
.env Local secrets and runtime environment variables
data/raw/GivenWhenThen.json Main BDD dataset used by RAG

๐Ÿ“ฆ Output Files

Generated artifacts are written under output/.

Path Contains
output/features/ Generated .feature files
output/tests/ Generated Java/Maven test project
output/reports/ Workflow summaries, execution summaries, and coverage JSON/YAML
output/jacoco/ JaCoCo execution data and HTML reports when produced

Useful generated documentation:

  • ๐Ÿ“Œ output/tests/START_HERE.md
  • ๐Ÿ“– output/tests/README.md
  • ๐Ÿงช output/tests/TEST_EXECUTION_GUIDE.md
  • ๐Ÿ“Š output/tests/TEST_GENERATION_SUMMARY.md
  • โœ… output/tests/VERIFICATION_REPORT.md

๐Ÿง  Local RAG

RAG is optional, but it improves scenario quality by giving agents examples from real-world BDD test corpora.

RAG sources

  • ๐Ÿ“š data/raw/GivenWhenThen.json
  • ๐Ÿ—ƒ๏ธ Optional E2EGit SQLite/CSV corpus
  • ๐Ÿ“ Optional user story text corpora

Build the Chroma index

python main.py ingest

Custom example:

python main.py ingest --givenwhenthen-json data/raw/GivenWhenThen.json --persist-dir chroma_db --collection tier3_rag

Query the local index

python main.py query "JWT authentication and leave approval"
python main.py query "invalid leave request boundary case" --k 3

RAG environment variables

$env:RAG_ENABLE="true"
$env:RAG_PERSIST_DIR="chroma_db"
$env:RAG_COLLECTION="tier3_rag"
$env:RAG_EMBEDDING_MODEL="sentence-transformers/all-MiniLM-L6-v2"

For more detail, see RAG_STRUCTURE.md.


๐Ÿ“Š Coverage Gates

Coverage targets are configured in business_requirements.yaml:

COVERAGE_TARGETS:
  LINE_COVERAGE: 50.0
  BRANCH_COVERAGE: 25.0
  METHOD_COVERAGE: 70.0

You can override them at runtime:

$env:MIN_LINE_COVERAGE="60"
$env:MIN_BRANCH_COVERAGE="40"
$env:MIN_METHOD_COVERAGE="60"
python run_pipeline_windows.py

Helpful flags:

Variable Effect
SKIP_TEST_EXECUTION=1 Generate tests without running Maven
FAIL_ON_COVERAGE_QG=1 Fail when coverage gates are not met
ALLOW_COVERAGE_QG_FAILURE=1 Continue even if quality gates fail
ENABLE_COVERAGE_IMPROVEMENT=false Disable the coverage improvement loop
MAX_HEALING_ATTEMPTS=3 Control failure-healing retries
MAX_GHERKIN_VALIDATION_RETRIES=2 Control Gherkin regeneration retries
MAX_COVERAGE_IMPROVEMENT_ATTEMPTS=1 Control coverage improvement retries

๐Ÿ—๏ธ Real Backend Coverage

By default, generated Maven tests can produce coverage for the generated test project. For coverage of the actual running backend services, start services with the JaCoCo Java agent and then run the pipeline.

Relevant helpers:

  • restart_services_with_jacoco.py
  • restart_services_with_jacoco.ps1
  • restart-services-with-jacoco.ps1
  • start_backend_with_jacoco.bat
  • collect_jacoco_coverage.py
  • run_real_coverage.ps1

High-level flow:

  1. โ˜• Start backend microservices with the JaCoCo agent
  2. ๐Ÿงช Run generated tests against live services
  3. ๐Ÿ“ฅ Dump .exec coverage files
  4. ๐Ÿ“Š Generate JaCoCo HTML reports

โš ๏ธ Some helper scripts may contain machine-specific paths. Review them before using them on another machine.

More notes are available in REAL_BACKEND_COVERAGE_SOLUTION.md.


๐ŸŒ Optional Dashboard

The project includes a React/Vite dashboard backed by app_server.py.

The dashboard can:

  • ๐Ÿงญ Load and edit config/services_matrix.yaml
  • ๐Ÿ“ Edit the main user story
  • ๐Ÿ“‹ Edit business requirements
  • โ–ถ๏ธ Start pipeline runs
  • ๐Ÿ“ก Show run status and logs
  • ๐Ÿ“Ž Link to latest output artifacts

Development UI

npm run dev

Build and serve through Python

npm run build
python app_server.py

Then open:

http://127.0.0.1:8000

Backend API routes:

Route Method Purpose
/api/state GET Load current UI/project state
/api/save POST Save edited config/story/requirements
/api/run POST Start a pipeline run
/api/run-status GET Poll run status
/files/{path} GET Serve safe project artifacts

๐Ÿงฐ CLI Commands

Pipeline

python run_pipeline.py
python run_pipeline.py --list
python run_pipeline.py --order
python run_pipeline.py --services auth,leave

RAG

python main.py ingest
python main.py query "API validation error scenario"
python main.py demo-gherkin

UI

npm run dev
npm run build
python app_server.py

Metrics and plots

powershell -ExecutionPolicy Bypass -File tools/run_metrics_and_plots.ps1
python tools/eval_metrics.py
python tools/plot_metrics.py

๐Ÿ“ˆ Metrics and Evaluation Tools

The tools/ folder includes scripts for inspecting generated outputs and comparing pipeline performance:

  • ๐Ÿ“Š tools/eval_metrics.py
  • ๐Ÿ“‰ tools/plot_metrics.py
  • ๐Ÿค– tools/run_llm_benchmark.py
  • ๐Ÿงฎ tools/recompute_benchmark_metrics.py
  • ๐Ÿ–ผ๏ธ tools/render_agent_benchmark_png.py
  • ๐Ÿงญ tools/render_pipeline_agents_png.py
  • ๐Ÿ“Œ tools/plot_loop_vs_no_loop.py
  • ๐Ÿ” tools/analyze_cucumber_failures.py

These are useful for research-style evaluation, debugging, and comparing LLM behavior.


๐Ÿงช Generated Test Project

The generated test project lives in:

output/tests/

Typical commands:

cd output/tests
mvn clean test
mvn test
mvn jacoco:report

On Windows, generated helper scripts may also be available:

.\run_tests.bat

Generated test docs usually include:

  • ๐ŸŽฏ quick start
  • ๐Ÿงช execution guide
  • ๐Ÿ“ file index
  • โœ… verification report
  • ๐Ÿ“Š generation summary

๐Ÿ” Security Notes

  • ๐Ÿšซ Do not commit .env
  • ๐Ÿšซ Do not commit real JWT tokens, passwords, database credentials, or API keys
  • ๐Ÿ”Ž Review config/services_matrix.yaml before sharing because it can contain local credentials or tokens
  • ๐Ÿงผ Treat generated tests as code that must be reviewed before production use
  • ๐Ÿ”’ Prefer environment variables for secrets instead of hard-coding them into YAML or scripts

๐Ÿž Troubleshooting

.env is missing

Create it from the example:

Copy-Item .env.example .env

Windows encoding problems

Use:

python run_pipeline_windows.py

Services are unreachable

Check:

  • the service is running
  • the port matches config/services_matrix.yaml
  • swagger_url returns JSON
  • firewalls or proxies are not blocking localhost

Swagger cannot be loaded

Either set:

swagger_spec: path/to/openapi.json

or:

swagger_url: http://localhost:9000/v3/api-docs

RAG returns no results

Run ingestion:

python main.py ingest

Then confirm chroma_db/ exists.

Maven is not found

Check Java and Maven:

java -version
mvn -version

If your Maven path is custom, update:

config/services_matrix.yaml

Tests fail because authentication expired

Refresh the JWT/login test data and verify:

  • login endpoint
  • credentials
  • token expiry
  • authorization headers generated by the test writer

Coverage gates fail

Lower thresholds temporarily or allow failure while debugging:

$env:ALLOW_COVERAGE_QG_FAILURE="1"
python run_pipeline_windows.py

๐Ÿ—บ๏ธ Key Files to Read First

File Why it matters
run_pipeline.py Main pipeline entry point
graph/workflow.py LangGraph node order and retry routing
graph/state.py Shared workflow state passed between agents
agents/scenario_designer.py Scenario planning logic
agents/test_writer.py Java/Maven test generation logic
agents/test_executor.py Maven execution logic
agents/coverage_analyst.py Coverage parsing and quality gates
tools/service_registry.py Reads and normalizes configured services
tools/swagger_parser.py OpenAPI parsing helpers
tools/rag_scenario_retriever.py RAG examples injected into prompts
main.py RAG CLI commands
app_server.py Optional dashboard backend

๐Ÿšฆ Recommended Workflow

For a clean run:

  1. โœ… Start your backend services
  2. โœ… Verify Swagger endpoints in the browser
  3. โœ… Activate Python venv
  4. โœ… Install Python and Node dependencies
  5. โœ… Fill .env
  6. โœ… Review config/services_matrix.yaml
  7. โœ… Optionally run python main.py ingest
  8. โœ… Run python run_pipeline_windows.py
  9. โœ… Inspect output/tests/START_HERE.md
  10. โœ… Review generated tests before using them in CI

๐Ÿงพ Current Domain Example

The included business requirements describe:

  • ๐Ÿ” authentication and user management
  • ๐Ÿข department and role management
  • ๐Ÿ“ leave request creation
  • โœ… approval and rejection workflows
  • ๐Ÿ›ก๏ธ JWT and role-based access rules
  • ๐Ÿ“† date validation, overlap detection, and leave balance constraints
  • ๐Ÿ”— integration between the auth and leave services

The project is not limited to that domain. Add or edit services and requirements to generate tests for other APIs.


โœจ Summary

This project is a complete AI-assisted QA pipeline:

  • ๐Ÿง  It understands requirements
  • ๐Ÿ“ก It reads OpenAPI specs
  • ๐Ÿงช It generates executable tests
  • ๐Ÿ” It retries and improves when things fail
  • ๐Ÿ“Š It checks coverage
  • ๐ŸŒ It can be controlled from a local dashboard
  • ๐Ÿ”Ž It can use local RAG to ground generation in real test examples

Use it as a test-generation accelerator, a research prototype, or a foundation for automated API QA across multiple microservices.

About

An Agentic AI and RAG-powered framework for automated microservices testing. Vortex leverages autonomous AI agents to generate, execute, and analyze end-to-end API test suites.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages