Website · Quick start · Docs · Roadmap · Contributing · Discord
PARTHA turns one repository revision into a sealed, queryable intelligence model, and serves architecture, dependencies, engineering review, insights, documentation, and AI context from that one model.
It is for staff and platform engineers, technical founders, and engineering leads who need an inspectable starting point for a codebase they did not write — or no longer fully trust their mental model of. It runs self-hosted; provider-backed AI is optional.
Understanding an unfamiliar or fast-moving codebase means reconstructing the same facts over and over: entry points from folders, dependencies from manifests, boundaries from imports, risk from partial tooling. Documentation, static analysis, and AI each build their own private interpretation, and those interpretations drift apart.
PARTHA builds the interpretation once. A bounded extraction pipeline turns the selected repository revision into a persistent Repository Intelligence snapshot, and every product surface — Architecture, Dependency Graph, Engineering Review, Insights, Documentation, exports, and optional AI — reads that shared model instead of re-parsing the code.
The result is a codebase view that is consistent across surfaces, bound to an exact Git commit or archive hash, inspectable back to source evidence, and explicit about what it could not determine. Where a fact cannot be proven, PARTHA says so rather than guessing.
Each surface reads the sealed snapshot for the analysed revision:
- Repository Intelligence — builds an immutable, revision-addressed structural snapshot from supported repository sources, with evidence and provenance on supported facts.
- Architecture — an interactive graph of modules and resolved relationships, plus an evidence-cited authentication explanation for supported Python/FastAPI patterns.
- Dependency Graph — direct declarations from
package.json,pyproject.toml, andrequirements.txt, with resolved pins from two lockfile formats recorded as resolutions, never as direct edges. - Engineering Review — findings that are each backed by a stored evidence span; unassessed categories stay visible. References into third-party or platform code, local names and non-code assets are not reported as findings. No score, grade, or health percentage.
- Repository Insights — defined counts, ratios, diagnostics, and coverage from one snapshot. No change-over-time claims.
- Repository Lineage — repeated imports of the same repository and branch grouped into a durable history, browsable through the API and UI. This is revision history, not cross-revision comparison.
- Documentation & export — structural documentation, and Review / Documentation / Architecture / Dependencies exported through one JSON / Markdown / HTML / PDF pipeline.
- Optional AI context — per-user provider configuration with encrypted keys and a deployment-owned egress allowlist. Providers receive structural facts and observed paths only — never source bytes or line spans — so answers carry no automatic citations.
The full contract, with every coverage and trust boundary stated per capability, is the capability matrix — generated from the code and drift-checked in CI.
PARTHA is organized as a small monorepo with a single shared backend and multiple product surfaces:
apps/backend— FastAPI service for auth, repository import, durable analysis jobs, snapshot storage, intelligence queries, and export generation.apps/frontend— React + TypeScript app for browsing architecture graphs, dependency views, review findings, insights, repository detail, and settings.apps/marketing— static landing site and scripted product walkthrough, intentionally decoupled from the application frontend.
The backend follows a layered pipeline that mirrors the product promise:
app/api/*— HTTP routes and OpenAPI contract.app/services/*— orchestration of repository, analysis, documentation, and AI operations.app/models/*andapp/repositories/*— persistence layer and owner-scoped repository access.app/extraction/*— repository inventory, language parsing, manifest and lockfile extraction, and IaC detection.app/intelligence/*— canonicalization, snapshot store, retention, and evidence-backed read models.app/analysis/*,app/graph/*,app/review/*, andapp/insights/*— architecture, dependency, review, and diagnostics builders.app/workers/*— durable background analysis execution and lease-based queue controls.app/ai/*— optional AI provider integration and provider-aware repository context assembly.
The project uses a shared repository intelligence snapshot (ri.v1) as the main read model, so downstream consumers read a single sealed interpretation instead of re-parsing source files. This design helps maintain consistency across architecture views, dependency graphs, engineering review, and exports.
The current stack is intentionally pragmatic rather than over-engineered:
- Backend: Python 3.12+, FastAPI, SQLAlchemy, Pydantic, Uvicorn, JWT, Redis, cryptography, tree-sitter.
- Frontend: React 18, TypeScript, Vite, Tailwind CSS, Zustand, React Router.
- Visualization:
@xyflow/react,@dagrejs/dagre,framer-motion,@monaco-editor/react. - Testing:
pytest,vitest,@testing-library/react, andplaywright.
PARTHA is a product-focused repository intelligence platform rather than a generic app template. The core ideas are already solid: a sealed repository snapshot, evidence-backed review, architecture and dependency extraction, and exports that reuse the same source-of-truth model. The main work remaining is operational maturity rather than basic product shape.
The highest leverage areas are:
- Improve backend type-safety and contract discipline.
- Move analysis execution beyond the single in-process worker.
- Reduce whole-repo re-analysis cost and expand semantic coverage.
This does not mean the codebase is immature in a chaotic sense; it means the product is transitioning from a strong proof-of-concept into a more scalable engineering platform.
A few principles show up repeatedly in this codebase:
- Shared facts, not duplicated parsers: if a repository fact is needed in multiple places, it should be built once and reused.
- Evidence-backed outputs: findings and architectural explanations should be traceable to source evidence rather than loose heuristics.
- Owner-scoped access: repositories and their derived data are treated as user-owned resources, not globally shared objects.
- Local-first development: the default dev setup uses SQLite, local storage, and a simple in-memory rate limiter so the project is easy to run and debug.
- Safety-conscious AI usage: optional AI is downstream from repository intelligence; it does not become a second interpretation layer.
The project is intentionally clear about where it is not yet production-hardened:
- In-process worker execution limits concurrency and makes scaling harder.
- Semantic coverage is strongest for Python and TypeScript/JavaScript; other languages mainly contribute inventory data.
- Whole-repository re-analysis remains full-cost and is not incremental.
- AI egress is optional and must be configured carefully to respect policy and trust boundaries.
- The development environment is not a hardened multi-tenant deployment; it should be treated as a trusted local environment.
These are conscious trade-offs in service of a focused product objective, not accidental gaps.
Import → Analyse → Explore → Export
- Import — upload a ZIP/TAR-family archive, or import a public GitHub repository over HTTPS.
- Analyse — PARTHA runs a durable, cancellable background job and seals a snapshot for that exact revision.
- Explore — inspect Architecture, Dependencies, Engineering Review, Insights, evidence, and lineage. A missing or stale snapshot shows an unavailable state, never a fallback interpretation.
- Export — generate structural documentation or export structured results.
flowchart LR
Input["Repository input<br/>archive · public GitHub"]
Import["Import<br/>safe storage · revision identity · file inventory"]
Analyse["Durable analysis<br/>Python · TypeScript/JavaScript · manifests<br/>lockfiles · service interactions · Docker Compose"]
RI[("Sealed ri.v1 snapshot<br/>facts · evidence · diagnostics · canonical hash")]
Product["Architecture · Dependencies · Review<br/>Insights · Documentation · Exports"]
AI["AI provider<br/>optional · structural context only"]
Input --> Import --> Analyse --> RI --> Product
RI -.-> AI
ri.v1 is PARTHA's versioned, sealed snapshot — the single read model. Each immutable snapshot describes one repository at one exact revision; supported facts carry a truth class and, where the contract requires it, provenance tied to an exact source location. The architectural rule is strict:
If a feature needs a repository fact, it belongs in the shared engine — never a second parser inside a consumer. AI is a downstream consumer of Repository Intelligence, never an independent interpreter of the repository.
Supported structural facts retain evidence and provenance back to their repository revision and source location; coverage is surface-dependent, and free-form AI is deliberately uncited. The snapshot's canonical graph hash detects content differences inside a deployment — it is not a digital signature.
Read more: System Overview · Repository Intelligence · RFC-0001 (ri.v1 contract) · RFC-0002 (Repository Lineage)
www.partha.uk is the public project site: it explains the product and includes a scripted walkthrough of an analysis, plus instructions for running PARTHA on your own code. The walkthrough uses fixed sample data and does not call a backend — there is no hosted PARTHA instance to analyse against. To analyse a real repository, run it locally with the quick start below.
| Tool | Version | Needed for |
|---|---|---|
| Python | 3.12 or 3.13 | Backend |
| Node.js | 22 | Frontend and workflow scripts |
| Git | recent | Checkout and public GitHub import |
Development uses SQLite, an in-memory rate limiter, and local filesystem storage. No container runtime or external database is required, and no .env file is needed.
git clone https://github.com/Second-Origin/PARTHA.git
cd PARTHA
# 1. Backend — http://localhost:8000 (OpenAPI at /docs, readiness at /ready)
cd apps/backend
python3.13 -m venv .venv && source .venv/bin/activate
pip install -e .
cd ../.. && npm run dev:backend
# 2. Frontend — http://localhost:5173 (second terminal)
npm ci --prefix apps/frontend
npm run dev:frontendOpen http://localhost:5173, register a local account, add a repository, and start analysis.
If you have Docker, one command builds and starts everything in a single container:
git clone https://github.com/Second-Origin/PARTHA.git
cd PARTHA
docker compose up --buildOpen http://localhost:8000. The first account you register becomes the owner of the instance; approve anyone else with docker compose exec partha python scripts/approve_email.py --email them@example.com. Data and generated secrets live in the partha-data volume, and the port is bound to 127.0.0.1 only. See docker-compose.yml for the details.
The development guide covers the full test / lint / build / benchmark / Docker / E2E commands and the local database and API-contract failures you are most likely to hit. Review the AI provider egress policy before configuring any custom or local provider endpoint.
- Trusted-environment use. PARTHA has not been operated or hardened for broad shared or multi-tenant hosting. Do not expose the development configuration to the public internet.
- Narrow semantic coverage. The deepest extraction is for supported Python and TypeScript/JavaScript constructs; other languages contribute file inventory. Role, module, layer, framework, and entry-point classification can be heuristic.
- Narrow dependency coverage. Three manifest formats and two lockfile formats; no transitive resolution, no vulnerability or outdated-version scanning.
- Whole-repository analysis. Every analysis re-reads the whole repository. There is no incremental re-analysis.
- No cross-revision comparison. Lineage preserves revision history; it does not diff two snapshots, detect renames or moves, or compute a historical blast radius. Change-impact analysis is single-snapshot structural traversal only.
- Optional AI can be external. Depending on configuration, AI calls a configured provider; only local providers keep everything on the host. See the egress policy.
- In-process worker. One daemon worker thread inside the API process handles one analysis job at a time; there is no separate worker service or job queue.
Non-auth product routes require authentication, repository access is owner-scoped, provider keys are Fernet-encrypted at rest, and AI egress is validated against a deployment-owned allowlist with DNS pinning — meaningful controls, but not a claim of production hardening. Registration does not verify email ownership, and in the default development environment any address may register. See SECURITY.md and docs/CAPABILITIES.md for the details.
- Documentation index — every guide, with reading paths
- Capability matrix — the detailed, generated capability contract
- System Overview — components, runtime flow, persistence, trust boundaries
- Repository Intelligence — extraction, snapshot, consumer, and evidence rules
- Local development — running, testing, and troubleshooting the stack
- Connecting an AI provider · AI provider egress policy
- Roadmap · Governance · Security policy
Issues and pull requests are welcome. Start with CONTRIBUTING.md for the fork-first workflow, branch conventions, and Definition of Ready / Done, then pick up a good first issue. Before changing analysis, parsing, or AI-grounding behaviour, read Repository Intelligence in full.
main carries the latest tagged release; dev is where active development happens and is the target of every pull request. A release is a promotion: prepare on dev, promote dev to main, verify, sync back, then tag from main. PARTHA follows Semantic Versioning and is pre-1.0, so minor versions may change behaviour — each release's notes say what moved. See all releases, the changelog, and CONTRIBUTING § Releases.
PARTHA is available under the Apache License 2.0.