A portfolio-grade Agentic RAG system for financial due diligence and risk intelligence that combines SEC EDGAR filings, XBRL facts, hybrid retrieval, deterministic financial tools, temporal disclosure comparison, a provenance-preserving risk graph, local Qwen3 reasoning, and claim-level citation verification.
Status: Portfolio-ready, evaluated, GitHub CI validated, and Hugging Face portfolio application deployed
Repository: unit-mole/filingsgraph-agentic-rag
Live application: Open the FilingsGraph Hugging Face Space
Primary stack: Python · Qwen3-8B · BGE-M3 · BGE reranker · BM25 · Reciprocal Rank Fusion · Qdrant · DuckDB · NetworkX · SEC EDGAR/XBRL · FastAPI · Gradio · pytest · Ruff
Mandatory paid external LLM/API dependency: $0.
The validated reasoning path uses locally executed open-source models. Hardware, electricity, storage, and internet access are not claimed to be free.
FilingsGraph is a research, engineering, and portfolio project. It is not investment, legal, accounting, or trading advice.
- Source material is restricted to publicly available filing and financial information.
- Filing text is treated as untrusted executable input: instructions embedded inside filings are never followed as agent commands.
- SEC/XBRL evidence is still treated as legitimate research evidence when cited and verified.
- Deterministic financial calculations are preferred over free-form model arithmetic.
- Generated synthesis is evidence-grounded and citation-checked.
- Human review remains appropriate for high-stakes financial conclusions.
Public-company due diligence is rarely a single retrieval problem.
An analyst may need to answer questions such as:
“How did NVIDIA revenue change while its export-control risk disclosure evolved across fiscal years?”
That requires multiple forms of reasoning at once:
- identify the correct company and filing period,
- retrieve narrative filing evidence,
- resolve exact XBRL financial facts,
- distinguish old versus new risk language,
- identify common risk exposure across companies,
- preserve filing/section provenance,
- calculate changes deterministically,
- and ensure every material claim is supported by evidence.
Traditional dense RAG is not enough for this workflow. FilingsGraph separates the problem into specialized deterministic and retrieval-backed tools and routes each research question to the appropriate evidence path.
Build an end-to-end financial research system that can:
- Resolve companies and SEC identifiers.
- Ingest and parse multi-year 10-K filings.
- Ingest SEC Company Facts / XBRL data.
- Preserve section, filing, fiscal-year, and company provenance.
- Build hierarchical textual chunks and extracted tables.
- Compare BM25, dense, hybrid, and reranked retrieval.
- Route exact financial questions to deterministic XBRL tools.
- Compare disclosure language across fiscal years.
- Build a temporal company-to-risk knowledge graph.
- Answer multi-company GraphRAG questions.
- Use local Qwen3 reasoning without mandatory paid APIs.
- Enforce claim-level citation attachment and prevent thinking leakage.
- Evaluate retrieval, financial accuracy, routing, temporal classification, graph quality, and grounding separately.
- Freeze the core retrieval architecture before final evaluation.
- Expose the system through FastAPI and Gradio.
- Publish the implementation and measured limitations transparently.
| Item | Implementation |
|---|---|
| Project name | filingsgraph-agentic-rag |
| Application | Temporal financial due diligence and filing-risk intelligence |
| Primary filings | SEC 10-K |
| Initial cohort | NVDA · AMD · INTC · AVGO · QCOM |
| Reasoning model | Qwen3-8B locally |
| Alternative model evaluated | Qwen3-14B |
| Dense embeddings | BAAI/bge-m3 |
| Sparse retrieval | BM25 |
| Fusion | Reciprocal Rank Fusion |
| Reranker | BAAI/bge-reranker-v2-m3 |
| Vector database | Qdrant |
| Structured financial store | DuckDB |
| Graph | NetworkX temporal risk graph |
| Financial calculations | Deterministic XBRL tools |
| Agent | Routed financial research orchestrator |
| Verification | Citation · numeric · temporal · entity · contradiction checks |
| API / UI | FastAPI + Gradio |
| Testing | pytest |
| Code quality | Ruff critical-error gate |
| Deployment | GitHub + Hugging Face Static Space |
| Cost posture | $0 mandatory paid external LLM/API dependency |
| Area | Technology |
|---|---|
| Language | Python 3.12 |
| Local reasoning | Qwen3-8B |
| Model comparison | Qwen3-8B vs Qwen3-14B |
| Dense retrieval | BAAI/bge-m3 |
| Sparse retrieval | BM25 |
| Fusion | Reciprocal Rank Fusion |
| Reranking | BAAI/bge-reranker-v2-m3 |
| Vector search | Qdrant |
| Financial database | DuckDB |
| Filing parsing | SEC filing sections + hierarchical chunks |
| Structured facts | SEC Company Facts / XBRL |
| Knowledge graph | NetworkX |
| API | FastAPI |
| UI | Gradio |
| Testing | pytest |
| Lint | Ruff |
| Automation | GitHub Actions |
| Deployment target | Hugging Face Static Space |
| Packaging | Git + tagged releases |
- SEC EDGAR ingestion under Fair Access constraints
- Multi-company / multi-year 10-K processing
- SEC Company Facts and XBRL ingestion
- Hierarchical filing chunking
- Table extraction
- Deterministic financial fact selection
- BM25 sparse retrieval
- BGE-M3 dense retrieval
- Dense + sparse hybrid retrieval
- Reciprocal Rank Fusion
- BGE reranking
- Query routing across financial, textual, temporal, graph, and mixed questions
- Temporal disclosure comparison
- Provenance-preserving company/risk graph construction
- Local Qwen3 reasoning
- Evidence-first generation
- Claim-level citation attachment checks
- Thinking-output sanitization
- Frozen-core evaluation methodology
- Fresh human-reviewed blind Temporal and Graph evaluation
- FastAPI
- Gradio
- GitHub Actions
- Reproducible JSON/CSV evaluation artifacts
The final local project processed the following measured dataset:
| Property | Measured value |
|---|---|
| Companies | 5 |
| SEC filings | 20 |
| XBRL / Company Facts records | 121,745 facts |
| Parsed filing sections | 97 |
| Hierarchical text chunks | 2,521 |
| Extracted tables | 2,348 |
| Final graph nodes | 223 |
| Final graph edges | 347 |
| Balanced benchmark questions | 200 |
| Benchmark DEV questions | 140 |
| Benchmark TEST questions | 60 |
NVIDIA (NVDA)
Advanced Micro Devices (AMD)
Intel (INTC)
Broadcom (AVGO)
Qualcomm (QCOM)
The cohort is intentionally small enough for reproducible local experimentation while still supporting cross-company financial and risk analysis.
User financial-research question
│
▼
Query understanding
│
▼
Financial Research Router
│
┌──────────┼──────────┬───────────┐
▼ ▼ ▼ ▼
NUMERIC TEXTUAL TEMPORAL GRAPH
│ │ │ │
│ │ │ │
▼ ▼ ▼ ▼
XBRL Tool Hybrid RAG Year-pair Temporal
/ DuckDB │ comparison risk graph
│ │ │ │
│ BM25 + BGE │ Provenance
│ │ │ edge evidence
│ ▼ ▼ │
│ RRF Topic-local │
│ │ comparison │
│ ▼ │ │
│ BGE reranker │ │
└──────────┴───────────┴───────────┘
│
▼
Evidence bundle builder
│
▼
Local Qwen3-8B
│
▼
Strict grounded output guard
│
┌───────┼────────┐
▼ ▼ ▼
Citation Numeric Temporal/
checking checks entity checks
└───────┼────────┘
▼
Evidence-grounded answer
│
▼
FastAPI / Gradio UI
flowchart TD
Q[Financial research question] --> R[Financial Research Router]
R -->|NUMERIC| X[XBRL / DuckDB deterministic tools]
R -->|TEXTUAL| H[Hybrid retrieval]
R -->|TEMPORAL| T[Temporal comparison engine]
R -->|GRAPH| G[Temporal risk graph]
R -->|MIXED| M[Multi-tool orchestration]
H --> B[BM25]
H --> D[BGE-M3]
B --> F[RRF fusion]
D --> F
F --> RR[BGE reranker]
X --> E[Evidence bundle]
RR --> E
T --> E
G --> E
M --> E
E --> L[Local Qwen3-8B]
L --> O[Output sanitizer / strict grounding guard]
O --> C[Citation verifier]
O --> N[Numeric verifier]
O --> TV[Temporal / entity checks]
C --> A[Grounded analyst answer]
N --> A
TV --> A
A --> API[FastAPI]
A --> UI[Gradio]
FilingsGraph does not force every question through one generic vector-search path.
Exact financial question -> deterministic XBRL tool
Narrative filing question -> hybrid text retrieval
Year-over-year risk -> temporal comparison
Cross-company risk -> graph retrieval
Mixed due-diligence task -> routed multi-tool evidence
This separation is the main engineering idea behind the system.
Query
├── BM25 sparse retrieval
└── BGE-M3 dense retrieval
│
▼
Reciprocal Rank Fusion
│
▼
BGE reranker
│
▼
provenance-rich filing evidence
The project evaluates retrieval by category rather than presenting one aggregate score as though all question types were ordinary chunk-retrieval tasks.
| Metric | Hybrid + reranker |
|---|---|
| Recall@5 | 0.4647 |
| Recall@10 | 0.4921 |
| Hit@10 | 0.6905 |
| MRR | 0.4681 |
| nDCG@10 | 0.4556 |
| Mean retrieval latency | ~118.9 ms |
The aggregate includes specialized question classes that are intentionally answered through XBRL, Temporal, or Graph tools rather than pure chunk retrieval.
For ordinary textual filing lookup, the selected system achieved:
| Metric | Result |
|---|---|
| Recall@5 | 1.0000 |
| Recall@10 | 1.0000 |
| Hit@10 | 1.0000 |
| MRR | 0.9444 |
| nDCG@10 | 0.9590 |
This distinction is important: the lower all-category retrieval aggregate is not interpreted as the quality of textual RAG alone.
Exact financial questions are routed to deterministic XBRL tooling rather than left to LLM extraction.
| Metric | Result |
|---|---|
| TEST financial questions | 12 |
| Fact-selection accuracy | 1.0000 |
| Unit accuracy | 1.0000 |
| Fact-ID accuracy | 1.0000 |
Example:
Question:
What was NVDA revenue in FY2025?
Deterministic result:
$130.497 billion
Evidence:
XBRL-NVDA-REVENUE-2025
This architecture prevents a language model from becoming the source of truth for exact financial calculations.
The router decides whether a question should use:
NUMERIC
TEXTUAL
TEMPORAL
GRAPH
MIXED
| Split | Accuracy |
|---|---|
| DEV | 98.41% |
| TEST | 100.00% |
Routing is deterministic/code-controlled; Qwen is used for synthesis rather than being allowed to freely choose arbitrary tools.
Temporal analysis compares topic-local evidence across fiscal years and classifies disclosure changes as:
NEW
REMOVED
UNCHANGED
EXPANDED
REDUCED
Two stages were used:
- a development diagnostic set for engineering iteration;
- a fresh human-reviewed blind set excluded from the previously reviewed Temporal rows.
| Metric | Result |
|---|---|
| Human-reviewed blind rows | 30 |
| Macro F1 | 0.4078 |
| Accuracy | 0.6000 |
The Temporal classifier is the primary measured limitation of the final system. This result is reported directly rather than tuned away after observing the holdout.
FilingsGraph constructs company-to-risk relationships with source provenance.
The graph supports questions such as:
Which selected semiconductor companies share exposure to export controls in FY2025?
Final graph:
| Property | Result |
|---|---|
| Nodes | 223 |
| Edges | 347 |
| Provenance-complete rate | 1.0000 |
| Fresh blind edge precision | 0.9000 |
| Graph path / QA accuracy | 0.7667 |
Graph edge precision was evaluated on a fresh 30-edge human-reviewed sample that excluded the previously reviewed diagnostic edges.
The final orchestrator combines specialized evidence rather than asking the language model to solve every subproblem itself.
Example mixed question:
How did NVDA revenue change from FY2024 to FY2025, and how did its export-controls risk disclosure change over the same period?
The agent can combine:
- deterministic XBRL facts,
- year-specific filing evidence,
- Temporal change evidence,
- graph evidence when relevant,
- and local Qwen synthesis.
Every material generated factual claim must carry one or more supplied citation IDs.
If evidence is insufficient, the system is designed to omit unsupported generated material rather than invent a supporting citation.
Four representative query classes were executed locally:
NUMERIC
TEMPORAL
GRAPH
MIXED
| Metric | Result |
|---|---|
| Citation precision | 1.0000 |
| Claim support rate | 1.0000 |
| Unsupported claim rate | 0.0000 |
| Thinking leak rate | 0.0000 |
| Non-empty answer rate | 1.0000 |
| Mean supported claims / answer | 1.75 |
The deterministic grounding metric measures citation attachment/validity. It is not presented as a fabricated semantic-entailment score.
Two Qwen3 sizes were executed locally.
| Model | Measured runtime |
|---|---|
| Qwen3-8B | ~28.19 s |
| Qwen3-14B | ~158.53 s |
Qwen3-8B was retained as the practical local production model because it fits the local RTX 5090 workflow substantially better. The project does not claim a formal 8B-versus-14B quality win because the comparison did not include a fully judged generation-quality benchmark.
| Version / stage | Purpose | Outcome |
|---|---|---|
| V0 | Dense retrieval baseline | Baseline |
| V1 | Sparse BM25 | Lexical evidence |
| V2 | Hybrid Dense + BM25 | Improved fusion |
| V3 | RRF + metadata filtering | More reliable retrieval |
| V4 | Reranked hybrid retrieval | Selected textual retrieval stack |
| V5 | Structured / Temporal / Graph tooling | Specialized financial intelligence |
| V6 | Routed agentic research + verification | Final system |
| Patch 3 | Output sanitization / citation evaluation | Thinking leak removed |
| Patch 4 | Strict grounded-answer guard | Unsupported claims removed; over-pruning diagnosed |
| Patch 5 | Evidence-first specialized finalization | Non-empty grounded answers restored; graph filtering improved |
The 17-file frozen core checksum gate remained unchanged after the final specialized-layer work.
Final measured ablation:
| Architecture | R@10 | MRR | Numeric Acc. | Temporal F1 | Graph QA | Citation Precision |
|---|---|---|---|---|---|---|
| Dense only | 0.0762 | 0.0606 | — | — | — | — |
| BM25 only | 0.2726 | 0.1716 | — | — | — | — |
| Hybrid | 0.3464 | 0.1289 | — | — | — | — |
| Hybrid + reranker | 0.4921 | 0.4681 | — | — | — | — |
| + structured XBRL | 0.4921 | 0.4681 | 1.0000 | — | — | — |
| + temporal retrieval | 0.4921 | 0.4681 | 1.0000 | 0.4078 | — | — |
| + graph retrieval | 0.4921 | 0.4681 | 1.0000 | 0.4078 | 0.7667 | — |
| Full routed system | 0.4921 | 0.4681 | 1.0000 | 0.4078 | 0.7667 | 1.0000 |
This table shows why FilingsGraph is not evaluated as a single generic retriever: deterministic financial tools, Temporal logic, GraphRAG, and generation verification each add a separate capability.
| Capability | Final measured result |
|---|---|
| Financial fact selection | 100% |
| Financial unit accuracy | 100% |
| Financial fact-ID accuracy | 100% |
| Router TEST accuracy | 100% |
| Textual TEST Recall@10 | 100% |
| Textual TEST MRR | 94.44% |
| Textual TEST nDCG@10 | 95.90% |
| Temporal blind Macro F1 | 40.78% |
| Temporal blind accuracy | 60.00% |
| Graph blind edge precision | 90.00% |
| Graph path / QA accuracy | 76.67% |
| Graph provenance completeness | 100% |
| Citation precision | 100% |
| Claim support rate | 100% |
| Unsupported claim rate | 0% |
| Thinking leak rate | 0% |
| Non-empty answer rate | 100% |
| Final frozen-core check | 17 / 17 unchanged |
The project uses a balanced 200-question benchmark:
200 total
├── 140 DEV
└── 60 TEST
Question classes include:
- exact financial,
- textual,
- temporal,
- graph,
- mixed,
- and no-answer cases.
The original TEST split was executed before the final archive process and therefore is not represented as a pristine never-observed holdout.
After TEST observation:
- the frozen retrieval architecture,
- reranker configuration,
- metadata filtering,
- XBRL selection,
- chunking,
- and router rules
were not tuned against TEST.
Specialized Temporal and Graph final metrics were instead measured using fresh human-reviewed samples excluded from their earlier reviewed development sets.
This limitation is disclosed intentionally.
FilingsGraph applies multiple verification layers:
Generated answer
│
├── Citation ID validity
├── Claim citation attachment
├── Numeric verification
├── Temporal consistency
├── Entity validation
└── Contradiction checks
Security principles include:
- SEC domain allowlisting,
- SEC Fair Access rate limiting,
- no execution of instructions embedded in filing text,
- secrets stored outside Git,
- local
.envexcluded from the repository, - model weights excluded from Git,
- deterministic financial calculations,
- and human-review boundaries for high-stakes conclusions.
See SECURITY.md.
Final local validation confirmed:
- complete pytest suite passing with 79 tests,
- dedicated security suite passing with 6/6 tests,
- frozen 17-file architecture integrity check passing,
- FastAPI startup successful,
GET /healthreturning HTTP 200,GET /companiesreturning HTTP 200,GET /metrics/summaryreturning HTTP 200,POST /researchreturning HTTP 200,- Gradio application launching successfully,
- local Qwen3-8B inference working on the RTX 5090,
- strict grounding returning non-empty answers across representative NUMERIC, TEMPORAL, GRAPH, and MIXED queries.
Run:
uvicorn filingsgraph.api.main:app --host 127.0.0.1 --port 8000Example health endpoint:
curl http://127.0.0.1:8000/healthRun:
python app/gradio_app.pyLocal URL:
http://127.0.0.1:7860
The Gradio interface contains:
Ask Question
Financial Facts
Risk Timeline
Graph Explorer
Evaluation
The full local-model option loads the GPU-backed model/index workflow.
The public portfolio application is available here:
Live application: https://huggingface.co/spaces/anmol-unitmole/filingsgraph-agentic-rag
The Hugging Face deployment is an interactive Static Space that presents the validated architecture, measured results, deployment scope, and representative locally executed workflow replays.
FilingsGraph public portfolio interface showing the project positioning, validated local-runtime status, engineering stack, and key measured outcomes.
Routed financial-research architecture showing deterministic XBRL tooling, hybrid retrieval, Temporal RAG, GraphRAG, Qwen3 synthesis, and claim-level verification.
Final evaluation view presenting structured-financial, textual retrieval, Temporal, Graph, grounding, and frozen-core measurements.
Representative validated workflow replay for Numeric, Temporal, Graph, and Mixed research modes.
The full Agentic RAG system is implemented and validated locally. The static nature of the public Hugging Face application is a hosting/runtime constraint, not a model-readiness limitation.
| Component | Validated local project | Public Hugging Face Static Space |
|---|---|---|
| Qwen3-8B reasoning | Live local GPU inference | Validated precomputed workflow replay |
| BGE-M3 retrieval | Live local embedding pipeline | Architecture / measured results |
| BM25 + RRF + reranker | Live local retrieval stack | Validated retrieval evidence |
| Qdrant | Live local vector index | Not exposed from static hosting |
| DuckDB / XBRL tools | Live deterministic financial tooling | Validated financial results |
| Temporal RAG | Live local comparison workflow | Validated Temporal replay / metrics |
| GraphRAG | Live local graph workflow | Validated graph replay / metrics |
| FastAPI / Gradio | Locally validated interfaces | Not executed server-side |
| User-entered arbitrary research questions | Supported by the full local backend | Not available without a hosted backend |
| Purpose | Full engineering implementation and evaluation | Public portfolio demonstration |
Why is the public demo static?
This Hugging Face Space uses the Static SDK, which does not provide the persistent Python/GPU runtime required by Qwen3-8B, BGE-M3, Qdrant, and the rest of the live backend stack. The model, retrieval pipeline, APIs, and local application are already implemented and validated. With an appropriate GPU-backed deployment environment, the same backend can be hosted as a fully live research application.
The public Space therefore does not pretend that free static hosting is executing Qwen3-8B. It presents genuine validated outputs and measurements while the GitHub repository preserves the complete implementation.
Measured limitations are retained as engineering evidence.
Final fresh blind result:
Macro F1 = 0.4078
Accuracy = 0.6000
Fine-grained EXPANDED versus UNCHANGED and other five-class disclosure distinctions remain the main final-system limitation.
Final measured accuracy:
0.7667
Graph traversal is useful but not perfect.
Fresh blind edge precision:
0.9000
Topic-specific evidence filtering materially improved precision, while a small false-positive rate remains.
An intermediate strict-grounding version produced only a 50% non-empty answer rate.
The final evidence-first fallback restored:
Non-empty answer rate = 1.0
Unsupported claim rate = 0.0
Development Macro F1 reached approximately 0.599, while the fresh blind score was 0.408.
This gap is reported rather than hidden through post-hoc tuning.
Important evidence artifacts include:
reports/final/summary.json
reports/final/
reports/experiments/grounding_evaluation.json
reports/experiments/patch4_diagnostic.json
reports/gold/temporal_review.csv
reports/gold/temporal_review_final.csv
reports/gold/graph_qa_review.csv
reports/gold/graph_edge_review.csv
reports/gold/graph_edge_review_final.csv
reports/temporal/
reports/graph/
reports/baseline/v6_final_frozen/
reports/baseline/v6_specialized_gold_baseline/
Generated model caches, raw SEC downloads, local Qdrant state, .venv, local databases, secrets, and other machine-specific artifacts are intentionally excluded from the public repository.
filingsgraph-agentic-rag/
├── .github/
│ └── workflows/
├── app/
│ └── gradio_app.py
├── configs/
├── data/
│ ├── graph/
│ ├── processed/
│ └── ...
├── docs/
├── reports/
│ ├── baseline/
│ ├── experiments/
│ ├── final/
│ ├── gold/
│ ├── graph/
│ └── temporal/
├── scripts/
├── src/
│ └── filingsgraph/
│ ├── agents/
│ ├── api/
│ ├── embeddings/
│ ├── graph/
│ ├── llm/
│ ├── retrieval/
│ ├── security/
│ ├── temporal/
│ ├── verification/
│ └── xbrl/
├── tests/
├── .env.example
├── BENCHMARK.md
├── DATA_SOURCES.md
├── MODEL_AND_DATA_LICENSES.md
├── SECURITY.md
├── VALIDATION.md
├── LICENSE
└── README.md
git clone https://github.com/unit-mole/filingsgraph-agentic-rag.git
cd filingsgraph-agentic-ragpython -m venv .venv
call .venv\Scripts\activate.bat
python -m pip install --upgrade pippython -m pip install -e ".[dev]"Copy:
.env.example -> .env
Provide a valid SEC Fair Access identity locally.
Never commit .env.
pytest -q
pytest tests\security -q
python -m scripts.check_v6_freezepython -m ruff check src scripts app tests --select E9,F63,F7,F82uvicorn filingsgraph.api.main:app --host 127.0.0.1 --port 8000python app\gradio_app.pyGitHub Actions validates deterministic checks without loading the local multi-GB Qwen model.
Workflows include:
tests
security
lint
The lint workflow intentionally gates critical Python/Ruff correctness classes:
E9
F63
F7
F82
rather than failing the entire public release on legacy style-only issues such as import sorting.
Formatting and broader style cleanup can be applied incrementally without weakening the deterministic test and security gates.
| Document | Purpose |
|---|---|
| BENCHMARK.md | Benchmark design and evaluation policy |
| DATA_SOURCES.md | SEC/data provenance |
| MODEL_AND_DATA_LICENSES.md | Model/data licensing notes |
| SECURITY.md | Security posture |
| VALIDATION.md | Validation scope and measured evidence |
| reports/final/summary.json | Final machine-readable results |
- The initial cohort contains five semiconductor companies and 20 filings.
- The final fresh blind Temporal Macro F1 is 0.4078; temporal five-class disclosure change detection remains the main weakness.
- Graph path / QA accuracy is 0.7667, so graph reasoning is not treated as perfect.
- Fresh blind graph edge precision is 0.9000, leaving a small residual false-positive rate.
- Grounding metrics verify citation attachment/validity; they are not presented as a semantic-entailment benchmark.
- The original TEST split was observed before final archival; core retrieval components were frozen afterward and not post-hoc tuned against those results.
- Qwen3-14B was substantially slower locally and was not selected for the default runtime.
- The public Hugging Face Static Space does not execute the full local Qwen3-8B/BGE-M3/Qdrant backend; it transparently presents validated local outputs because the current static hosting environment does not provide the required GPU-backed Python runtime.
- This project supports research and portfolio demonstration; it is not investment advice.
Potential extensions include:
- larger multi-sector SEC cohorts,
- quarterly 10-Q temporal analysis,
- richer XBRL concept normalization,
- improved five-class Temporal disclosure classification,
- semantic entailment evaluation for generated claims,
- stronger graph relation extraction,
- Neo4j-backed graph exploration,
- portfolio-level multi-company comparison,
- automated filing-event alerts,
- authenticated private research workspaces,
- hosted GPU inference,
- and richer analyst dashboards.
- Generative AI
- Retrieval-Augmented Generation
- Agentic AI
- Financial RAG
- SEC EDGAR
- XBRL
- Financial data engineering
- Hybrid retrieval
- BGE-M3
- BM25
- Reciprocal Rank Fusion
- BGE reranking
- Qdrant
- DuckDB
- Temporal RAG
- GraphRAG
- NetworkX
- Knowledge graphs
- Query routing
- Deterministic tool use
- Qwen3 local inference
- Evidence-grounded generation
- Citation verification
- Numeric verification
- Failure analysis
- Frozen-evaluation methodology
- Human-reviewed gold evaluation
- FastAPI
- Gradio
- pytest
- Ruff
- GitHub Actions
- Hugging Face deployment
- Portfolio-focused AI engineering
Performance claims in this README come from locally generated evaluation artifacts.
The project distinguishes:
DEV diagnostics
frozen/observed TEST results
fresh human-reviewed specialized blind results
No final Temporal or Graph claim is presented as stronger than the measured evidence supports.
The public application will be described according to what it actually executes on the deployed hardware.
One-line description: Temporal financial due-diligence Agentic RAG system combining SEC filings, XBRL, hybrid retrieval, Temporal RAG, GraphRAG, local Qwen3 reasoning, and claim-level evidence verification.
Pinned repository description: Portfolio-grade Agentic Financial RAG project with SEC EDGAR/XBRL ingestion, BGE-M3 + BM25/RRF hybrid retrieval, deterministic financial tools, temporal disclosure intelligence, provenance-preserving GraphRAG, local Qwen3-8B synthesis, fresh human-reviewed evaluation, FastAPI/Gradio interfaces, and GitHub CI.
FilingsGraph is not:
“Embed SEC filings and ask an LLM financial questions.”
It is:
SEC Filing Ingestion
+ XBRL Financial Facts
+ Hierarchical Parsing
+ Table Extraction
+ Sparse Retrieval
+ Dense Retrieval
+ Reciprocal Rank Fusion
+ Reranking
+ Deterministic Financial Tools
+ Temporal Disclosure Comparison
+ Provenance-Preserving Risk Graph
+ Routed Agentic Research
+ Local Qwen3 Synthesis
+ Strict Citation Grounding
+ Multi-Layer Verification
+ Human-Reviewed Evaluation
The central engineering question is not simply whether an LLM can summarize a filing.
It is whether a financial research system can route each question to the right evidence source, calculate exact facts deterministically, compare changing disclosure language over time, connect cross-company risk evidence, and prove which filing evidence supports the final answer.
Project code: MIT.
SEC filings are public regulatory disclosures. Model weights and third-party dependencies retain their own licenses. See MODEL_AND_DATA_LICENSES.md.
Anmol Tripathi
Quality Data Scientist building portfolio projects across Data Science, Machine Learning, Applied AI, Generative AI, Agentic RAG, Natural Language Processing, Analytics Engineering, and Quality Analytics.



