Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Universal RAG Engine (AdamantiumAI)

An enterprise-grade Retrieval-Augmented Generation (RAG) pipeline designed specifically for dense, complex documentation like engineering spec books, legal contracts, manufacturing manuals, and codebases. Watch the 90 second demo here: DEMO

The Problem vs. The Solution

Out-of-the-box RAG pipelines often fail in professional environments. They get stuck routing users to Table of Contents pages, they strip out critical technical syntax, and they suffer from "Context Amnesia" when a legal clause is separated from the contract's title page.

The Universal RAG Engine solves this:

Intelligent Reranking: A two-tier ranking system (Heuristic Header Boost + LLM Reranking) ensures the engine returns the actual definition of a clause or code, completely ignoring Table of Contents and Index references.

Contextual Source Injection: The reranker reads the file provenance of every snippet, ensuring that "Page 5" of an agreement is properly linked to the company named on "Page 1".

Hybrid Search (BM25 + Dense Vectors): Preserves strict technical syntax (like str in Python or 33 09 01 in construction specs) while maintaining semantic understanding for broader questions.

Key Features

Multi-Format Ingestion: Seamlessly processes PDFs (with OCR fallback for scanned images), CSVs, JSON, Excel, Word, HTML, and raw code files (.py, .js, .cpp, etc.).

Visual Source Verification: Converts PDF pages to images and stores them in secure AWS S3 buckets. The UI provides users with secure, short-lived pre-signed URLs to view the exact original document page.

Multi-Tenant Architecture: Built from the ground up for agencies/B2B providers. Data is strictly isolated using Pinecone Namespaces and distinct S3 folders.

FastAPI Backend: Secure, rate-limited REST API to handle queries concurrently.

Streamlit Admin Console: A sleek, dark-mode chat interface with session management, source expanders, and transcript downloads.

System Architecture

Ingestion (ingest.py / main.py):
Files are parsed, OCR is applied if necessary, and documents are chunked using LangChain.
Raw text and generated images are uploaded to AWS S3.
Text is embedded (OpenAI text-embedding-3-small) and sparse-encoded (BM25) before being upserted to Pinecone Serverless.

Retrieval (query.py):
User queries are rewritten for context and run through a Pinecone Hybrid Search.
Results pass through a custom LLM Reranker (gpt-4o-mini) to filter out noise (TOCs/Glossaries) and prioritize substantive clauses.

Serving (server.py):
FastAPI serves the query via Server-Sent Events (Streaming) and generates secure AWS presigned URLs for source images.

Client (app.py):
Streamlit frontend handles authentication, client-ID routing, and rendering markdown/images.

Tech Stack

Backend: Python, FastAPI, Uvicorn
Frontend: Streamlit
AI/ML: LangChain, OpenAI (GPT-4o-mini, Embeddings)
Vector Database: Pinecone (Serverless)
Object Storage: AWS S3, Boto3
Document Processing: PDFPlumber, pdf2image, PyTesseract (OCR)

Author

Josh Gilstrap
LinkedIn
jgilstrap23@gmail.com

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages