Skip to content

Latest commit

Β 

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

Document Intelligence RAG

A document-grounded question-answering application built with Retrieval-Augmented Generation (RAG). It allows users to ask questions about indexed PDF documents and receive answers grounded in the document content.

Project Overview

Large documents can contain useful information that is difficult to find manually.

This project provides a simple way to ask natural-language questions about a document and receive an answer based onldocument-y on the information available in that document.

The application:

  • Extracts text from PDF documents
  • Splits documents into smaller chunks
  • Generates semantic embeddings
  • Stores embeddings using FAISS
  • Retrieves the most relevant document chunks
  • Sends retrieved context to an LLM through the Groq API
  • Generates document-grounded answers
  • Displays relevant source pages
  • Refuses to answer when the requested information is not found in the document

Architecture

                    PDF DOCUMENT
                         β”‚
                         β–Ό
                β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                β”‚ Document Loader β”‚
                β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                         β”‚
                         β–Ό
                   Text Chunking
                         β”‚
                         β–Ό
              Sentence Transformers
                   Embeddings
                         β”‚
                         β–Ό
                β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                β”‚      FAISS      β”‚
                β”‚  Vector Store   β”‚
                β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                         β”‚
                         β–Ό
                  Persistent Store
                  index.faiss
                  chunks.json
                         β”‚
                         β–Ό
                  User Question
                         β”‚
                         β–Ό
                  Query Embedding
                         β”‚
                         β–Ό
                    FAISS Search
                         β”‚
                         β–Ό
                 Top-k Relevant
                    Chunks
                         β”‚
                         β–Ό
                 β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                 β”‚   Groq API   β”‚
                 β”‚  GPT-OSS-20B β”‚
                 β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜
                        β”‚
                        β–Ό
                  Grounded Answer
                        β”‚
                        β–Ό
                   Streamlit UI


RAG Workflow

1. Document Ingestion

The PDF document is loaded and its text is extracted.

PDF β†’ Extracted Text

2. Text Chunking

The extracted text is divided into smaller chunks for efficient retrieval.

Document β†’ Text Chunks

3. Embedding Generation

Each text chunk is converted into a vector using the all-MiniLM-L6-v2 sentence-transformer model.

Text Chunk β†’ 384-dimensional embedding

4. Vector Storage

The embeddings are indexed using FAISS.

The vector store is persisted locally as:

index.faiss
chunks.json

5. Query Retrieval

When a user asks a question:

Question
   ↓
Question Embedding
   ↓
FAISS Similarity Search
   ↓
Top-k Relevant Chunks

6. Answer Generation

The retrieved document chunks are provided as context to GPT-OSS-20B through the Groq API.

The model is instructed to answer only using the retrieved document context.

Question + Retrieved Context
            ↓
         Groq API
            ↓
       GPT-OSS-20B
            ↓
          Answer

Tech Stack

Technology Purpose
Python Core programming language
Streamlit Web application interface
PyPDF PDF text extraction
Sentence Transformers Text embeddings
FAISS Vector similarity search
Groq API LLM inference
GPT-OSS-20B Answer generation
python-dotenv Environment variable management
NumPy Numerical operations

Features

  • PDF document question answering
  • Semantic vector search
  • Sentence Transformer embeddings
  • Persistent FAISS vector store
  • Groq API-powered answer generation
  • Source page display
  • Out-of-document question handling
  • Streamlit interface

Evaluation

The RAG system was tested with six representative questions.

Test Case Expected Behavior Result
What is the Internet of Things? Answer βœ… PASS
What are the main challenges of IoT? Answer βœ… PASS
How is IoT used in smart grids? Answer βœ… PASS
What is GPS used for in smartphones? Answer βœ… PASS
How do airplanes fly? Refuse βœ… PASS
What is photosynthesis? Refuse βœ… PASS

Result: 6/6 expected behavioral outcomes passed.

Note: This is a small functional evaluation set and should not be interpreted as a general accuracy benchmark.

Application Screenshots

Document-Grounded Answer

The application retrieves relevant information from the indexed document and generates a grounded answer using the Groq API.

RAG Answer

Out-of-Document Question

The system refuses to answer when the requested information is not available in the indexed document.

Out-of-Document Question

Project Structure

document-intelligence-rag/
β”‚
β”œβ”€β”€ app.py
β”œβ”€β”€ evaluate_rag.py
β”œβ”€β”€ requirements.txt
β”œβ”€β”€ README.md
β”œβ”€β”€ .gitignore
β”‚
β”œβ”€β”€ data/
β”‚   β”œβ”€β”€ documents/
β”‚   └── vector_store/
β”‚
β”œβ”€β”€ screenshot/
β”‚   β”œβ”€β”€ rag-answer.png
β”‚   └── out-of-document.png
β”‚
β”œβ”€β”€ src/
β”‚   β”œβ”€β”€ __init__.py
β”‚   β”œβ”€β”€ document_loader.py
β”‚   β”œβ”€β”€ embeddings.py
β”‚   β”œβ”€β”€ rag_pipeline.py
β”‚   β”œβ”€β”€ text_splitter.py
β”‚   └── vector_store.py
β”‚
β”œβ”€β”€ test_loader.py
β”œβ”€β”€ test_groq.py
└── test_groq_models.py

Installation

1. Clone the repository

git clone https://github.com/pratiksha18r/document-intelligence-rag.git
cd document-intelligence-rag

2. Create a virtual environment

python -m venv .venv

On Windows PowerShell:

.venv\Scripts\activate

3. Install dependencies

pip install -r requirements.txt

4. Configure the Groq API key

Create a .env file in the project root:

GROQ_API_KEY=your_groq_api_key

Never commit the .env file to GitHub.

Prepare the Vector Store

The FAISS vector store is generated locally and is excluded from GitHub.

After setting up the project, run the ingestion pipeline:

python test_loader.py

This processes the document, generates embeddings, creates the FAISS index, and saves the vector store locally.

The generated files are:

data/vector_store/
β”œβ”€β”€ index.faiss
└── chunks.json

Run the Application

After creating the vector store, start the Streamlit application:

streamlit run app.py

The application will open in your browser.

Security

Sensitive and generated files are excluded from Git using .gitignore.

The Groq API key is loaded through an environment variable and is not stored directly in the source code.

Current Limitations

  • Currently designed around a locally indexed document.
  • The evaluation dataset is intentionally small.
  • Retrieval currently uses a fixed top-k value.
  • The application does not yet provide an in-app document upload workflow.
  • Retrieval quality depends on chunking and embedding choices.

Future Improvements

  • Multiple-document support
  • In-app PDF upload and indexing
  • Improved retrieval evaluation
  • Similarity thresholding
  • Reranking retrieved chunks
  • Conversation history
  • Streaming LLM responses
  • More detailed source citations
  • Automated test suite
  • Cloud deployment

Author

Pratiksha Bodkhe

MCA | Artificial Intelligence & Machine Learning Interested in Artificial Intelligence, Machine Learning, Data Science, and Software Development.

About

Document-grounded RAG application using FAISS, Sentence Transformers, Groq API and Streamlit.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Contributors

Languages