A document-grounded question-answering application built with Retrieval-Augmented Generation (RAG). It allows users to ask questions about indexed PDF documents and receive answers grounded in the document content.
Large documents can contain useful information that is difficult to find manually.
This project provides a simple way to ask natural-language questions about a document and receive an answer based onldocument-y on the information available in that document.
The application:
- Extracts text from PDF documents
- Splits documents into smaller chunks
- Generates semantic embeddings
- Stores embeddings using FAISS
- Retrieves the most relevant document chunks
- Sends retrieved context to an LLM through the Groq API
- Generates document-grounded answers
- Displays relevant source pages
- Refuses to answer when the requested information is not found in the document
PDF DOCUMENT
β
βΌ
βββββββββββββββββββ
β Document Loader β
ββββββββββ¬βββββββββ
β
βΌ
Text Chunking
β
βΌ
Sentence Transformers
Embeddings
β
βΌ
βββββββββββββββββββ
β FAISS β
β Vector Store β
ββββββββββ¬βββββββββ
β
βΌ
Persistent Store
index.faiss
chunks.json
β
βΌ
User Question
β
βΌ
Query Embedding
β
βΌ
FAISS Search
β
βΌ
Top-k Relevant
Chunks
β
βΌ
ββββββββββββββββ
β Groq API β
β GPT-OSS-20B β
ββββββββ¬ββββββββ
β
βΌ
Grounded Answer
β
βΌ
Streamlit UI
The PDF document is loaded and its text is extracted.
PDF β Extracted Text
The extracted text is divided into smaller chunks for efficient retrieval.
Document β Text Chunks
Each text chunk is converted into a vector using the all-MiniLM-L6-v2 sentence-transformer model.
Text Chunk β 384-dimensional embedding
The embeddings are indexed using FAISS.
The vector store is persisted locally as:
index.faiss
chunks.json
When a user asks a question:
Question
β
Question Embedding
β
FAISS Similarity Search
β
Top-k Relevant Chunks
The retrieved document chunks are provided as context to GPT-OSS-20B through the Groq API.
The model is instructed to answer only using the retrieved document context.
Question + Retrieved Context
β
Groq API
β
GPT-OSS-20B
β
Answer
| Technology | Purpose |
|---|---|
| Python | Core programming language |
| Streamlit | Web application interface |
| PyPDF | PDF text extraction |
| Sentence Transformers | Text embeddings |
| FAISS | Vector similarity search |
| Groq API | LLM inference |
| GPT-OSS-20B | Answer generation |
| python-dotenv | Environment variable management |
| NumPy | Numerical operations |
- PDF document question answering
- Semantic vector search
- Sentence Transformer embeddings
- Persistent FAISS vector store
- Groq API-powered answer generation
- Source page display
- Out-of-document question handling
- Streamlit interface
The RAG system was tested with six representative questions.
| Test Case | Expected Behavior | Result |
|---|---|---|
| What is the Internet of Things? | Answer | β PASS |
| What are the main challenges of IoT? | Answer | β PASS |
| How is IoT used in smart grids? | Answer | β PASS |
| What is GPS used for in smartphones? | Answer | β PASS |
| How do airplanes fly? | Refuse | β PASS |
| What is photosynthesis? | Refuse | β PASS |
Result: 6/6 expected behavioral outcomes passed.
Note: This is a small functional evaluation set and should not be interpreted as a general accuracy benchmark.
The application retrieves relevant information from the indexed document and generates a grounded answer using the Groq API.
The system refuses to answer when the requested information is not available in the indexed document.
document-intelligence-rag/
β
βββ app.py
βββ evaluate_rag.py
βββ requirements.txt
βββ README.md
βββ .gitignore
β
βββ data/
β βββ documents/
β βββ vector_store/
β
βββ screenshot/
β βββ rag-answer.png
β βββ out-of-document.png
β
βββ src/
β βββ __init__.py
β βββ document_loader.py
β βββ embeddings.py
β βββ rag_pipeline.py
β βββ text_splitter.py
β βββ vector_store.py
β
βββ test_loader.py
βββ test_groq.py
βββ test_groq_models.py
git clone https://github.com/pratiksha18r/document-intelligence-rag.git
cd document-intelligence-ragpython -m venv .venvOn Windows PowerShell:
.venv\Scripts\activatepip install -r requirements.txtCreate a .env file in the project root:
GROQ_API_KEY=your_groq_api_key
Never commit the .env file to GitHub.
The FAISS vector store is generated locally and is excluded from GitHub.
After setting up the project, run the ingestion pipeline:
python test_loader.pyThis processes the document, generates embeddings, creates the FAISS index, and saves the vector store locally.
The generated files are:
data/vector_store/
βββ index.faiss
βββ chunks.json
After creating the vector store, start the Streamlit application:
streamlit run app.pyThe application will open in your browser.
Sensitive and generated files are excluded from Git using .gitignore.
The Groq API key is loaded through an environment variable and is not stored directly in the source code.
- Currently designed around a locally indexed document.
- The evaluation dataset is intentionally small.
- Retrieval currently uses a fixed top-k value.
- The application does not yet provide an in-app document upload workflow.
- Retrieval quality depends on chunking and embedding choices.
- Multiple-document support
- In-app PDF upload and indexing
- Improved retrieval evaluation
- Similarity thresholding
- Reranking retrieved chunks
- Conversation history
- Streaming LLM responses
- More detailed source citations
- Automated test suite
- Cloud deployment
Pratiksha Bodkhe
MCA | Artificial Intelligence & Machine Learning Interested in Artificial Intelligence, Machine Learning, Data Science, and Software Development.

