Professional teaching/demo project for explaining how Retrieval-Augmented Generation (RAG) works with a single PDF.
- Backend: FastAPI + LangChain + FAISS
- Frontend: Streamlit
- Models: OpenAI embeddings + OpenAI chat model
- Architecture Overview
- Project Structure
- Configuration
- Quick Start
- How the RAG Pipeline Works
- API Reference
- Frontend Behavior
- Teaching Flow
- Troubleshooting
The system is split into two services:
- Backend (FastAPI)
- Accepts PDF uploads
- Chunks and embeds the document
- Stores vectors in FAISS
- Answers questions both with and without retrieval
- Frontend (Streamlit)
- UI for upload + question asking
- Visual comparison:
- Without RAG (general model answer)
- With RAG (grounded answer from retrieved chunks)
- Displays retrieved chunks with source/page/similarity score
Data flow:
PDF -> Chunking -> Embeddings -> FAISS -> Retrieval -> Prompt -> LLM -> Answer
rag-pdf-teaching-demo/
├── backend/
│ ├── main.py
│ ├── config.py
│ ├── rag_service.py
│ ├── schemas.py
│ └── utils.py
├── frontend/
│ └── app.py
├── data/
│ └── uploads/
├── storage/
│ └── faiss_index/
├── .env
├── requirements.txt
└── README.md
Runtime settings are defined in backend/config.py:
CHUNK_SIZE = 300CHUNK_OVERLAP = 50TOP_K = 4EMBEDDING_MODEL = "text-embedding-3-small"LLM_MODEL = "gpt-4.1-mini"TEMPERATURE = 0UPLOAD_DIR = "data/uploads"FAISS_INDEX_PATH = "storage/faiss_index"
Environment variable required in .env:
OPENAI_API_KEY=your_real_openai_api_keyUse the following versions for a stable run.
Python 3.13.12
fastapi==0.136.0uvicorn==0.45.0streamlit==1.57.0requests==2.32.5python-multipart==0.0.26python-dotenv==1.2.1langchain==1.2.15langchain-openai==1.1.16langchain-community==0.4.1langchain-text-splitters==1.1.2faiss-cpu==1.13.2pypdf==6.10.2
python --version
pip show fastapi uvicorn streamlit requests python-multipart python-dotenv langchain langchain-openai langchain-community langchain-text-splitters faiss-cpu pypdfcd "D:\Workshop Material\Jupyter Notebooks\RAGApp\rag-pdf-teaching-demo"
pip install -r requirements.txtcd "D:\Workshop Material\Jupyter Notebooks\RAGApp\rag-pdf-teaching-demo"
uvicorn backend.main:app --reloadExpected backend URL:
http://localhost:8000
Health check:
cd "D:\Workshop Material\Jupyter Notebooks\RAGApp\rag-pdf-teaching-demo"
streamlit run frontend/app.pyExpected frontend URL:
http://localhost:8501
- Upload PDF from frontend.
- Backend saves file in
data/uploads. PyPDFLoaderreads pages into documents.RecursiveCharacterTextSplittercreates chunks.OpenAIEmbeddingsconverts chunks to vectors.FAISS.from_documents(...)builds vector index.- Index is persisted under
storage/faiss_index.
For each question, backend runs two paths:
- Without RAG: send question directly to LLM.
- With RAG:
- run
similarity_search_with_score(question, k=TOP_K) - take top chunks + their similarity scores
- build grounded prompt from retrieved chunks
- ask LLM to answer using only retrieved context
- run
Each retrieved chunk includes a score, so students can see ranking quality and understand:
"The retriever ranks chunks by semantic similarity."
Response:
{"status":"ok"}Request:
multipart/form-data- field:
file(PDF)
Response example:
{
"filename": "Employee-Handbook.pdf",
"pages": 34,
"chunks": 242,
"message": "PDF indexed successfully"
}Request body:
{"question":"What is the attendance policy?"}Response fields:
questionno_rag_answerrag_answerretrieved_chunks[]:contentsourcepagescore
The Streamlit app contains:
- RAG Flow diagram
- Upload + index section
- Comparison section
- Without RAG shown with
st.error(...) - With RAG shown with
st.success(...)
- Without RAG shown with
- Retrieved chunks section
- Displays source/page/content/
Similarity Score
- Displays source/page/content/
Backend URL used by frontend:
BACKEND_URL = "http://localhost:8000"infrontend/app.py
Recommended classroom order:
- Upload PDF and build index.
- Explain chunking and embeddings.
- Explain FAISS retrieval and ranking.
- Ask a question.
- Compare no-RAG vs RAG answers.
- Open retrieved chunks and inspect scores.
- Emphasize grounding and traceability.
-
Frontend cannot call backend
- verify backend is running on
http://localhost:8000 - verify
BACKEND_URLinfrontend/app.py
- verify backend is running on
-
No useful retrieved chunks
- re-index after changing chunk settings
- adjust
TOP_Kfor broader context
-
Dependency issues on Windows
- activate your environment
- run
pip install -r requirements.txtagain
The LLM does not read the PDF directly.
FAISS retrieves the most semantically relevant chunks first.
The LLM generates the answer using retrieved context.