B.Tech Final Year Engineering Project
A full-stack AI platform combining Computer Vision (Salesforce BLIP) and Large Language Models (Qwen2.5-3B via Ollama) to convert images into context-aware captions, social media posts, multi-language descriptions, and overlay-captioned visuals.
This project implements an end-to-end multimodal content generation system. Users can upload an image, which is first processed by Salesforce BLIP (a state-of-the-art Vision-Language model running locally via PyTorch/Transformers) to extract a raw visual description. This description is then passed to Qwen2.5-3B-Instruct (served via Ollama's local API) to synthesize rich, context-aware content adapted for target platforms (Instagram, LinkedIn, X, Facebook), writing styles (Creative, Marketing, Humorous, etc.), and languages (English, Hindi, Marathi).
Finally, the generated text can be dynamically overlaid onto the original image using Pillow image processing to produce a shareable captioned output image.
+---------------------------------------+
| Vite + React Frontend (UI) |
| - Image Upload & Interactive Preview |
| - Style, Language & Platform Controls |
+-------------------+-------------------+
|
| HTTP REST API
v
+---------------------------------------+
| FastAPI Python Backend |
| - Async Route Controllers |
| - Input Validation & Pydantic Models |
+---------+-------------------+---------+
| |
+--------------------+ +--------------------+
| |
v v
+-----------------------+ +-----------------------+
| BLIP Vision Model | | Ollama Service API |
| - Salesforce BLIP Base| | - Qwen2.5-3B-Instruct |
| - PyTorch Engine | | - Prompt Engineering |
+-----------+-----------+ +-----------+-----------+
| |
| Visual Description | Rewritten Caption
+----------------------------------+--------------------------+
|
v
+-----------------------+
| Pillow Image Service |
| (Caption Text Overlay)|
+-----------+-----------+
|
v
+-----------------------+
| SQLite Database |
| (History & Metadata) |
+-----------------------+
- Framework: React.js (v18) initialized with Vite
- Styling: Modern Vanilla CSS with CSS custom properties (Variables), dark mode support, glassmorphism, and responsive layouts
- HTTP Client: Axios
- Framework: FastAPI (Python 3.10+)
- Server: Uvicorn (ASGI)
- Validation: Pydantic v2
- Image Processing: Pillow (PIL), OpenCV
- Vision-Language Model:
Salesforce/blip-image-captioning-base(PyTorch + Hugging Face Transformers) - Local LLM Engine: Ollama running
qwen2.5:3b(Qwen2.5-3B-Instruct)
- Storage: SQLite3 (persisting caption history, parameters, and generated image paths)
- Frameworks: Unsloth, Hugging Face PEFT (LoRA/QLoRA)
AI-Image-Captioning/
│
├── frontend/ # React + Vite Frontend
│ ├── src/
│ │ ├── components/ # Reusable UI components
│ │ ├── pages/ # Main page layouts
│ │ ├── services/ # API integration layer (Axios)
│ │ ├── App.jsx # Main app container
│ │ ├── main.jsx # React entry point
│ │ └── index.css # Design system & styles
│ ├── package.json
│ └── vite.config.js
│
├── backend/ # FastAPI Backend
│ ├── app/
│ │ ├── main.py # FastAPI entry point
│ │ ├── config.py # App configuration & constants
│ │ ├── api/ # Router endpoints (caption, content, image, history)
│ │ ├── services/ # Core AI & Image services (blip, ollama, image)
│ │ ├── models/ # Schemas & SQLite Database models
│ │ └── utils/ # Helper utilities (file handling, image formatting)
│ ├── uploads/ # Temporary uploaded user images
│ ├── outputs/ # Rendered captioned images
│ ├── requirements.txt # Python dependencies
│ └── .env # Environment config
│
├── training/ # LLM Fine-Tuning module (Unsloth / LoRA)
│ ├── dataset/ # Custom dataset storage
│ ├── scripts/ # Training scripts
│ └── notebooks/ # Jupyter notebooks for experimentation
│
├── data/ # Local SQLite database file storage
├── README.md # Project documentation
└── .gitignore # Version control exclusion rules
- Python: 3.10 or higher
- Node.js: v18.0 or higher
- Ollama: Installed locally (https://ollama.com)
Pull the Qwen 2.5 3B Instruct model using Ollama:
ollama pull qwen2.5:3bEnsure the Ollama service is running in the background (default port 11434).
- Open a terminal and navigate to the backend folder:
cd backend - Activate your virtual environment (e.g. the existing
devenvironment or create one):# On Windows (PowerShell): ..\dev\Scripts\Activate.ps1
- Install dependencies:
pip install -r requirements.txt
- Run the FastAPI development server:
uvicorn app.main:app --reload --port 8000
- Open a second terminal and navigate to the frontend folder:
cd frontend - Install Node dependencies:
npm install
- Start the Vite dev server:
npm run dev
- Access the web application at
http://localhost:3000.
The development of this project follows a 15-phase incremental workflow:
- Phase 1: Project Structure & Configuration Setup (Current)
- Phase 2: FastAPI Server Skeleton & Health Endpoint
- Phase 3: Dependency Installation & Validation
- Phase 4: BLIP Vision Service Implementation
- Phase 5: Image Captioning API Endpoint (
/api/caption) - Phase 6: API Testing via Swagger UI (
http://localhost:8000/docs) - Phase 7: Ollama Integration Service (
ollama_service.py) - Phase 8: Content Generation API Endpoint (
/api/generate) - Phase 9: Combined End-to-End Pipeline Endpoint (
/api/generate-from-image) - Phase 10: React Frontend Components & State Integration
- Phase 11: Image Caption Overlay Service (Pillow)
- Phase 12: History & Persistence Engine (SQLite)
- Phase 13: Modern UI Polish & Micro-animations
- Phase 14: System Verification & End-to-End Testing
- Phase 15: Model Fine-Tuning Pipeline Setup (
training/)
This project is developed for B.Tech Final Year Academic Dissertation & Capstone Demonstration. All rights reserved.