AI-powered handwritten answer sheet evaluation using Google ADK 2.0, Gemini 2.5 Flash, Agent Skills, and Google Drive MCP.
This repository is prepared for Kaggle evaluation and includes pre-packaged reference materials and answer sheets to demonstrate ScribeAI's capabilities across different evaluation scenarios.
- OBJECT ORIENTED PROGRAMMING USING C++.pdf: The official C++ OOP exam question paper.
- Answer booklet.pdf: The printed answer booklet template containing student instructions and blank pages where students write their exams.
The grading engine uses a structured C++ OOP marking scheme located at app/marking_scheme.json.
Important
AI Grading Axiom: The grading capability of the AI is directly proportional to how detailed and robust the marking scheme JSON is. Detailed concept lists lead to extremely accurate, objective, and consistent evaluations.
The sample/ directory contains 5 distinct student answer sheets for testing:
sample/Answersheet_Robin_Danie_CS001.pdfsample/Answersheet_clearFail_CS002.pdfsample/Answersheet_average_CS003.pdfsample/Answersheet_PromptInjection_CS004.pdfsample/Answersheet_worstHandwriting_CS005.pdf
Follow these steps to set up the ScribeAI grading pipeline on your local machine:
git clone <repository-url>
cd scribeaiScribeAI uses uv for python environment and dependency management. Run the following command to sync and install the environment:
uv syncCreate a .env file in the project root folder and provide your Gemini API key:
GEMINI_API_KEY=your_api_key_hereWarning
API Key Safety: Do not commit your .env file to Git. Every evaluator must use their own API key.
ScribeAI can automatically upload results and reports to Google Drive. To configure this optional integration:
- Create Google Cloud Project: Go to the Google Cloud Console, create a project, and enable the Google Drive API.
- Setup OAuth Consent & Client: Configure the OAuth Consent Screen and create client credentials for a Desktop app.
- Save Credentials: Download the client credentials JSON file, rename it to
credentials.json, and place it in the root folder of this project. - Install Local Dependencies: On Windows, to prevent temporary cache issues, install the MCP server locally in the workspace:
npm install @modelcontextprotocol/server-gdrive
- Authenticate: Upon your first pipeline run, the Google Drive MCP server will prompt you to authenticate via your web browser.
Once setup is complete, you can run evaluations in two ways:
Start the local server:
uv run adk web --port 8000Open http://localhost:8000 in your browser. From here, you can:
- Chat with ScribeAI, upload any PDF/image answer sheet, and watch it grade in real-time.
- Batch Grade All Papers at Once: Simply send the word
samplein the Web UI chat. ScribeAI will automatically detect it as a folder, grade all 5 papers sequentially, and save results.
If you want to run the pipeline programmatically in your terminal:
-
Evaluate a Single Paper:
uv run python scratch/test_run.py
(Grades Robin Danie's paper CS001)
-
Batch Evaluate All 5 Sample Papers:
uv run python scratch/test_batch.py
(Grades all student papers in the
sample/folder sequentially)
Once an evaluation runs, ScribeAI stores results in:
spreadsheets/results.xlsx(Excel results logging)reports/report_<student_id>.md(Question-by-concept markdown grading reports)batch_reports/batch_<timestamp>.md(Master run batch summary, when batch processed)
ScribeAI is a multi-agent system that automates the evaluation of handwritten student answer sheets.
The system reads scanned answer sheets, evaluates answers against a professor-defined marking scheme, performs confidence-based quality control, generates detailed feedback, and stores results automatically.
The goal is not to replace professors, but to reduce repetitive grading workload while ensuring that uncertain cases are reviewed by a human.
Professors often evaluate dozens of handwritten answer sheets for every examination. This process is time-consuming, repetitive, and can lead to inconsistencies when large batches of papers must be graded under time constraints.
Students usually receive only a final score and rarely receive detailed feedback explaining where marks were gained or lost.
ScribeAI explores how AI agents can assist with routine evaluation while keeping human judgment involved whenever needed.
The system follows a four-agent workflow:
Responsibilities:
- Read scanned images and PDFs
- Extract handwritten answers
- Identify student information
- Generate handwriting confidence score
- Detect suspicious content and prompt injection attempts
Output:
{
"student_name": "...",
"roll_number": "...",
"answers": [...]
}Responsibilities:
- Read professor marking scheme
- Evaluate answers concept-by-concept
- Award marks based on rubric
- Detect keyword stuffing
- Generate evaluation confidence score
Match Levels:
- Full Match = 100%
- Partial Match = 50%
- No Match = 0%
Responsibilities:
- Quality control
- Routing decisions
- Human review escalation
Routing Rules:
- Legibility Confidence β₯ 80
- Evaluation Confidence β₯ 80
- Score between 10% and 95%
- Confidence between 70 and 79
- Extreme scores
- Any uncertain evaluation
- Confidence below 70
- Suspicious content detected
- Major disagreement between evaluators
Responsibilities:
- Independent re-evaluation
- Compare scores
- Generate final reports
- Upload results via MCP
- Maintain grading logs
Student Answer Sheet
β
Agent 1
Handwriting Extraction
β
Agent 2
Concept Evaluation
β
Agent 3
Routing Logic
β β
Human Agent 4
Review Second Opinion
β
Final Report
β
Google Drive MCP
β
Results Spreadsheet
Reads handwritten answers from scanned images and PDFs using Gemini Vision.
Grades based on concepts rather than exact wording.
Escalates uncertain evaluations to professors.
Processes entire folders of answer sheets.
Generates question-wise feedback for every student.
Maintains:
- results.xlsx
- results.csv
Uploads reports and spreadsheets automatically using Google Drive MCP.
- Blind grading
- Prompt injection detection
- Output validation
- Confidence-based routing
scribeai/
βββ AGENTS.md
βββ CONTEXT.md
βββ Makefile
βββ pyproject.toml
β
βββ app/
β βββ agent.py
β
βββ assets/
β βββ workflow.png
β
βββ uploads/
βββ extracted/
βββ evaluations/
βββ reports/
βββ spreadsheets/
βββ batch_reports/
β
βββ .agents/
β βββ skills/
β βββ handwriting-extractor/
β β βββ SKILL.md
β β
β βββ concept-evaluator/
β β βββ SKILL.md
β β
β βββ report-generator/
β βββ SKILL.md
β
βββ tests/
βββ test_agent.py
- Google ADK 2.0
- Gemini 2.5 Flash
- Google Drive MCP
- Agent Skills
- Python 3.11+
- Docker
- uv
- Agents CLI
Run the test suite using uv to automatically load dependencies:
uv run --with pytest pytest tests/The test suite includes:
- Perfect paper
- Near-zero paper
- Borderline pass
- Borderline fail
- Keyword stuffing
- Prompt injection attempt
- Unattempted questions
- Unclear handwriting
Current version supports:
- Text-based answers
- Handwritten text
- PDF and image input
Current version does not support:
- Circuit diagrams
- Graphs
- Engineering drawings
- Diagram evaluation
- Professor-specific grading adaptation
- University ERP integration
- Web dashboard
- Advanced analytics
- Multi-language support
As a student, I have experienced the long wait for examination results and the lack of meaningful feedback after assessments.
ScribeAI explores how AI agents can assist with one of the most repetitive academic tasks while still keeping important decisions under human supervision.
The objective is not to replace educators, but to help them focus on teaching and mentoring rather than repetitive grading work.
