Skip to content

Latest commit

 

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

DeepProbe

It doesn't test what you learned. It tests whether you can engineer with it.

DeepProbe is an adaptive AI technical interviewer for graduates of a 31-day AI Engineering cohort (a course that builds a RAG-powered healthcare-plan chatbot: structured/unstructured data → embeddings → vector search → RAG → prompting → function calling → fine-tuning → agents → multi-agent orchestration → MCP → evaluation → security → deployment).

It studies each candidate's actual learning journey — completions, attempts, skips — before asking a single question, then conducts a real multi-turn conversational interview that adapts its next question to what the candidate just said.


Why this is not a questionnaire

Every other take on "AI interviewer" tends to be a fixed list of questions with an LLM wrapper. DeepProbe's interview loop is a closed loop, not a script:

Candidate Profile → Interview Planner → Question → Candidate Answer
        ↑                                                    │
        └──────────── Update Knowledge State ← Answer Analysis

Two concrete things make this real:

  1. Completion ≠ mastery. A mission passed on attempt 1 and a mission passed on attempt 5 both show up as "passed" in the raw data — but DeepProbe treats them very differently (strong vs uncertain). Skipped topics are marked evidence: unavailable, not silently treated as failures.
  2. The next question depends on the last answer. A correct, deep answer escalates toward system design or a cross-topic scenario. A partial answer gets one probing follow-up on the missing piece. An incorrect answer gets a diagnostic follow-up from a different angle (max 2 in a row) before the interview moves on. This is implemented in server/services/interviewPlanner.js — you can watch it happen in the "How your interview adapted" screen at the end of every interview.

System Architecture

DeepProbe System Architecture


Quick start

You'll need Node.js 18+ and a free Gemini API key from Google AI Studio.

# 1. Install dependencies for both server and client
npm run install:all

# 2. Configure your API key
cp server/.env.example server/.env
# edit server/.env and set GEMINI_API_KEY=...

# 3. Run the backend (terminal 1)
npm run dev:server        # http://localhost:8787

# 4. Run the frontend (terminal 2)
npm run dev:client        # http://localhost:5173

Open http://localhost:5173. Pick a candidate from the roster, review their learning profile, start the interview, answer a few questions, and watch the final report.


Demo flow (what to show judges)

  1. Select candidate — 20 real candidates are in server/data/candidates.json, ranging from a candidate who passed almost everything on the first try (Emily Chen) to one who skipped most core AI topics (Mia Alvarez).
  2. Learning profile — shows exactly what DeepProbe extracted from that candidate's mission history: strong / uncertain / weak / skipped topics and an estimated starting difficulty.
  3. Start the interview — answer honestly, then try answering one question either very well or very poorly and watch the next question shift accordingly.
  4. Cross-topic question — appears when the candidate is doing well; combines 2–3 curriculum topics into one realistic engineering scenario (e.g. vector DB + retrieval + MCP).
  5. Final report — overall score, five category scores, strengths, gaps, next steps.
  6. "How your interview adapted" — a node-by-node replay of every question with a plain-language reason it was asked, proving the interview really was adaptive and not a fixed script.

Run the same candidate twice, or run two different candidates back to back — the questions and the path through the interview will differ because the underlying knowledge state differs.


Architecture

deep-probe/
├── server/
│   ├── data/
│   │   ├── curriculum.json         # 31-day cohort curriculum (supplied)
│   │   └── candidates.json         # 20 candidate mission records (supplied)
│   ├── utils/
│   │   └── curriculum.js           # day/module lookup, cross-topic combo map
│   ├── services/
│   │   ├── candidateProfiler.js    # deterministic: raw missions → Knowledge Profile
│   │   ├── interviewPlanner.js     # the adaptive core: topic selection, difficulty
│   │   │                           #   control, follow-up logic, session orchestration
│   │   ├── questionGenerator.js    # LLM: writes the next interviewer message
│   │   ├── answerEvaluator.js      # LLM: grades correctness/depth/misconceptions
│   │   ├── feedbackGenerator.js    # LLM: final structured assessment
│   │   └── llmClient.js            # Gemini API wrapper, retries, JSON repair
│   ├── state/
│   │   └── sessionStore.js         # in-memory Map<sessionId, InterviewState>
│   ├── routes/
│   │   ├── interview.js            # POST /api/interview (the required contract)
│   │   └── candidates.js           # GET endpoints for the demo candidate picker
│   └── index.js
└── client/                         # React + Vite
    └── src/
        ├── pages/
        │   ├── CandidateSelect.jsx
        │   ├── ProfileView.jsx     # pre-interview knowledge profile
        │   ├── InterviewScreen.jsx # the conversational interview UI
        │   └── ReportScreen.jsx    # final scorecard + adaptation path
        └── services/api.js

Why the profiler isn't an LLM call

candidateProfiler.js is deterministic code, not a prompt. Mission pass/fail/attempt data is a factual record — computing "passed on attempt 1 = strong evidence, passed on attempt 5 = weak evidence" is a pure function of that data, and running it through an LLM would only add latency and a chance of misreading the numbers. The LLM budget is spent where judgment is actually required: writing the next question, grading an open-ended answer, and writing the final narrative feedback.

Why difficulty/follow-up control isn't a separate LLM call

The Answer Evaluator already returns a structured correctness + depthScore + misconceptions verdict. The Difficulty Controller and Follow-up Controller are deterministic functions over that structured output (nextDifficulty() and decideNextDirective() in interviewPlanner.js) rather than a second free-text LLM call. This keeps the loop fast (2 LLM calls per turn instead of 3–4), makes the adaptive logic auditable/testable, and removes a source of inconsistency between "what the evaluator found" and "what the controller decided to do about it."

Topic selection & coverage

interviewPlanner.js builds a topic pool from the candidate's own mission history (plus a handful of core AI- engineering days that are important but weren't in the candidate's record, marked unassessed). Each turn it weights untouched topics by status (weak > uncertain > skipped > unassessed > strong) times a curated importance score, with light randomization so runs aren't perfectly deterministic. It enforces the hackathon minimums (≥8 questions, ≥4 distinct curriculum days) and caps at 12–14 questions so the interview doesn't ramble.

Privacy / no leaking internals

The candidate never sees a score, a topic label, a difficulty level, or a reason during the interview — only the next question. The /api/interview response never includes anything beyond reply/done/feedback. The "how your interview adapted" rationale strings are template-based plain-language explanations (rationaleFor() in interviewPlanner.js), generated after the fact — never a dump of internal scores or chain-of-thought.


API

POST /api/interview

No authentication. State is keyed by sessionId.

Start a session (first request for a given sessionId):

{ "sessionId": "abc-123", "candidate": { "member": {...}, "missions": [...], "signals": {...} } }

{ "reply": "Welcome. Let's begin your interview.", "done": false }

Continue (every subsequent request):

{ "sessionId": "abc-123", "message": "RAG retrieves relevant documents and feeds them to the LLM as context." }

{ "reply": "Good. Now suppose retrieval quality is high but the answer still hallucinates...", "done": false }

Finish:

{
  "reply": "Interview completed.",
  "done": true,
  "feedback": {
    "summary": "...",
    "strengths": ["..."],
    "gaps": ["..."],
    "next": ["..."]
  },
  "adaptationPath": [{ "index": 1, "topic": "...", "level": 2, "rationale": "..." }]
}

adaptationPath is additive (not required by the spec) and powers the "How your interview adapted" view; the feedback object always contains at minimum summary, strengths, gaps, next as required.

Errors: 400 missing/invalid sessionId, candidate, or message; 409 if the session is already complete; 502 if the LLM backend is unavailable; 500 for anything unexpected. The server never crashes on a malformed LLM response — every LLM call has a structural fallback (see llmClient.js).

GET /api/candidates

Returns the demo roster (id, name, role, experience) for the candidate picker.

GET /api/candidates/:id

Returns the full supplied candidate record (used to kick off POST /api/interview).

GET /api/candidates/:id/profile

Returns the same Candidate Knowledge Profile the interview engine itself computes — shown pre-interview so judges can see the "why" behind the questions that follow.

GET /api/interview/:sessionId/adaptation

Convenience re-fetch of the adaptation path for an existing/completed session.


Configuration

server/.env:

GEMINI_API_KEY=              # required — free key at https://aistudio.google.com/apikey
LLM_MODEL=gemini-2.5-flash   # optional override
PORT=8787                    # optional override

Never hardcode the key; it's read from the environment only.

Note: the Gemini free tier is rate-limited (requests/minute and requests/day). If you hit the limit mid-demo, the server returns a 502 with a clear message rather than crashing — just wait a few seconds and retry, or upgrade to a paid Gemini key for heavier use.


Known trade-offs (hackathon scope)

  • Sessions are stored in-memory (server/state/sessionStore.js). Swapping in Redis/Postgres means implementing the same get/set/has/remove interface — nothing else in the codebase needs to change.
  • The interview target length (10–11 questions, min 8, hard cap 14) is randomized slightly per session for variety rather than fully dynamic; it was a deliberate choice to guarantee the hackathon's coverage requirements are always met without needing a second LLM call just to decide "should we stop now."

About

AI-powered adaptive technical interviewer with candidate intelligence and Gemini-driven interview adaptation.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages