Skip to content

feat: matcher rework, task-routed LLM calls, and beta hardening - #108

Merged
harmandeep2993 merged 24 commits into
mainfrom
feat/improvements
Sep 1, 2026
Merged

harmandeep2993 merged 24 commits into
mainfrom
feat/improvements

Conversation

@harmandeep2993

Copy link
Copy Markdown
Owner

Summary

Rebuilds the scoring engine around actionable sections plus hard gates, routes
every LLM call to a per-task provider/model tier (adding an Anthropic provider),
reworks the ATS optimiser into a verified full-resume generator, and lands a
round of security and privacy hardening for the public beta. 24 commits.

Matcher and scoring

  • Score the five sections a candidate can act on (required skills,
    responsibilities, preferred skills, education, certifications); move years of
    experience and language level to hard pass/fail gates instead of weighted
    dimensions.
  • Gate on relevant experience rather than total years; count experience in
    full-time-equivalent years.
  • Drop the related-skill partial-credit band: a skill matches or it does not,
    because the embedding model cannot reliably separate a related skill from a
    coincidentally similar one.
  • Responsibility coverage is judged by an LLM per duty rather than by keyword
    overlap.
  • Robust seniority detection for the entry-level gate, including seniority
    stated without a number, and a post-score entry gate that reads the extracted
    JD rather than the title.
  • New modules: matcher/gates.py, matcher/experience.py,
    matcher/responsibility_coverage.py; old per-section scorers removed.

LLM routing and providers

  • call_llm() routes each task to its own provider and model tier. High-volume
    extraction runs on a cheap fast model; low-volume, user-facing or
    single-score work runs on a stronger model. A tier with no API key falls back
    to the active provider instead of failing.
  • New Anthropic (Claude) provider.
  • Stop reasoning models returning empty responses; remove the unused Ollama
    health module.

ATS

  • Optimise is now a full-resume generator: verifies the rewritten resume
    against the source, adds a contact header, renders a two-column DOCX without
    tables.
  • Fix: optimise crashed on every call because format() consumed the JSON
    schema braces.
  • Remove the ATS Check tab from the sidebar nav (routes and tab component
    unchanged).

Job fetch

  • Configurable default country list, seeded with DE and NL.
  • Ask Adzuna for the newest jobs within the age window; seen-stop pagination so
    no job in the age window is skipped.
  • Fix a race where the scheduler and a manual run could both drive the same
    run-status counter past its total.

Analyzer UI

  • Results panel shows evidence, job context, the score math, and what each gap
    is worth.
  • Render summary emphasis as bold and focus steps as a list.
  • Fix: summary crashed on the new match() shape, which was failing every
    analysis.

Security and privacy

  • Enforce HTTPS, add CSP and HSTS, reject unsafe CORS in production.
  • GDPR erasure endpoint; keep PII out of tokens; stop storing contact data.
  • Stop exposing an API key suffix and the database host.

Docs and tests

  • README updated to the current engine and the task-routed call_llm();
    new Architecture section with a diagram under docs/.
  • Contract and matcher suites expanded (58 tests): new gate behaviour,
    per-duty coverage, tiered routing, and the None-sentinel engine paths.

Responsibilities and experience were 43 percent of the weight and both were
cosine comparisons between resume text and JD text. Measured: genuine
matches between duty-phrased JD lines and achievement-phrased resume bullets
land at 0.30-0.47 cosine, while bullets from unrelated professions reach
0.34, so the distributions overlap and no threshold separates signal from
noise. Both sections also measured the same thing, counting one weak signal
twice. Strong resumes capped around 60-70.

Scores, weighted and all actionable: required_skills 0.40, responsibilities
0.35, preferred_skills 0.15, education 0.05, certifications 0.05.
Responsibility coverage is now judged by the LLM per duty (yes/partial/no
with the evidence line), replacing the cosine scorer. It is opt-in via
llm_judge, so the Analyser gets it while bulk job scoring stays
deterministic and cheap.

Gates, pass/fail warnings and never points: years of experience parsed from
the JD in EN/DE and compared against total_experience_years, which
extraction already computed and the old scorer ignored entirely (a 1-year
and a 12-year candidate both scored 81.7 against a 5+ years requirement);
and language level via CEFR comparison.

The same realistic resume that scored 64.6 now scores 82.9, with the real
blockers stated plainly instead of silently shaving points.
Secret-safety pass before deployment. Two real leaks found and closed:

- provider_catalog() returned key_hint, the last 4 characters of the live
  OpenAI/Groq API key, to every authenticated user via GET /api/llm-settings.
  The frontend never used it. has_key is now a boolean and no fragment of any
  key reaches a client.
- database init logged the full Turso URL, putting the live database endpoint
  into logs and any crash report that captures them. It now logs the backend
  kind only.

.env.example was missing three variables the code reads (APP_USERNAME,
APP_PASSWORD, SESSION_SECRET) and used a stale Turso name; it now documents
every variable and states that the frontend reads none of them.

README gains a secrets section with the audit result and an explicit warning
to rotate any key that ever reached a commit, since git history is permanent.

No hardcoded secrets exist in the tree and .env was never committed.
…ct data

Personal-data audit before deployment. The resume is the densest PII payload
in the app, and three real problems were found:

No way to delete an account (GDPR Art. 17). DELETE /api/auth/account now
erases everything: resume files from disk, then rows in analyses,
analysis_cache, resumes, user_resume, matches, events, seen_jobs,
user_settings, users, plus the in-memory session. It requires the current
password, so a stolen token alone cannot destroy an account.

The JWT carried an email claim, and the token lives in localStorage where any
script on the page can read it. The server never used the claim (it resolves
users by sub), so it is removed; TopBar now fetches the email from
GET /api/auth/me. No personal data remains in browser storage.

Extraction collected the candidate name, email, phone, and personal links into
resumes.extracted_json, and nothing read them. That is data we neither need
nor use, so the schema drops them and the prompt now forbids emitting contact
details anywhere in the output. Scoring never needed them.

Also: the APP_PASSWORD session cookie gains secure=True in production, and the
dead ollama_health module (which printed raw LLM responses via print) is
deleted.

Passwords remain bcrypt-hashed and are never logged or returned. Logs record
lengths, never content.
…n production

Protects data in transit. TLS itself is terminated by the platform, but the app
now refuses to be used insecurely when APP_ENV=production:

- HTTPSRedirectMiddleware sends plain-HTTP requests to https before a body is
  read, so a resume or password is never transmitted in clear text
- Strict-Transport-Security (1 year, includeSubDomains) makes browsers refuse
  http:// for the domain entirely. Deliberately not sent in dev, where it would
  pin localhost to https in the developer browser for a year
- CORS boot check: a wildcard origin with allow_credentials would let any site
  issue authenticated requests as a logged-in user and read the response, so the
  server exits on '*' or any http:// origin

Adds a Content-Security-Policy on every response. The JWT is in localStorage, so
an injected script is the realistic path to stealing it: script-src 'self' blocks
that, connect-src 'self' removes the exfiltration channel, and object-src 'none'
plus base-uri and form-action lock the remaining injection vectors. Google Fonts
is allowed for styles and fonts only, and blob: for the resume preview iframe.
Permissions-Policy disables geolocation, microphone, and camera.

All outbound calls were already HTTPS. _IS_PRODUCTION moves above the middleware
that reads it - it was defined after first use.
…y analysis

/api/analyze returned 500 after every LLM call had already been paid for.
generate_summary still indexed breakdown[experience], but experience became a
gate rather than a scored section, so the key no longer exists: KeyError.

The experience label is now derived from the gate (meets / below the N+ years
required, or no explicit requirement), which is what the reader actually wants
to know. Section labels also tolerate None, since preferred_skills,
certifications, and education legitimately come back as None when the JD says
nothing about them, and comparing None to a number would have been the next
crash.

A regression test drives generate_summary with the real match() output,
including None sections.
The summary prompt asks the LLM to wrap key terms in strong tags, but the panel
rendered each item directly. React escapes raw HTML, so users saw the literal
markup instead of bold text.

Parsing the marker in JS and emitting real elements fixes it without
dangerouslySetInnerHTML, which would be unsafe here: the text is LLM output, and
injecting it as markup would turn prompt injection into an XSS vector. Only the
strong element is ever produced, so any other tag a model emits stays inert
plain text that React escapes.

Recommended focus also joined its array with a space, which read as one run-on
sentence. Each item is a separate action, so it now renders as a numbered list.
…t models

Reported: a resume showed 6.1 years of experience that no recruiter would
recognise. The cause was our own arithmetic, not the LLM - the prompt tells the
model to emit 0 and the total is computed in Python - so it credited every role
at its full calendar span. A two-year Werkstudent contract (about 20h/week) and
a six-month internship counted the same as two years and six months of full-time
work.

Years are now full-time-equivalent. The timeline is walked one month at a time
and each month is credited at the weight of the most substantial role covering
it: full-time, freelance, apprenticeship, research, and teaching at 1.0;
part-time, working student, and internship at 0.5; volunteer at 0.25; anything
unrecognised at 1.0 so an unknown type is never silently discounted. Taking the
maximum weight per month also means a working student job held during a
full-time role adds nothing instead of double-counting. This matters directly
for the experience gate, which was passing candidates on inflated totals.

Extraction now runs on its own model. Every score depends on the structured JSON
pulled out of the resume and JD, so the provider can nominate a stronger model
for it via extraction_model in config.yaml (gpt-5-mini) while scoring summaries
and the relevance gate stay on the cheaper default (gpt-4o-mini). An explicit
admin pin in the UI still wins over the config default.

The reasoning-model family takes different request parameters: gpt-5, o1, o3,
and o4 reject max_tokens in favour of max_completion_tokens and only accept the
default temperature, so sending the chat-model payload would have returned 400
on the first extraction call. The OpenAI provider now branches on the model.
… gap is worth

The analysis page showed a number and three flat keyword lists. It never
explained where a match came from, what the user was being scored against, or
what to do next. Three things were already computed and simply never surfaced.

Skills now carry provenance. The scorer already knew how each skill resolved, so
it now reports it: listed in the skills section, proven by a bullet, a close
name match, or a related skill. The UI quotes the exact resume line that proves
a skill, which turns a chip into evidence.

That exposes the insight the product exists to give: a skill proven in an
experience bullet but absent from the skills section is one a recruiter's
keyword search will never find. Those are pulled out into a Quick win panel with
the advice to add them verbatim, which is free score with no dishonesty.

Preferred skills were computed, sent, and rendered nowhere - 15 percent of the
score with no explanation. They now get their own panel.

The job context (title, company, work mode, employment type, seniority) was
extracted by the LLM and thrown away. It now sits beside the score, so the user
can see what they are being judged against.

Two new panels make the number actionable and defensible. What would raise your
score ranks each missing required skill by the exact points it would add, which
is pure arithmetic on the weights the engine already used. How this score was
calculated shows section by section how the total was built, including which
sections the job ad never mentioned and were therefore excluded.

Skill scorers now return a fifth element (the evidence map); the engine and the
monkeypatched tests are updated accordingly. Cached results from older runs
still render: the panel falls back to the flat keyword lists when the richer
breakdown is absent.
Resume extraction started failing with 422. The log showed gpt-5-mini returning
0 chars after 48 seconds with finish_reason=length: max_completion_tokens is a
budget for reasoning PLUS output, so passing the plain output budget through let
the model spend the entire allowance thinking and emit nothing at all. Reasoning
models now get their own token headroom on top of the output budget, plus a low
reasoning_effort, since our tasks are structured extraction and bounded
judgement rather than open-ended problem solving. A live call that previously
returned 0 chars in 48s now returns valid JSON in 1.5s.

The model tiers were also wrong on two counts. Putting extraction on the strong
model silently made every job fetch hundreds of times more expensive, because
the pipeline extracts one JD per fetched job. And extraction is the task that
benefits least: it is structured parsing, where reasoning adds nothing while
giving up both speed and a deterministic temperature.

Models are now tiered by the shape of the work, not by whether it is extraction:

  bulk (gpt-4o-mini) - runs in loops: one JD extraction per fetched job, one
  relevance call per 30 titles. Hundreds per run, so it must stay cheap.

  quality (gpt-5-mini) - low volume, and either read by the user, sent to an
  employer, or decides a single job's score: profile summary, ATS optimise,
  resume rewrite, and the responsibility coverage judge, which alone determines
  35 percent of the score.

An explicit admin model pin still overrides both tiers.
Reported from a live summary: "AI/ML Engineer with 6.1 years of experience".
Total years says nothing about how much of it is relevant. A career changer with
five years in mechanical engineering and one in ML has 6.1 years of experience
and roughly one year of what the employer actually asked for - and would have
sailed through a 5+ years gate on a number that was almost entirely irrelevant.

Experience is now measured against the job. A past role counts only if it
evidences at least a quarter of the JD's required skills, reusing the same
evidence search the skill scorer already uses, so the judgement is deterministic,
free, and consistent with how skills are matched elsewhere. The gate compares the
requirement against those relevant years and reports both figures, so the user
sees the difference rather than a flattering total.

The career-changer case now correctly fails: "This role asks for 5+ years. You
have 6.3 years of experience, but only 1.3 in roles matching this job." A genuine
eight-year ML engineer still clears the same bar. The gates panel lists which
roles counted and which were a different field.

The summary prompt was also implying unrelated experience counted toward the
role. It now receives both numbers and is told to cite the relevant years.

Duration maths (FTE weighting, date parsing, interval merging) moves out of the
resume extractor into services/matcher/experience.py, since the gates need it
too and were otherwise reaching into the extractor's private helpers.
… stop citing the objective

Three defects visible in a real analysis, all fixed.

The related-skill band was awarding half credit for skills the candidate does
not have. Measured against our embedding model:

    java vs fastapi        0.822   nonsense, and the highest score of the three
    aws vs azure           0.706   genuinely related
    kubernetes vs docker   0.412   genuinely related, scored as unrelated

The wrong pair outranks both right ones, so no threshold can separate them - the
same failure that killed cosine for responsibilities. The model has no skill
ontology; it matches surface tokens. The band is removed: a skill is present or
it is not. Embeddings are still used, but only to catch a rephrasing of the same
skill (python vs python programming, 0.93). The closest-skill hint on a missing
entry goes too, since the same unreliable measurement would offer fastapi as the
nearest thing to java, and a wrong hint is worse than none. In the reported case
this drops required skills from 50 to 30, which is the honest number: four of the
five "matches" were fabricated.

The JD extractor was pulling personality traits into preferred_skills
("team-oriented", "pragmatic implementation", "interest in ai trends"), which
then scored 0 of 3 and cost a section carrying 15 percent of the weight. Nobody
lists those as resume keywords. Both extraction prompts now state that skills are
things, not traits, and must exclude attitudes and work styles entirely.

Skill evidence quoted the objective line back as proof: "artificial intelligence
- proven in your experience: Searching for job opportunities in AI/ML". A
statement of intent is not evidence of doing the work. The summary is excluded
from evidence quotes; a skill appearing only there is still found by the corpus
search, so the ATS advice still fires, but it can no longer be cited as
experience.

Scoring version bumped so cached analyses are not served under the old rules.
…act header

The generated resume was neither sendable nor trustworthy.

It had no contact header. The DOCX began at "Summary" - no name, no email, no
phone - so it was not a resume anyone could send, and an ATS that cannot find an
email has nothing to attach an application to. The writer now extracts the
contact block from the source resume and both the DOCX and the plain-text preview
lead with it. Contact details are processed in memory and returned to the client;
nothing is persisted, so the data-minimisation position is unchanged.

More seriously, "never invent anything" was an instruction to the model and
nothing more - no code checked whether the output was true. It is now enforced.
Every skill in the generated resume must be evidenced in the original text via
the same alias-aware search the scorer uses, and every employer or job title must
appear in the source; anything else is stripped and reported. A test drives a
deliberately hallucinating response: three invented skills and a fabricated
"Senior ML Engineer at Google" role are all removed, and only what the resume
actually supports survives.

The UI shows what was stripped, so the user can see both that the tool caught the
fabrication and that everything remaining is backed by their real experience -
which matters on a document they put their name to.

The writer is also now given the JD's required keywords explicitly, with the
instruction to use only those the resume genuinely supports: a missing skill stays
missing. Generation already runs on the quality model.
…ma braces

ATS optimise has never worked. The generation prompt embeds a literal JSON
schema to show the model the expected shape, and the code built it with
str.format(), which reads every brace in that schema as a placeholder. The call
raised KeyError before it ever reached the LLM. The bug predates the contact and
verification work; adding a "contact" block to the schema only changed which key
appeared in the error.

Placeholders are now substituted with str.replace(), so the literal braces of the
JSON schema pass through untouched.

Verified end to end against the live API: check returns 75 percent coverage with
kubernetes correctly reported missing, optimise returns 200 with the contact
block populated, the skills list containing only what the resume evidences, and
the DOCX rendering to a real document.
Two-column resumes normally break ATS parsing because they are built with a
table, and a parser reads a grid cell by cell - which is how job titles end up
interleaved with skills, or dropped. This uses Word's column layout instead
(the w:cols element on the section), which is a rendering instruction only: the
paragraphs stay in linear order in the XML, so a text extractor still walks them
top to bottom.

The header and summary stay full width, then a continuous section runs two
columns with a column break between the sidebar (Skills, Education) and the main
content (Work Experience). Verified by reading the generated file back: zero
tables, every standard heading present, and the reading order intact. Single
column remains the default and the safest option; the layout is a request
parameter, validated, with both buttons in the UI.

Adds Anthropic as a fourth provider. The Messages API differs from the OpenAI
chat API in ways the provider module hides: x-api-key auth with a version header,
a required max_tokens, and no response_format json_object - JSON is forced by
prefilling the assistant turn with an opening brace and prepending it back onto
the reply. Claude Sonnet 5 is wired as its quality model, Haiku 4.5 as the bulk
model, and gpt-5 is added to the OpenAI list as the stronger option there.

Set ANTHROPIC_API_KEY in .env and select the provider in Settings to use it.
Model selection was per-provider: both tiers had to come from whichever vendor
was active, so running cheap bulk work on one and generation on another was
impossible. Tasks now carry their own provider as well as their model, declared
in task_models in config.yaml:

  bulk     one JD extraction per fetched job, one relevance call per 30 titles.
           Hundreds per run, so it must stay fast and cheap.
  quality  the profile summary, the ATS resume rewrite, and the responsibility
           coverage judge - read by the user, sent to an employer, or deciding a
           single job's score.

Callers name the task rather than looking a model up themselves, so the routing
lives in one place. A task pointed at a provider whose API key is absent logs a
warning and falls back to the active provider, so a half-configured deployment
degrades instead of failing every request. An explicit admin pin in Settings
still overrides everything.

Verified live with bulk on OpenAI and quality on Anthropic in the same run:
gpt-4o-mini answered the bulk call in 1.2s, claude-sonnet-5 the quality call in
1.4s.

Two real Anthropic API constraints surfaced while testing and are now handled.
The current models reject `temperature` outright, and they reject the
assistant-prefill trick used to force JSON, so temperature is not sent and JSON
is requested in the prompt instead, with the router's existing fence-stripping
parser cleaning up whatever survives. Both failures had been caught by the Groq
fallback, which is the behaviour intended, but they would have meant every
quality call silently degrading to llama.
The deterministic seniority net - the only reliable layer, since the LLM is
inconsistent and skips titles in a batch - caught 3 of 19 realistic senior
titles. It knew a handful of English words and nothing else.

Now catches the ways senior roles actually hide from a keyword filter:
level numbers (Engineer II, SDE 3, L5), compound leads (Teamlead, TechLead),
missing seniority words (expert, chief, VP, vice president, distinguished,
fellow), and German experience markers (Abteilungsleitung, Erfahrener, mit
Berufserfahrung). 25 of 25 test titles caught.

Version numbers are not seniority: a bare level digit counts only right after
a role noun (Engineer 3), never after a technology (Python 3, Angular 2, Web3,
S3). An explicit entry word (junior, graduate, Werkstudent, Azubi) always
overrides, so Junior Engineer II is not misread as senior. Zero false
positives across the version-number and entry-word traps.

The cheap pre-LLM keyword gate (config exclude_keywords) gains the same words
so more seniors are dropped before any token is spent.
…title

The title-only fetch gate cannot tell whether a plain 'ML Engineer' posting
is actually entry level - the body might say '5+ years' or the job_level might
be 'senior'. But by the time a job is scored, its full JD has been extracted,
so that evidence is already in hand.

A second entry gate now runs after scoring, when entry-only is on: if the
extracted JD names a senior job_level (senior/lead/principal/staff/director)
or asks for more years than an entry candidate has (max_experience_years,
default 2), the job is dropped and its id blocked from re-appearing - even
though it was already scored. mid-level within the year bar is kept, since
many such roles accept strong juniors and the years check is the real bar.

Deterministic and free: it reads data the scorer already produced, no extra
LLM call. This is what catches the seniors that slip past a title that gives
nothing away.
Many JDs never write a year count. They say 'several years of experience',
'proven track record', 'mehrjaehrige/einschlaegige Berufserfahrung', or
'fundierte Kenntnisse' - and the numeric parser returned None on all of them,
so the post-score entry gate let them through.

The gate now also matches these non-numeric seniority phrases (EN and DE,
including ae transliterations of umlauts), scanning both the extracted
experience_requirements and the job summary. An explicit entry phrase always
overrides - 'no prior experience', 'Berufseinsteiger willkommen', 'erste
Berufserfahrung', 'graduates welcome' - so first-job postings that mention the
word 'experience' are not misread as senior. A stated number stays
authoritative: '1 year of experience' is entry and never triggers the phrase
rule. Nine phrasings caught, zero false positives across the entry traps.

The JD extraction prompt is updated to capture these non-numeric phrases into
experience_requirements verbatim rather than dropping anything without a
digit, so the gate has them to read.
…e + nl

Adzuna already fetched any country via its per-country endpoint, but the default
for a user who never chose their own was a single country ('de'). Netherlands
worked only if the user typed it into Search Criteria.

The default is now a list, default_countries in config.yaml, seeded with de and
nl. get_countries returns it for users with no saved choice; users who set their
own countries are unaffected, and Arbeitnow/Bundesagentur still only run when de
is included since they are German boards.
Adzuna was queried with its default relevance ordering and no age limit, which broke the incremental-fetch model in two ways: a fresh posting ranked below the page window by relevance was never seen at all, while stale jobs we would discard in the recency filter anyway occupied slots in it.

sort_by=date and max_days_old=MAX_AGE_DAYS make the window mean the newest jobs within our 30-day rule. Repeated runs walk new postings from the top and the seen_jobs dedup makes each run pay only for the genuinely new ones. Verified live: newest-first, all within the window, and the server-side age filter shrank one query pool from 780 relevance-ranked jobs to 113 recent ones.
The Adzuna window was a fixed per-title budget (50 of possibly hundreds), and every run re-fetched the same top slice: jobs ranked below the window were never seen by anyone, ever. With results now date-sorted, exhaustive coverage becomes cheap to guarantee.

fetch_adzuna_jobs gains a seen-stop mode: given the user's already-processed ids, it walks the date-sorted pool page by page and stops at the first page containing nothing new - everything older was covered by an earlier run. The first run walks the whole 30-day pool (bounded by max_pages_per_query, default 10 pages / 500 jobs per query); each later run downloads only the pages holding new postings. Budget mode (seen_ids=None) is unchanged for existing callers and tests.

Wired through fetch_combined into both the manual run route and the scheduler, which pass the user's seen_jobs ids. Verified live: run 1 fetched all 115 jobs of a query pool across 3 pages; run 2 downloaded one page, found nothing new, and stopped with zero LLM cost. A unit test pins the walk/stop/new-jobs-at-top behaviour with a stubbed API.
The scheduler checked has_resume/get_run_status and then started a fetch without claiming the run slot, so a manual fetch begun during the slow paginated fetch could pass begin_run() and start a second scoring loop. Both loops then incremented the same status["checked"], driving it past status["total"] (the "688 of 446 checked" symptom).

begin_run() now does the check-and-set under a lock, the scheduler claims the slot via begin_run() before fetching and releases it on failure, and discover_and_score() refuses to enter when a scoring loop is already in an active phase for that user.
Drops the ATS Check entry from the sidebar navigation. The ATS routes and tab component stay in place; only the nav entry point is removed.
… section

Updates the README to match the current engine: five scored sections plus two pass/fail gates instead of seven weighted dimensions, no related-skill partial credit, and the task-tiered call_llm() router. Adds an Architecture section with a diagram (docs/architecture.png) and a docs/README.md explaining where the diagram lives and why it must not go under images/.
Copilot AI lite review requested due to automatic review settings September 1, 2026 06:50

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The review found a few verified correctness/documentation issues in changed code paths that should be fixed before approval.

Once you've addressed the issues Copilot identified, you can request another Copilot review.

Pull request overview

This PR substantially rebuilds JobsFitAI’s matching/scoring pipeline around actionable sections plus hard pass/fail gates, introduces task-routed LLM calls with an Anthropic provider option, upgrades ATS “optimise” into a verified full-resume generator with DOCX layout options, and adds multiple security/privacy hardening measures for the public beta.

Changes:

  • Reworked matcher: five scored sections + deterministic gates for experience/language, plus optional LLM-judged responsibility coverage.
  • Added task-based LLM routing (call_llm(task=...)) with provider/model tiering and a new Anthropic provider; improved handling for OpenAI reasoning models.
  • ATS optimisation now generates a full resume, verifies output against the source, and supports single- and two-column DOCX rendering; plus beta security/privacy improvements (CSP/HSTS/HTTPS enforcement, account erasure, no PII in JWT).
File summaries
File Description
README.md Updates product description, routing, architecture, and security/deployment documentation.
docs/README.md Adds docs folder guidance and architecture diagram conventions.
frontend/src/lib/errors.js Updates user-facing auth error text.
frontend/src/lib/auth.js Removes JWT-decoding helper; relies on /api/auth/me for email.
frontend/src/components/TopBar.jsx Fetches and displays user email via authenticated API call.
frontend/src/components/tabs/Settings.jsx Adds GDPR account deletion UI modal and flow.
frontend/src/components/tabs/ATS.jsx Adds DOCX layout selection and displays removed unsupported claims.
frontend/src/components/Sidebar.jsx Removes ATS Check tab from navigation.
frontend/src/components/AnalysisResults.jsx Reworks analyzer UI: gates panel, evidence-based skills, per-duty coverage, score math, rich-text rendering.
backend/tests/test_matcher.py Expands matcher tests for gates, experience FTE math, LLM routing tiers, and coverage behavior (stubbed).
backend/tests/test_contract.py Adds contract tests for ATS verification, DOCX layouts, account deletion, JWT claims, and CSP assertions.
backend/services/resume_rewriter.py Routes resume-improvement LLM calls through the quality tier.
backend/services/prompts/schemas/resume_schema.json Removes contact fields from resume extraction schema.
backend/services/prompts/resume_prompt.py Adds strict “never extract contact details” rule and refines skill extraction constraints.
backend/services/prompts/jd_prompt.py Tightens skills extraction to exclude “traits” and expands experience requirement capture (numeric + non-numeric).
backend/services/profile_summary.py Adapts summary generation to gates (relevant vs total years) and new match output shape; routes via quality tier.
backend/services/matcher/skill_aliases.py Adds line-level evidence extraction for UI quoting.
backend/services/matcher/scores/skills.py Removes related-skill partial credit; adds evidence mapping and “similar rephrasing” only.
backend/services/matcher/scores/responsibilities.py Removes cosine-similarity responsibilities scorer (replaced by LLM-judge module).
backend/services/matcher/scores/languages.py Converts language logic into parsing helpers (consumed by gates), removes scoring.
backend/services/matcher/scores/experiences.py Removes experience scorer (replaced by gates + FTE math).
backend/services/matcher/scores/init.py Updates exports to scored sections only (no experience/languages/responsibilities scorer).
backend/services/matcher/responsibility_coverage.py Adds LLM-judged per-duty responsibility coverage module.
backend/services/matcher/gates.py Adds deterministic gates: relevant experience years and language proficiency; includes seniority detection logic.
backend/services/matcher/experience.py Introduces FTE-weighted, non-double-counting experience duration math.
backend/services/matcher/engine.py Reworks matcher orchestration to scores + gates; responsibility coverage becomes opt-in.
backend/services/llm/providers/openai.py Adds reasoning-model detection and correct parameterization/token budgeting; removes key hint exposure.
backend/services/llm/providers/groq.py Removes key hint exposure.
backend/services/llm/providers/anthropic.py Adds Anthropic provider implementation for quality tier.
backend/services/llm/ollama_health.py Removes unused Ollama health module.
backend/services/llm/caller.py Adds task routing (provider/model selection) and supports Anthropic in provider resolution.
backend/services/job_relevance.py Strengthens deterministic seniority detection with structural patterns and explicit entry-word override.
backend/services/job_matcher.py Fixes run-status race via locking; adds seen-stop pagination support and post-score entry-level gate.
backend/services/fetchers/job_fetcher.py Adds Adzuna newest-first + max-age filters and “seen-stop” pagination, bounded by config.
backend/services/extractors/resume_extractor.py Reuses shared experience-math module and clarifies extraction model intent.
backend/services/extractors/jd_extractor.py Clarifies extraction as cheap/bulk task behavior.
backend/services/ats.py ATS generation becomes full-resume JSON + verification against source; adds contact header in plain-text render.
backend/schemas/auth.py Adds request schema for account deletion.
backend/schemas/ats.py Adds layout field for DOCX generation request.
backend/repositories/settings_repo.py Switches default countries to configurable list.
backend/models/user.py Adds hard-delete account erasure across all per-user tables and cached analyses.
backend/main.py Enforces production HTTPS/CORS safety, adds CSP/HSTS/Permissions-Policy, and hardens scheduler run claiming.
backend/core/state.py Adds task model routing and Anthropic provider support; removes key hint from provider catalog.
backend/core/security.py Removes email from JWT payload; token contains only user id (sub).
backend/core/database.py Stops logging Turso endpoint URL/host details.
backend/core/config.py Adds Anthropic config, task routing config, default countries list, and Adzuna paging cap.
backend/core/config_validator.py Updates matcher weights schema to scored sections only (gates excluded).
backend/api/routes/resume_analyzer.py Updates scoring version; enables LLM judge for analyzer; returns job context, score math, evidence, gates, duties, and skill impact.
backend/api/routes/job_matches.py Enables seen-stop pagination and enables LLM judge for single-job scoring endpoint.
backend/api/routes/auth.py Adds account deletion endpoint, clears in-memory state, removes email from JWT creation.
backend/api/routes/ats_maker.py Adds single/two-column DOCX generation with columns (no tables) and contact header.
.env.example Documents production safety requirements and optional legacy password gate variables.
Review details
  • Files reviewed: 52/53 changed files
  • Comments generated: 3
  • Review effort level: Lite

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread README.md
Comment on lines +141 to +143
<!-- Diagram not committed yet. Export your architecture image to docs/architecture.png
(create the docs/ folder at the repo root). Do NOT use a folder literally named
"images" - .gitignore ignores it. SVG also works: docs/architecture.svg. -->
Comment thread backend/services/ats.py
for skill in parsed.get("skills") or []:
if not isinstance(skill, str) or not skill.strip():
continue
if found_in_corpus(skill, corpus) or skill.lower().strip() in corpus:
"blocking_count": 2,
},
"matched_required": ["python", "git"],
"partial_required": ["tensorflow"], # related skill, half credit
@harmandeep2993
harmandeep2993 merged commit 17338e6 into main Sep 1, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants