feat: matcher rework, task-routed LLM calls, and beta hardening - #108
Merged
Merged
Conversation
Responsibilities and experience were 43 percent of the weight and both were cosine comparisons between resume text and JD text. Measured: genuine matches between duty-phrased JD lines and achievement-phrased resume bullets land at 0.30-0.47 cosine, while bullets from unrelated professions reach 0.34, so the distributions overlap and no threshold separates signal from noise. Both sections also measured the same thing, counting one weak signal twice. Strong resumes capped around 60-70. Scores, weighted and all actionable: required_skills 0.40, responsibilities 0.35, preferred_skills 0.15, education 0.05, certifications 0.05. Responsibility coverage is now judged by the LLM per duty (yes/partial/no with the evidence line), replacing the cosine scorer. It is opt-in via llm_judge, so the Analyser gets it while bulk job scoring stays deterministic and cheap. Gates, pass/fail warnings and never points: years of experience parsed from the JD in EN/DE and compared against total_experience_years, which extraction already computed and the old scorer ignored entirely (a 1-year and a 12-year candidate both scored 81.7 against a 5+ years requirement); and language level via CEFR comparison. The same realistic resume that scored 64.6 now scores 82.9, with the real blockers stated plainly instead of silently shaving points.
Secret-safety pass before deployment. Two real leaks found and closed: - provider_catalog() returned key_hint, the last 4 characters of the live OpenAI/Groq API key, to every authenticated user via GET /api/llm-settings. The frontend never used it. has_key is now a boolean and no fragment of any key reaches a client. - database init logged the full Turso URL, putting the live database endpoint into logs and any crash report that captures them. It now logs the backend kind only. .env.example was missing three variables the code reads (APP_USERNAME, APP_PASSWORD, SESSION_SECRET) and used a stale Turso name; it now documents every variable and states that the frontend reads none of them. README gains a secrets section with the audit result and an explicit warning to rotate any key that ever reached a commit, since git history is permanent. No hardcoded secrets exist in the tree and .env was never committed.
…ct data Personal-data audit before deployment. The resume is the densest PII payload in the app, and three real problems were found: No way to delete an account (GDPR Art. 17). DELETE /api/auth/account now erases everything: resume files from disk, then rows in analyses, analysis_cache, resumes, user_resume, matches, events, seen_jobs, user_settings, users, plus the in-memory session. It requires the current password, so a stolen token alone cannot destroy an account. The JWT carried an email claim, and the token lives in localStorage where any script on the page can read it. The server never used the claim (it resolves users by sub), so it is removed; TopBar now fetches the email from GET /api/auth/me. No personal data remains in browser storage. Extraction collected the candidate name, email, phone, and personal links into resumes.extracted_json, and nothing read them. That is data we neither need nor use, so the schema drops them and the prompt now forbids emitting contact details anywhere in the output. Scoring never needed them. Also: the APP_PASSWORD session cookie gains secure=True in production, and the dead ollama_health module (which printed raw LLM responses via print) is deleted. Passwords remain bcrypt-hashed and are never logged or returned. Logs record lengths, never content.
…n production Protects data in transit. TLS itself is terminated by the platform, but the app now refuses to be used insecurely when APP_ENV=production: - HTTPSRedirectMiddleware sends plain-HTTP requests to https before a body is read, so a resume or password is never transmitted in clear text - Strict-Transport-Security (1 year, includeSubDomains) makes browsers refuse http:// for the domain entirely. Deliberately not sent in dev, where it would pin localhost to https in the developer browser for a year - CORS boot check: a wildcard origin with allow_credentials would let any site issue authenticated requests as a logged-in user and read the response, so the server exits on '*' or any http:// origin Adds a Content-Security-Policy on every response. The JWT is in localStorage, so an injected script is the realistic path to stealing it: script-src 'self' blocks that, connect-src 'self' removes the exfiltration channel, and object-src 'none' plus base-uri and form-action lock the remaining injection vectors. Google Fonts is allowed for styles and fonts only, and blob: for the resume preview iframe. Permissions-Policy disables geolocation, microphone, and camera. All outbound calls were already HTTPS. _IS_PRODUCTION moves above the middleware that reads it - it was defined after first use.
…y analysis /api/analyze returned 500 after every LLM call had already been paid for. generate_summary still indexed breakdown[experience], but experience became a gate rather than a scored section, so the key no longer exists: KeyError. The experience label is now derived from the gate (meets / below the N+ years required, or no explicit requirement), which is what the reader actually wants to know. Section labels also tolerate None, since preferred_skills, certifications, and education legitimately come back as None when the JD says nothing about them, and comparing None to a number would have been the next crash. A regression test drives generate_summary with the real match() output, including None sections.
The summary prompt asks the LLM to wrap key terms in strong tags, but the panel rendered each item directly. React escapes raw HTML, so users saw the literal markup instead of bold text. Parsing the marker in JS and emitting real elements fixes it without dangerouslySetInnerHTML, which would be unsafe here: the text is LLM output, and injecting it as markup would turn prompt injection into an XSS vector. Only the strong element is ever produced, so any other tag a model emits stays inert plain text that React escapes. Recommended focus also joined its array with a space, which read as one run-on sentence. Each item is a separate action, so it now renders as a numbered list.
…t models Reported: a resume showed 6.1 years of experience that no recruiter would recognise. The cause was our own arithmetic, not the LLM - the prompt tells the model to emit 0 and the total is computed in Python - so it credited every role at its full calendar span. A two-year Werkstudent contract (about 20h/week) and a six-month internship counted the same as two years and six months of full-time work. Years are now full-time-equivalent. The timeline is walked one month at a time and each month is credited at the weight of the most substantial role covering it: full-time, freelance, apprenticeship, research, and teaching at 1.0; part-time, working student, and internship at 0.5; volunteer at 0.25; anything unrecognised at 1.0 so an unknown type is never silently discounted. Taking the maximum weight per month also means a working student job held during a full-time role adds nothing instead of double-counting. This matters directly for the experience gate, which was passing candidates on inflated totals. Extraction now runs on its own model. Every score depends on the structured JSON pulled out of the resume and JD, so the provider can nominate a stronger model for it via extraction_model in config.yaml (gpt-5-mini) while scoring summaries and the relevance gate stay on the cheaper default (gpt-4o-mini). An explicit admin pin in the UI still wins over the config default. The reasoning-model family takes different request parameters: gpt-5, o1, o3, and o4 reject max_tokens in favour of max_completion_tokens and only accept the default temperature, so sending the chat-model payload would have returned 400 on the first extraction call. The OpenAI provider now branches on the model.
… gap is worth The analysis page showed a number and three flat keyword lists. It never explained where a match came from, what the user was being scored against, or what to do next. Three things were already computed and simply never surfaced. Skills now carry provenance. The scorer already knew how each skill resolved, so it now reports it: listed in the skills section, proven by a bullet, a close name match, or a related skill. The UI quotes the exact resume line that proves a skill, which turns a chip into evidence. That exposes the insight the product exists to give: a skill proven in an experience bullet but absent from the skills section is one a recruiter's keyword search will never find. Those are pulled out into a Quick win panel with the advice to add them verbatim, which is free score with no dishonesty. Preferred skills were computed, sent, and rendered nowhere - 15 percent of the score with no explanation. They now get their own panel. The job context (title, company, work mode, employment type, seniority) was extracted by the LLM and thrown away. It now sits beside the score, so the user can see what they are being judged against. Two new panels make the number actionable and defensible. What would raise your score ranks each missing required skill by the exact points it would add, which is pure arithmetic on the weights the engine already used. How this score was calculated shows section by section how the total was built, including which sections the job ad never mentioned and were therefore excluded. Skill scorers now return a fifth element (the evidence map); the engine and the monkeypatched tests are updated accordingly. Cached results from older runs still render: the panel falls back to the flat keyword lists when the richer breakdown is absent.
Resume extraction started failing with 422. The log showed gpt-5-mini returning 0 chars after 48 seconds with finish_reason=length: max_completion_tokens is a budget for reasoning PLUS output, so passing the plain output budget through let the model spend the entire allowance thinking and emit nothing at all. Reasoning models now get their own token headroom on top of the output budget, plus a low reasoning_effort, since our tasks are structured extraction and bounded judgement rather than open-ended problem solving. A live call that previously returned 0 chars in 48s now returns valid JSON in 1.5s. The model tiers were also wrong on two counts. Putting extraction on the strong model silently made every job fetch hundreds of times more expensive, because the pipeline extracts one JD per fetched job. And extraction is the task that benefits least: it is structured parsing, where reasoning adds nothing while giving up both speed and a deterministic temperature. Models are now tiered by the shape of the work, not by whether it is extraction: bulk (gpt-4o-mini) - runs in loops: one JD extraction per fetched job, one relevance call per 30 titles. Hundreds per run, so it must stay cheap. quality (gpt-5-mini) - low volume, and either read by the user, sent to an employer, or decides a single job's score: profile summary, ATS optimise, resume rewrite, and the responsibility coverage judge, which alone determines 35 percent of the score. An explicit admin model pin still overrides both tiers.
Reported from a live summary: "AI/ML Engineer with 6.1 years of experience". Total years says nothing about how much of it is relevant. A career changer with five years in mechanical engineering and one in ML has 6.1 years of experience and roughly one year of what the employer actually asked for - and would have sailed through a 5+ years gate on a number that was almost entirely irrelevant. Experience is now measured against the job. A past role counts only if it evidences at least a quarter of the JD's required skills, reusing the same evidence search the skill scorer already uses, so the judgement is deterministic, free, and consistent with how skills are matched elsewhere. The gate compares the requirement against those relevant years and reports both figures, so the user sees the difference rather than a flattering total. The career-changer case now correctly fails: "This role asks for 5+ years. You have 6.3 years of experience, but only 1.3 in roles matching this job." A genuine eight-year ML engineer still clears the same bar. The gates panel lists which roles counted and which were a different field. The summary prompt was also implying unrelated experience counted toward the role. It now receives both numbers and is told to cite the relevant years. Duration maths (FTE weighting, date parsing, interval merging) moves out of the resume extractor into services/matcher/experience.py, since the gates need it too and were otherwise reaching into the extractor's private helpers.
… stop citing the objective
Three defects visible in a real analysis, all fixed.
The related-skill band was awarding half credit for skills the candidate does
not have. Measured against our embedding model:
java vs fastapi 0.822 nonsense, and the highest score of the three
aws vs azure 0.706 genuinely related
kubernetes vs docker 0.412 genuinely related, scored as unrelated
The wrong pair outranks both right ones, so no threshold can separate them - the
same failure that killed cosine for responsibilities. The model has no skill
ontology; it matches surface tokens. The band is removed: a skill is present or
it is not. Embeddings are still used, but only to catch a rephrasing of the same
skill (python vs python programming, 0.93). The closest-skill hint on a missing
entry goes too, since the same unreliable measurement would offer fastapi as the
nearest thing to java, and a wrong hint is worse than none. In the reported case
this drops required skills from 50 to 30, which is the honest number: four of the
five "matches" were fabricated.
The JD extractor was pulling personality traits into preferred_skills
("team-oriented", "pragmatic implementation", "interest in ai trends"), which
then scored 0 of 3 and cost a section carrying 15 percent of the weight. Nobody
lists those as resume keywords. Both extraction prompts now state that skills are
things, not traits, and must exclude attitudes and work styles entirely.
Skill evidence quoted the objective line back as proof: "artificial intelligence
- proven in your experience: Searching for job opportunities in AI/ML". A
statement of intent is not evidence of doing the work. The summary is excluded
from evidence quotes; a skill appearing only there is still found by the corpus
search, so the ATS advice still fires, but it can no longer be cited as
experience.
Scoring version bumped so cached analyses are not served under the old rules.
…act header The generated resume was neither sendable nor trustworthy. It had no contact header. The DOCX began at "Summary" - no name, no email, no phone - so it was not a resume anyone could send, and an ATS that cannot find an email has nothing to attach an application to. The writer now extracts the contact block from the source resume and both the DOCX and the plain-text preview lead with it. Contact details are processed in memory and returned to the client; nothing is persisted, so the data-minimisation position is unchanged. More seriously, "never invent anything" was an instruction to the model and nothing more - no code checked whether the output was true. It is now enforced. Every skill in the generated resume must be evidenced in the original text via the same alias-aware search the scorer uses, and every employer or job title must appear in the source; anything else is stripped and reported. A test drives a deliberately hallucinating response: three invented skills and a fabricated "Senior ML Engineer at Google" role are all removed, and only what the resume actually supports survives. The UI shows what was stripped, so the user can see both that the tool caught the fabrication and that everything remaining is backed by their real experience - which matters on a document they put their name to. The writer is also now given the JD's required keywords explicitly, with the instruction to use only those the resume genuinely supports: a missing skill stays missing. Generation already runs on the quality model.
…ma braces ATS optimise has never worked. The generation prompt embeds a literal JSON schema to show the model the expected shape, and the code built it with str.format(), which reads every brace in that schema as a placeholder. The call raised KeyError before it ever reached the LLM. The bug predates the contact and verification work; adding a "contact" block to the schema only changed which key appeared in the error. Placeholders are now substituted with str.replace(), so the literal braces of the JSON schema pass through untouched. Verified end to end against the live API: check returns 75 percent coverage with kubernetes correctly reported missing, optimise returns 200 with the contact block populated, the skills list containing only what the resume evidences, and the DOCX rendering to a real document.
Two-column resumes normally break ATS parsing because they are built with a table, and a parser reads a grid cell by cell - which is how job titles end up interleaved with skills, or dropped. This uses Word's column layout instead (the w:cols element on the section), which is a rendering instruction only: the paragraphs stay in linear order in the XML, so a text extractor still walks them top to bottom. The header and summary stay full width, then a continuous section runs two columns with a column break between the sidebar (Skills, Education) and the main content (Work Experience). Verified by reading the generated file back: zero tables, every standard heading present, and the reading order intact. Single column remains the default and the safest option; the layout is a request parameter, validated, with both buttons in the UI. Adds Anthropic as a fourth provider. The Messages API differs from the OpenAI chat API in ways the provider module hides: x-api-key auth with a version header, a required max_tokens, and no response_format json_object - JSON is forced by prefilling the assistant turn with an opening brace and prepending it back onto the reply. Claude Sonnet 5 is wired as its quality model, Haiku 4.5 as the bulk model, and gpt-5 is added to the OpenAI list as the stronger option there. Set ANTHROPIC_API_KEY in .env and select the provider in Settings to use it.
Model selection was per-provider: both tiers had to come from whichever vendor
was active, so running cheap bulk work on one and generation on another was
impossible. Tasks now carry their own provider as well as their model, declared
in task_models in config.yaml:
bulk one JD extraction per fetched job, one relevance call per 30 titles.
Hundreds per run, so it must stay fast and cheap.
quality the profile summary, the ATS resume rewrite, and the responsibility
coverage judge - read by the user, sent to an employer, or deciding a
single job's score.
Callers name the task rather than looking a model up themselves, so the routing
lives in one place. A task pointed at a provider whose API key is absent logs a
warning and falls back to the active provider, so a half-configured deployment
degrades instead of failing every request. An explicit admin pin in Settings
still overrides everything.
Verified live with bulk on OpenAI and quality on Anthropic in the same run:
gpt-4o-mini answered the bulk call in 1.2s, claude-sonnet-5 the quality call in
1.4s.
Two real Anthropic API constraints surfaced while testing and are now handled.
The current models reject `temperature` outright, and they reject the
assistant-prefill trick used to force JSON, so temperature is not sent and JSON
is requested in the prompt instead, with the router's existing fence-stripping
parser cleaning up whatever survives. Both failures had been caught by the Groq
fallback, which is the behaviour intended, but they would have meant every
quality call silently degrading to llama.
The deterministic seniority net - the only reliable layer, since the LLM is inconsistent and skips titles in a batch - caught 3 of 19 realistic senior titles. It knew a handful of English words and nothing else. Now catches the ways senior roles actually hide from a keyword filter: level numbers (Engineer II, SDE 3, L5), compound leads (Teamlead, TechLead), missing seniority words (expert, chief, VP, vice president, distinguished, fellow), and German experience markers (Abteilungsleitung, Erfahrener, mit Berufserfahrung). 25 of 25 test titles caught. Version numbers are not seniority: a bare level digit counts only right after a role noun (Engineer 3), never after a technology (Python 3, Angular 2, Web3, S3). An explicit entry word (junior, graduate, Werkstudent, Azubi) always overrides, so Junior Engineer II is not misread as senior. Zero false positives across the version-number and entry-word traps. The cheap pre-LLM keyword gate (config exclude_keywords) gains the same words so more seniors are dropped before any token is spent.
…title The title-only fetch gate cannot tell whether a plain 'ML Engineer' posting is actually entry level - the body might say '5+ years' or the job_level might be 'senior'. But by the time a job is scored, its full JD has been extracted, so that evidence is already in hand. A second entry gate now runs after scoring, when entry-only is on: if the extracted JD names a senior job_level (senior/lead/principal/staff/director) or asks for more years than an entry candidate has (max_experience_years, default 2), the job is dropped and its id blocked from re-appearing - even though it was already scored. mid-level within the year bar is kept, since many such roles accept strong juniors and the years check is the real bar. Deterministic and free: it reads data the scorer already produced, no extra LLM call. This is what catches the seniors that slip past a title that gives nothing away.
Many JDs never write a year count. They say 'several years of experience', 'proven track record', 'mehrjaehrige/einschlaegige Berufserfahrung', or 'fundierte Kenntnisse' - and the numeric parser returned None on all of them, so the post-score entry gate let them through. The gate now also matches these non-numeric seniority phrases (EN and DE, including ae transliterations of umlauts), scanning both the extracted experience_requirements and the job summary. An explicit entry phrase always overrides - 'no prior experience', 'Berufseinsteiger willkommen', 'erste Berufserfahrung', 'graduates welcome' - so first-job postings that mention the word 'experience' are not misread as senior. A stated number stays authoritative: '1 year of experience' is entry and never triggers the phrase rule. Nine phrasings caught, zero false positives across the entry traps. The JD extraction prompt is updated to capture these non-numeric phrases into experience_requirements verbatim rather than dropping anything without a digit, so the gate has them to read.
…e + nl
Adzuna already fetched any country via its per-country endpoint, but the default
for a user who never chose their own was a single country ('de'). Netherlands
worked only if the user typed it into Search Criteria.
The default is now a list, default_countries in config.yaml, seeded with de and
nl. get_countries returns it for users with no saved choice; users who set their
own countries are unaffected, and Arbeitnow/Bundesagentur still only run when de
is included since they are German boards.
Adzuna was queried with its default relevance ordering and no age limit, which broke the incremental-fetch model in two ways: a fresh posting ranked below the page window by relevance was never seen at all, while stale jobs we would discard in the recency filter anyway occupied slots in it. sort_by=date and max_days_old=MAX_AGE_DAYS make the window mean the newest jobs within our 30-day rule. Repeated runs walk new postings from the top and the seen_jobs dedup makes each run pay only for the genuinely new ones. Verified live: newest-first, all within the window, and the server-side age filter shrank one query pool from 780 relevance-ranked jobs to 113 recent ones.
The Adzuna window was a fixed per-title budget (50 of possibly hundreds), and every run re-fetched the same top slice: jobs ranked below the window were never seen by anyone, ever. With results now date-sorted, exhaustive coverage becomes cheap to guarantee. fetch_adzuna_jobs gains a seen-stop mode: given the user's already-processed ids, it walks the date-sorted pool page by page and stops at the first page containing nothing new - everything older was covered by an earlier run. The first run walks the whole 30-day pool (bounded by max_pages_per_query, default 10 pages / 500 jobs per query); each later run downloads only the pages holding new postings. Budget mode (seen_ids=None) is unchanged for existing callers and tests. Wired through fetch_combined into both the manual run route and the scheduler, which pass the user's seen_jobs ids. Verified live: run 1 fetched all 115 jobs of a query pool across 3 pages; run 2 downloaded one page, found nothing new, and stopped with zero LLM cost. A unit test pins the walk/stop/new-jobs-at-top behaviour with a stubbed API.
The scheduler checked has_resume/get_run_status and then started a fetch without claiming the run slot, so a manual fetch begun during the slow paginated fetch could pass begin_run() and start a second scoring loop. Both loops then incremented the same status["checked"], driving it past status["total"] (the "688 of 446 checked" symptom). begin_run() now does the check-and-set under a lock, the scheduler claims the slot via begin_run() before fetching and releases it on failure, and discover_and_score() refuses to enter when a scoring loop is already in an active phase for that user.
Drops the ATS Check entry from the sidebar navigation. The ATS routes and tab component stay in place; only the nav entry point is removed.
… section Updates the README to match the current engine: five scored sections plus two pass/fail gates instead of seven weighted dimensions, no related-skill partial credit, and the task-tiered call_llm() router. Adds an Architecture section with a diagram (docs/architecture.png) and a docs/README.md explaining where the diagram lives and why it must not go under images/.
There was a problem hiding this comment.
🟡 Changes recommended
The review found a few verified correctness/documentation issues in changed code paths that should be fixed before approval.
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Pull request overview
This PR substantially rebuilds JobsFitAI’s matching/scoring pipeline around actionable sections plus hard pass/fail gates, introduces task-routed LLM calls with an Anthropic provider option, upgrades ATS “optimise” into a verified full-resume generator with DOCX layout options, and adds multiple security/privacy hardening measures for the public beta.
Changes:
- Reworked matcher: five scored sections + deterministic gates for experience/language, plus optional LLM-judged responsibility coverage.
- Added task-based LLM routing (
call_llm(task=...)) with provider/model tiering and a new Anthropic provider; improved handling for OpenAI reasoning models. - ATS optimisation now generates a full resume, verifies output against the source, and supports single- and two-column DOCX rendering; plus beta security/privacy improvements (CSP/HSTS/HTTPS enforcement, account erasure, no PII in JWT).
File summaries
| File | Description |
|---|---|
| README.md | Updates product description, routing, architecture, and security/deployment documentation. |
| docs/README.md | Adds docs folder guidance and architecture diagram conventions. |
| frontend/src/lib/errors.js | Updates user-facing auth error text. |
| frontend/src/lib/auth.js | Removes JWT-decoding helper; relies on /api/auth/me for email. |
| frontend/src/components/TopBar.jsx | Fetches and displays user email via authenticated API call. |
| frontend/src/components/tabs/Settings.jsx | Adds GDPR account deletion UI modal and flow. |
| frontend/src/components/tabs/ATS.jsx | Adds DOCX layout selection and displays removed unsupported claims. |
| frontend/src/components/Sidebar.jsx | Removes ATS Check tab from navigation. |
| frontend/src/components/AnalysisResults.jsx | Reworks analyzer UI: gates panel, evidence-based skills, per-duty coverage, score math, rich-text rendering. |
| backend/tests/test_matcher.py | Expands matcher tests for gates, experience FTE math, LLM routing tiers, and coverage behavior (stubbed). |
| backend/tests/test_contract.py | Adds contract tests for ATS verification, DOCX layouts, account deletion, JWT claims, and CSP assertions. |
| backend/services/resume_rewriter.py | Routes resume-improvement LLM calls through the quality tier. |
| backend/services/prompts/schemas/resume_schema.json | Removes contact fields from resume extraction schema. |
| backend/services/prompts/resume_prompt.py | Adds strict “never extract contact details” rule and refines skill extraction constraints. |
| backend/services/prompts/jd_prompt.py | Tightens skills extraction to exclude “traits” and expands experience requirement capture (numeric + non-numeric). |
| backend/services/profile_summary.py | Adapts summary generation to gates (relevant vs total years) and new match output shape; routes via quality tier. |
| backend/services/matcher/skill_aliases.py | Adds line-level evidence extraction for UI quoting. |
| backend/services/matcher/scores/skills.py | Removes related-skill partial credit; adds evidence mapping and “similar rephrasing” only. |
| backend/services/matcher/scores/responsibilities.py | Removes cosine-similarity responsibilities scorer (replaced by LLM-judge module). |
| backend/services/matcher/scores/languages.py | Converts language logic into parsing helpers (consumed by gates), removes scoring. |
| backend/services/matcher/scores/experiences.py | Removes experience scorer (replaced by gates + FTE math). |
| backend/services/matcher/scores/init.py | Updates exports to scored sections only (no experience/languages/responsibilities scorer). |
| backend/services/matcher/responsibility_coverage.py | Adds LLM-judged per-duty responsibility coverage module. |
| backend/services/matcher/gates.py | Adds deterministic gates: relevant experience years and language proficiency; includes seniority detection logic. |
| backend/services/matcher/experience.py | Introduces FTE-weighted, non-double-counting experience duration math. |
| backend/services/matcher/engine.py | Reworks matcher orchestration to scores + gates; responsibility coverage becomes opt-in. |
| backend/services/llm/providers/openai.py | Adds reasoning-model detection and correct parameterization/token budgeting; removes key hint exposure. |
| backend/services/llm/providers/groq.py | Removes key hint exposure. |
| backend/services/llm/providers/anthropic.py | Adds Anthropic provider implementation for quality tier. |
| backend/services/llm/ollama_health.py | Removes unused Ollama health module. |
| backend/services/llm/caller.py | Adds task routing (provider/model selection) and supports Anthropic in provider resolution. |
| backend/services/job_relevance.py | Strengthens deterministic seniority detection with structural patterns and explicit entry-word override. |
| backend/services/job_matcher.py | Fixes run-status race via locking; adds seen-stop pagination support and post-score entry-level gate. |
| backend/services/fetchers/job_fetcher.py | Adds Adzuna newest-first + max-age filters and “seen-stop” pagination, bounded by config. |
| backend/services/extractors/resume_extractor.py | Reuses shared experience-math module and clarifies extraction model intent. |
| backend/services/extractors/jd_extractor.py | Clarifies extraction as cheap/bulk task behavior. |
| backend/services/ats.py | ATS generation becomes full-resume JSON + verification against source; adds contact header in plain-text render. |
| backend/schemas/auth.py | Adds request schema for account deletion. |
| backend/schemas/ats.py | Adds layout field for DOCX generation request. |
| backend/repositories/settings_repo.py | Switches default countries to configurable list. |
| backend/models/user.py | Adds hard-delete account erasure across all per-user tables and cached analyses. |
| backend/main.py | Enforces production HTTPS/CORS safety, adds CSP/HSTS/Permissions-Policy, and hardens scheduler run claiming. |
| backend/core/state.py | Adds task model routing and Anthropic provider support; removes key hint from provider catalog. |
| backend/core/security.py | Removes email from JWT payload; token contains only user id (sub). |
| backend/core/database.py | Stops logging Turso endpoint URL/host details. |
| backend/core/config.py | Adds Anthropic config, task routing config, default countries list, and Adzuna paging cap. |
| backend/core/config_validator.py | Updates matcher weights schema to scored sections only (gates excluded). |
| backend/api/routes/resume_analyzer.py | Updates scoring version; enables LLM judge for analyzer; returns job context, score math, evidence, gates, duties, and skill impact. |
| backend/api/routes/job_matches.py | Enables seen-stop pagination and enables LLM judge for single-job scoring endpoint. |
| backend/api/routes/auth.py | Adds account deletion endpoint, clears in-memory state, removes email from JWT creation. |
| backend/api/routes/ats_maker.py | Adds single/two-column DOCX generation with columns (no tables) and contact header. |
| .env.example | Documents production safety requirements and optional legacy password gate variables. |
Review details
- Files reviewed: 52/53 changed files
- Comments generated: 3
- Review effort level: Lite
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Comment on lines
+141
to
+143
| <!-- Diagram not committed yet. Export your architecture image to docs/architecture.png | ||
| (create the docs/ folder at the repo root). Do NOT use a folder literally named | ||
| "images" - .gitignore ignores it. SVG also works: docs/architecture.svg. --> |
| for skill in parsed.get("skills") or []: | ||
| if not isinstance(skill, str) or not skill.strip(): | ||
| continue | ||
| if found_in_corpus(skill, corpus) or skill.lower().strip() in corpus: |
| "blocking_count": 2, | ||
| }, | ||
| "matched_required": ["python", "git"], | ||
| "partial_required": ["tensorflow"], # related skill, half credit |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Rebuilds the scoring engine around actionable sections plus hard gates, routes
every LLM call to a per-task provider/model tier (adding an Anthropic provider),
reworks the ATS optimiser into a verified full-resume generator, and lands a
round of security and privacy hardening for the public beta. 24 commits.
Matcher and scoring
responsibilities, preferred skills, education, certifications); move years of
experience and language level to hard pass/fail gates instead of weighted
dimensions.
full-time-equivalent years.
because the embedding model cannot reliably separate a related skill from a
coincidentally similar one.
overlap.
stated without a number, and a post-score entry gate that reads the extracted
JD rather than the title.
matcher/gates.py,matcher/experience.py,matcher/responsibility_coverage.py; old per-section scorers removed.LLM routing and providers
call_llm()routes each task to its own provider and model tier. High-volumeextraction runs on a cheap fast model; low-volume, user-facing or
single-score work runs on a stronger model. A tier with no API key falls back
to the active provider instead of failing.
health module.
ATS
against the source, adds a contact header, renders a two-column DOCX without
tables.
format()consumed the JSONschema braces.
unchanged).
Job fetch
no job in the age window is skipped.
run-status counter past its total.
Analyzer UI
is worth.
match()shape, which was failing everyanalysis.
Security and privacy
Docs and tests
call_llm();new Architecture section with a diagram under
docs/.per-duty coverage, tiered routing, and the None-sentinel engine paths.