feat: classify finding severity with custom rubrics - #791
Open
kmbroai wants to merge 5 commits into
Open
Conversation
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
ianw-oai
approved these changes
Sep 3, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Let CLI and SDK users classify existing findings under a custom severity rubric and resume interrupted classification without repeating successful model calls. Each successful finding is checkpointed in SQLite before the next finding starts. Omitting a rubric inherits the original severity without a model call.
Changes
codex-security classify-severitywith exactly one of--scan ID(including unique prefixes andlatest) or--scan-dir PATH, optional--rubric PATH, repeatable--knowledge-base PATHand--finding-id ID, and existing model/effort controls.--reprocess(default false) and SDKreprocess: trueto bypass matching checkpoints for the selected findings. Normal runs reuse assessments only when evidence, rubric, and context hashes match. Exclusions are reusable assessments with null severity.classifySeverityprimitive and the database-backedclassifyScanSeverityandclassifyScanDirectorySeveritywrappers. Each rubric evaluation uses the supplied report/context in a separate read-only Codex turn without tools or source inspection.finding_severity_assessmentsstores the latest assessment per finding ID, andscan_severity_classificationsstores each scan's requested selection and policy/context hashes. Save each successful row independently; a failed reassessment retains that finding's previous row. Preserve original findings and sealed artifacts.severity-classification.jsonafter a successful run. Publication rejects incomplete or stale selections, omits exclusions, and uses assessed severity for ticket priority/title. SDK callers can still supply an assessment explicitly.publish scan --to linear --finding-id IDfor direct selection while retaining full sealed-scan membership checks. Update help, SDK types, documentation, package fixtures, and bundled-file contracts.Testing
pnpm run typesandpnpm run format: passed.pnpm packandpnpm run check:package -- <tarball>: passed. Validated 411 archive entries, installed exports and strict NodeNext consumer types, CLI/SDK lifecycle, credential locking, plugin/MCP startup, and nested worker startup.classify-severity --helpandgit diff --check: passed.Risk and rollout
This adds public CLI/SDK surface and an additive local database migration. Classification of external scan directories also uses the configured local state database. Existing JSON-only assessments require one classification run to populate checkpoints; the JSON export is no longer read as an authoritative assessment. Classification and publication must use the same state directory.
Reprocessing replaces successful rows individually and retains rows outside the selected set. There is no assessment history or whole-scan rollback. Changing only the model or effort requires
--reprocess. StandaloneclassifySeverityremains an in-memory operation. Original scan severity and existing Linear tickets are unchanged;--skip-existingretains its current behavior.Public disclosure review