diff --git a/.agents/skills/insight-add-entity-pattern/SKILL.md b/.agents/skills/insight-add-entity-pattern/SKILL.md new file mode 100644 index 0000000..6acd81d --- /dev/null +++ b/.agents/skills/insight-add-entity-pattern/SKILL.md @@ -0,0 +1,9 @@ +--- +name: insight-add-entity-pattern +description: Add an Insight_Extractor regex entity across PatternLabel, patterns, tests and documentation. Use only in cgfixit/Insight_Extractor. +--- + +Read the repository's root `AGENTS.md` and its delegated `.codex/AGENTS.md`. +Then follow [the maintained workflow](../../../.codex/skills/add-entity-pattern/SKILL.md). +Resolve command paths from the repository root and use its activated virtual environment. +Report actual verification results and explicit skips; preserve existing user authorization. diff --git a/.agents/skills/insight-optimize/SKILL.md b/.agents/skills/insight-optimize/SKILL.md new file mode 100644 index 0000000..4b7e2e3 --- /dev/null +++ b/.agents/skills/insight-optimize/SKILL.md @@ -0,0 +1,9 @@ +--- +name: insight-optimize +description: Measure and optimize an Insight_Extractor hot path while preserving output, spans and lazy loading. Use only in cgfixit/Insight_Extractor. +--- + +Read the repository's root `AGENTS.md` and its delegated `.codex/AGENTS.md`. +Then follow [the maintained workflow](../../../.codex/skills/optimize/SKILL.md). +Resolve command paths from the repository root and use its activated virtual environment. +Report actual verification results and explicit skips; preserve existing user authorization. diff --git a/.agents/skills/insight-preflight/SKILL.md b/.agents/skills/insight-preflight/SKILL.md new file mode 100644 index 0000000..f7c2d66 --- /dev/null +++ b/.agents/skills/insight-preflight/SKILL.md @@ -0,0 +1,9 @@ +--- +name: insight-preflight +description: Validate Insight_Extractor lint, types, offline unit tests and staging before commit or push. Use only in cgfixit/Insight_Extractor. +--- + +Read the repository's root `AGENTS.md` and its delegated `.codex/AGENTS.md`. +Then follow [the maintained workflow](../../../.codex/skills/preflight/SKILL.md). +Resolve command paths from the repository root and use its activated virtual environment. +Report actual verification results and explicit skips; preserve existing user authorization. diff --git a/.agents/skills/insight-verify-no-model/SKILL.md b/.agents/skills/insight-verify-no-model/SKILL.md new file mode 100644 index 0000000..8692473 --- /dev/null +++ b/.agents/skills/insight-verify-no-model/SKILL.md @@ -0,0 +1,9 @@ +--- +name: insight-verify-no-model +description: Verify Insight_Extractor regex, keyword and fake-model orchestration without downloading models. Use only in cgfixit/Insight_Extractor. +--- + +Read the repository's root `AGENTS.md` and its delegated `.codex/AGENTS.md`. +Then follow [the maintained workflow](../../../.codex/skills/verify-no-model/SKILL.md). +Resolve command paths from the repository root and use its activated virtual environment. +Report actual verification results and explicit skips; preserve existing user authorization. diff --git a/.codex/AGENTS.md b/.codex/AGENTS.md index 0f18d66..0a6aef9 100644 --- a/.codex/AGENTS.md +++ b/.codex/AGENTS.md @@ -20,10 +20,11 @@ Important paths: - `src/insight_extractor/` — package code; - `tests/unit/` — model-free unit tests; -- `tests/integration/` — real-model tests; +- `tests/integration/` — optional legacy tests; inspect mocks and API compatibility before running; - `.github/workflows/ci.yml` — required lint, type, unit, and smoke gates; - `requirements.txt`, `constraints.txt`, `pyproject.toml` — dependency declarations and pins; -- `.codex/` — Codex instructions and project skills; +- `.codex/` — Codex instructions and workflow references; +- `.agents/skills/` — discoverable, repository-specific skill entrypoints; - `.claude/skills/` — existing detailed implementation checklists. ## Setup and validation @@ -32,7 +33,7 @@ Use a Python 3.12+ virtual environment when possible: ```powershell python -m pip install -r requirements.txt -c constraints.txt -python -m pip install -e ".[dev]" +python -m pip install -e ".[dev]" -c constraints.txt ``` Required gates: @@ -44,7 +45,13 @@ python -m mypy src/insight_extractor python -m pytest tests/unit/ -v --tb=short ``` -The CI smoke path must remain model-free. Real BERT downloads belong only to `tests/integration/`, manual workflow dispatch, or commits containing `[run-integration]`. +The CI smoke path uses fake BERT boundaries and production CLI orchestration. +Optional integration CI runs on workflow dispatch or a qualifying push head commit +containing `[run-integration]`; a PR commit message alone does not enable it. +The current integration files contain stale mocks/API calls and do not establish +real-model compatibility. Run with offline guards during diagnosis and report failures. + +Historical manual corrections: the live constraints already include the accelerate compatibility fix; `load_state()` invalidates embeddings lazily rather than recomputing them. Check current source and manifests before following historical sequences in `CLAUDE.md`. ## Non-negotiable invariants @@ -69,9 +76,9 @@ The CI smoke path must remain model-free. Real BERT downloads belong only to `te Read `.codex/README.md` and `.codex/codex_custom_instructions.md`, then use the relevant skill: -- `preflight` — before commit/push; -- `verify-no-model` — validate pipeline changes without downloading BERT; -- `add-entity-pattern` — add a static regex entity end to end; -- `optimize` — measured, minimal optimization of a real hot path. +- `insight-preflight` — before commit/push; +- `insight-verify-no-model` — validate pipeline changes without downloading BERT; +- `insight-add-entity-pattern` — add a static regex entity end to end; +- `insight-optimize` — measured, minimal optimization of a real hot path. For fuller implementation checklists, consult the matching file under `.claude/skills/`. diff --git a/.codex/README.md b/.codex/README.md index 746e940..e2be4bd 100644 --- a/.codex/README.md +++ b/.codex/README.md @@ -7,20 +7,36 @@ Codex-specific onboarding for `Insight_Extractor`. - `AGENTS.md` — detailed project guidance loaded through the root `AGENTS.md` entrypoint. - `codex_custom_instructions.md` — review-first behavior, minimal diffs, and secret safety. - `ponytail-plugin.json` — existing optional Ponytail metadata. -- `skills/` — Codex-native project workflows. +- `skills/` — maintained workflow references. +- `../.agents/skills/` — discoverable repository skill entrypoints with `insight-` names. The root `AGENTS.md` is Codex's auto-discovered entrypoint and delegates here. `CLAUDE.md` remains the detailed historical operating manual. Do not duplicate runtime architecture or dependency pins here. +## Discovery + +Codex discovers repository skills under `.agents/skills`; `.codex/skills` alone is +not the documented repository discovery location. The four thin entrypoints link +to the maintained workflows here, avoiding duplicate checklists and collisions with +personal skills such as `preflight` and `optimize`. See the +[official skill discovery documentation](https://developers.openai.com/codex/skills/). +Open a new task in this checkout if the new skills are missing from the current catalog. + +`skills/Insight-Optimizer/SKILL.md` is an optional extended optimization reference +(`name: insight-optimizer`); `skills/fable-protocol/SKILL.md` is historical owner +guidance, not a default project entrypoint. Its cross-project and model claims are +not evidence about this codebase. `ponytail-plugin.json` is reference metadata, +not proof that a plugin or hooks are installed. + ## Skill dispatch | Skill | Use it for | |---|---| -| `preflight` | Ruff, formatting, strict mypy, unit tests, smoke checks, and staging hygiene before a commit or push | -| `verify-no-model` | Regex/dynamic/full-fake pipeline verification without a HuggingFace download | -| `add-entity-pattern` | Adding a new static regex entity and synchronizing enum, pattern, tests, docs, and smoke coverage | -| `optimize` | A small, measured improvement to regex, stemmer, state, or dynamic-expansion hot paths | +| `insight-preflight` | Ruff, formatting, strict mypy, unit tests, smoke checks, and staging hygiene before a commit or push | +| `insight-verify-no-model` | Regex/dynamic/full-fake pipeline verification without a HuggingFace download | +| `insight-add-entity-pattern` | Adding a new static regex entity and synchronizing enum, pattern, tests, docs, and smoke coverage | +| `insight-optimize` | A small, measured improvement to regex, stemmer, state, or dynamic-expansion hot paths | The skills reuse the existing `.claude/skills/` checklists where those are more detailed. The Codex versions are the entrypoints and add Windows-friendly commands plus explicit stop conditions. diff --git a/.codex/skills/Insight-Optimizer/SKILL.md b/.codex/skills/Insight-Optimizer/SKILL.md index 69f8baf..e6010d4 100644 --- a/.codex/skills/Insight-Optimizer/SKILL.md +++ b/.codex/skills/Insight-Optimizer/SKILL.md @@ -1,5 +1,5 @@ --- -name: optimize +name: insight-optimizer description: Make 1-3 measured, minimal optimization to Insight_Extractor's regex, dynamic-stemmer, state, or keyword-expansion paths without changing public behavior or lazy model loading. --- @@ -8,7 +8,7 @@ description: Make 1-3 measured, minimal optimization to Insight_Extractor's rege This is a measured optimization workflow, not a broad refactor. Target repo: `CGFixIT/Insight_Extractor` (Python 3.12+, Pydantic v2, lazy SentenceTransformer). Source of truth order: live source → `CLAUDE.md` → `docs/SPEC.md` → `README.md`. -Always start from `main`. Trace every caller of a shared helper before touching it. +Start a new change from current `origin/main` on a `codex/` branch; preserve an existing task branch. Trace every caller of a shared helper before touching it. ## Hard Constraints (never violate) @@ -18,7 +18,7 @@ Always start from `main`. Trace every caller of a shared helper before touching - Do **not** remove or weaken lazy model loading (`model` / `tokenizer` properties). - Do **not** break exact match order, deduplication, span fidelity, or exception paths. - Preserve the keyword-mutation sequence after any change: - `stemmer.set_keywords(...)` → `registry.regenerate_dynamic_patterns(...)` → `_recompute_keyword_embeddings()` → `_auto_categorize_keywords()`. + Follow the current mutation path in `update_thread_keywords`; `load_state` rebuilds regex runtime and marks embeddings dirty without loading a model. Do not make state loading eager. - Line length ≤ 100. Prefer `pathlib.Path` + `encoding="utf-8"`. No f-strings in log calls. - Use `# ponytail:` comments for deliberate ceilings. @@ -94,3 +94,4 @@ Do not manufacture a cache, configuration knob, or “money-mode” behavior to # ponytail: TfidfVectorizer rebuilt every update_thread_keywords call; # rolling corpus only after multi-document expansion is measured +``` diff --git a/.codex/skills/preflight/SKILL.md b/.codex/skills/preflight/SKILL.md index 2ceae94..876ce91 100644 --- a/.codex/skills/preflight/SKILL.md +++ b/.codex/skills/preflight/SKILL.md @@ -13,12 +13,15 @@ Run from the repository root in PowerShell: python -m ruff check src/ tests/ python -m ruff format --check src/ tests/ python -m mypy src/insight_extractor +$env:HF_HUB_OFFLINE = "1" +$env:TRANSFORMERS_OFFLINE = "1" +$env:HF_HUB_DISABLE_TELEMETRY = "1" python -m pytest tests/unit/ -v --tb=short git diff --check git status --short ``` -For changes to `config.py`, `constants.py`, packaging, or import boundaries, also run the model-free smoke check from the detailed checklist. Do not run the CLI from the repository root because it writes `insights_extracted.md` and `insight_extractor_state.json`. +For changes to `config.py`, `constants.py`, packaging, or import boundaries, also run the lightweight import check from the detailed checklist. That check is not the full CI smoke: inspect the `smoke-test` job in `.github/workflows/ci.yml`, which runs production CLI orchestration with fake model/tokenizer boundaries and verifies the report and state. Do not run the CLI from the repository root because it writes `insights_extracted.md` and `insight_extractor_state.json`. Before staging, reject generated artifacts: diff --git a/.codex/skills/verify-no-model/SKILL.md b/.codex/skills/verify-no-model/SKILL.md index 106453a..7c678a4 100644 --- a/.codex/skills/verify-no-model/SKILL.md +++ b/.codex/skills/verify-no-model/SKILL.md @@ -11,7 +11,9 @@ Rules: - Run from a disposable directory, never the repository root. - Regex and dynamic keyword paths must not access `.model` or `.tokenizer`. -- For full-pipeline tests, inject `extractor._model` and `extractor._tokenizer` before calling `extract()` or `load_state()`. +- `load_state()` must stay model-free; when a saved model name changes it clears both cached model and tokenizer. For subsequent full-pipeline checks, inject fake `_model` and `_tokenizer` after loading state. +- Model-free means no model download, not no ML dependencies: importing `extractor.py` imports `sentence_transformers`. +- Set `HF_HUB_OFFLINE=1` and `TRANSFORMERS_OFFLINE=1` for offline verification. - Do not use `seed_keywords=[]` to mean “no keywords”; the implementation treats falsy input as the default seed bank. - Assert concrete values such as a known CVE, IP, or entity, not merely non-empty output. - Confirm the Markdown report and JSON state round-trip in the disposable directory, then confirm the repository has no generated artifacts. @@ -29,4 +31,4 @@ try { } ``` -Escalate to `tests/integration/` only when changing model loading, real tokenization, embedding behavior, dependency compatibility, or model-path error handling. Use `[run-integration]` or workflow dispatch for that gate, and report it explicitly if skipped. \ No newline at end of file +Escalate to `tests/integration/` only when changing model loading, real tokenization, embedding behavior, dependency compatibility, or model-path error handling. Inspect the optional suite for stale mocks before relying on it. Workflow dispatch enables its CI job; `[run-integration]` only works on qualifying push head commits, not ordinary PR events. Report failures or skipped model verification explicitly. \ No newline at end of file diff --git a/README.md b/README.md index a3eeef5..102dd96 100644 --- a/README.md +++ b/README.md @@ -338,7 +338,7 @@ Insight_Extractor/ ├── requirements.txt # Runtime deps with transformers compatibility note ├── constraints.txt # Pinned known-good versions ├── README.md # This file -├── SPEC.md # Full technical specification +├── docs/SPEC.md # Technical design reference; source is behavioral authority ├── src/ │ └── insight_extractor/ │ ├── __init__.py # Package entry point with lazy imports @@ -366,6 +366,12 @@ Insight_Extractor/ --- +## Codex setup + +Start at [AGENTS.md](AGENTS.md) and [the Codex workflow map](.codex/README.md). +Repository skills are discoverable under `.agents/skills/` with `insight-` names. +See [the setup review](docs/CODEX_SETUP_REVIEW.md) for verified scope and follow-ups. + ## Development Setup ```cmd diff --git a/docs/CODEX_SETUP_REVIEW.md b/docs/CODEX_SETUP_REVIEW.md new file mode 100644 index 0000000..13ae5aa --- /dev/null +++ b/docs/CODEX_SETUP_REVIEW.md @@ -0,0 +1,100 @@ +# Codex setup review + +Reviewed 2026-09-06 against origin/main at +`8097406` (before this documentation-only setup change). + +## Architecture and scope + +The packaged application is `src/insight_extractor`, installed with hatchling and +exposed through `insight-extract` and `python -m insight_extractor`. +The review covered package modules, unit and integration tests, README, SPEC, +dependency manifests, CI/security workflows, and every `.codex` file. +`docs/insight_extractor.py` is a separate legacy script with working-directory +inputs; its Porter-like logic is not the packaged stemmer implementation. + +- `__init__` lazily exposes classes; lightweight models/config/constants/utilities + remain usable without importing torch. `extractor.py` itself imports ML libraries. +- The stemmer generates regexes, compiles growing keyword banks incrementally, + resolves original keywords, and maintains match spans. Registry regeneration is + explicit; it is not an observer of arbitrary stemmer mutations. +- `extract()` expands keywords with TF-IDF, prepares embeddings, then combines + static/dynamic regex, semantic hits, sentence scores, and keyword statistics. +- Reports escape Markdown and HTML. State loading restores settings and invalidates + embeddings lazily; changing the saved model name also clears model/tokenizer objects. + +## Setup corrections + +- Added four `insight-` skill entrypoints under `.agents/skills`, linking to existing + `.codex/skills` workflows. This preserves maintained references and avoids collisions + with generic personal skills. Repository discovery is documented in the + [official Codex skills guide](https://developers.openai.com/codex/skills/). +- Gave the extended optimizer a distinct name and closed its unfinished code fence. +- Corrected offline verification around state restoration and clarified that the + lightweight regex check does not reproduce the production CLI smoke job. +- Documented optional integration trigger limitations and stale test fixtures. +- Kept constraints on the editable development install in Codex onboarding. + +The root AGENTS entrypoint already delegates to `.codex/AGENTS.md`. No project +config file or additional plugin installation is needed for these workflows. +Existing Ponytail JSON is reference metadata, not an installed plugin declaration. +Open a new task if its skill catalog does not show the new entrypoints. + +## Follow-up findings + +1. **Tokenizer termination and coverage:** `SentenceTokenizer.chunk_text` advances + by `max_tokens - overlap` without validating either argument. A long input with + equal budget/overlap never advances; `tokenize_sentences` fixes overlap at 50, + so small caller budgets also need review. Existing tokenizer tests cover normal + chunking but do not cover invalid budget/overlap combinations. + Add a bounded fake-tokenizer regression, then define validation and test normal, + zero, negative, and oversized overlap without downloading a model. +2. **Integration fixtures have drifted:** both integration files patch + `insight_extractor.tokenizer.AutoTokenizer`, which is only imported under + `TYPE_CHECKING` at module scope. `test_e2e.py` additionally uses unsupported + constructor flags and obsolete result fields. Some extraction occurs after + patch contexts exit. Repair fixtures against current signatures, use correctly + sized deterministic vectors, and distinguish offline orchestration from any + deliberate real-model test. A passing required CI gate does not certify this suite. +3. **State validation deserves a separate fix:** syntactically valid JSON with a + non-object root or invalid enum values can escape `StateLoadError`; field-by-field + assignment can leave partially changed state. Validate before mutation and test + failure atomicity while retaining lazy loading and the legacy category mapping. +4. **Documentation debt:** README advertises Porter/lemmatization and character-error + fuzzy matching, while the packaged stemmer uses suffix expansion and substring + regex matching. It also references a missing LICENSE. Packaging URLs use a hyphen + instead of the actual repository underscore. Historical CLAUDE/SPEC details and + security-audit conclusions require rechecking against current source/advisories. + This review is not a fresh dependency vulnerability audit. + +## Suggested next skills + +- `insight-tokenizer-check`: reproduce chunk-budget/overlap failures with an injected + tokenizer, enforce termination and token preservation, and verify lazy imports. +- `insight-state-contract`: validate state round trips, malformed schema handling, + failure atomicity, legacy categories, relative-path behavior, and model invalidation. + +These are proposals for separate changes, not newly implemented runtime behavior. + +## Local validation + +The editable package was installed in `.venv` with system packages disabled and +`-c constraints.txt`. The available interpreter is Windows Python 3.12.0rc3; +this is not evidence of a clean stable-Python 3.13 installation or real-model inference. + +- `pip check`: no broken requirements. +- Ruff lint and format: passed. +- Strict mypy: passed, 10 source files. +- Offline unit tests: 118 passed. +- Optional E2E diagnostic (`pytest tests/integration/test_e2e.py -x -q --tb=short`): + fails at fixture setup with `AttributeError: module 'insight_extractor.tokenizer' + has no attribute 'AutoTokenizer'`; stopped after the first failure. No real-model + download was attempted. The optional suite remains a separate repair task. +- Production CI smoke Python body, extracted from `ci.yml` and run in an ignored + scratch directory: passed (console, escaped Markdown, keyword expansion, state). +- Four new skill entrypoints: metadata validation passed; every workflow link resolves. +- Lightweight root import resolves to this checkout and does not import torch. +- `git diff --check`: passed. + +The local Codex CLI is 0.153.4. Its sandbox diagnostic uses a separate sandbox home; +missing credentials and `TERM=dumb` there are not evidence of broken desktop auth. +The GitHub connector successfully resolved this repository and its access permissions.