Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 9 additions & 0 deletions .agents/skills/insight-add-entity-pattern/SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,9 @@
---
name: insight-add-entity-pattern
description: Add an Insight_Extractor regex entity across PatternLabel, patterns, tests and documentation. Use only in cgfixit/Insight_Extractor.
---

Read the repository's root `AGENTS.md` and its delegated `.codex/AGENTS.md`.
Then follow [the maintained workflow](../../../.codex/skills/add-entity-pattern/SKILL.md).
Resolve command paths from the repository root and use its activated virtual environment.
Report actual verification results and explicit skips; preserve existing user authorization.
9 changes: 9 additions & 0 deletions .agents/skills/insight-optimize/SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,9 @@
---
name: insight-optimize
description: Measure and optimize an Insight_Extractor hot path while preserving output, spans and lazy loading. Use only in cgfixit/Insight_Extractor.
---

Read the repository's root `AGENTS.md` and its delegated `.codex/AGENTS.md`.
Then follow [the maintained workflow](../../../.codex/skills/optimize/SKILL.md).
Resolve command paths from the repository root and use its activated virtual environment.
Report actual verification results and explicit skips; preserve existing user authorization.
9 changes: 9 additions & 0 deletions .agents/skills/insight-preflight/SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,9 @@
---
name: insight-preflight
description: Validate Insight_Extractor lint, types, offline unit tests and staging before commit or push. Use only in cgfixit/Insight_Extractor.
---

Read the repository's root `AGENTS.md` and its delegated `.codex/AGENTS.md`.
Then follow [the maintained workflow](../../../.codex/skills/preflight/SKILL.md).
Resolve command paths from the repository root and use its activated virtual environment.
Report actual verification results and explicit skips; preserve existing user authorization.
9 changes: 9 additions & 0 deletions .agents/skills/insight-verify-no-model/SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,9 @@
---
name: insight-verify-no-model
description: Verify Insight_Extractor regex, keyword and fake-model orchestration without downloading models. Use only in cgfixit/Insight_Extractor.
---

Read the repository's root `AGENTS.md` and its delegated `.codex/AGENTS.md`.
Then follow [the maintained workflow](../../../.codex/skills/verify-no-model/SKILL.md).
Resolve command paths from the repository root and use its activated virtual environment.
Report actual verification results and explicit skips; preserve existing user authorization.
23 changes: 15 additions & 8 deletions .codex/AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,10 +20,11 @@ Important paths:

- `src/insight_extractor/` — package code;
- `tests/unit/` — model-free unit tests;
- `tests/integration/` — real-model tests;
- `tests/integration/` — optional legacy tests; inspect mocks and API compatibility before running;
- `.github/workflows/ci.yml` — required lint, type, unit, and smoke gates;
- `requirements.txt`, `constraints.txt`, `pyproject.toml` — dependency declarations and pins;
- `.codex/` — Codex instructions and project skills;
- `.codex/` — Codex instructions and workflow references;
- `.agents/skills/` — discoverable, repository-specific skill entrypoints;
- `.claude/skills/` — existing detailed implementation checklists.

## Setup and validation
Expand All @@ -32,7 +33,7 @@ Use a Python 3.12+ virtual environment when possible:

```powershell
python -m pip install -r requirements.txt -c constraints.txt
python -m pip install -e ".[dev]"
python -m pip install -e ".[dev]" -c constraints.txt
```

Required gates:
Expand All @@ -44,7 +45,13 @@ python -m mypy src/insight_extractor
python -m pytest tests/unit/ -v --tb=short
```

The CI smoke path must remain model-free. Real BERT downloads belong only to `tests/integration/`, manual workflow dispatch, or commits containing `[run-integration]`.
The CI smoke path uses fake BERT boundaries and production CLI orchestration.
Optional integration CI runs on workflow dispatch or a qualifying push head commit
containing `[run-integration]`; a PR commit message alone does not enable it.
The current integration files contain stale mocks/API calls and do not establish
real-model compatibility. Run with offline guards during diagnosis and report failures.

Historical manual corrections: the live constraints already include the accelerate compatibility fix; `load_state()` invalidates embeddings lazily rather than recomputing them. Check current source and manifests before following historical sequences in `CLAUDE.md`.

## Non-negotiable invariants

Expand All @@ -69,9 +76,9 @@ The CI smoke path must remain model-free. Real BERT downloads belong only to `te

Read `.codex/README.md` and `.codex/codex_custom_instructions.md`, then use the relevant skill:

- `preflight` — before commit/push;
- `verify-no-model` — validate pipeline changes without downloading BERT;
- `add-entity-pattern` — add a static regex entity end to end;
- `optimize` — measured, minimal optimization of a real hot path.
- `insight-preflight` — before commit/push;
- `insight-verify-no-model` — validate pipeline changes without downloading BERT;
- `insight-add-entity-pattern` — add a static regex entity end to end;
- `insight-optimize` — measured, minimal optimization of a real hot path.

For fuller implementation checklists, consult the matching file under `.claude/skills/`.
26 changes: 21 additions & 5 deletions .codex/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,20 +7,36 @@ Codex-specific onboarding for `Insight_Extractor`.
- `AGENTS.md` — detailed project guidance loaded through the root `AGENTS.md` entrypoint.
- `codex_custom_instructions.md` — review-first behavior, minimal diffs, and secret safety.
- `ponytail-plugin.json` — existing optional Ponytail metadata.
- `skills/` — Codex-native project workflows.
- `skills/` — maintained workflow references.
- `../.agents/skills/` — discoverable repository skill entrypoints with `insight-` names.

The root `AGENTS.md` is Codex's auto-discovered entrypoint and delegates here.
`CLAUDE.md` remains the detailed historical operating manual. Do not duplicate runtime
architecture or dependency pins here.

## Discovery

Codex discovers repository skills under `.agents/skills`; `.codex/skills` alone is
not the documented repository discovery location. The four thin entrypoints link
to the maintained workflows here, avoiding duplicate checklists and collisions with
personal skills such as `preflight` and `optimize`. See the
[official skill discovery documentation](https://developers.openai.com/codex/skills/).
Open a new task in this checkout if the new skills are missing from the current catalog.

`skills/Insight-Optimizer/SKILL.md` is an optional extended optimization reference
(`name: insight-optimizer`); `skills/fable-protocol/SKILL.md` is historical owner
guidance, not a default project entrypoint. Its cross-project and model claims are
not evidence about this codebase. `ponytail-plugin.json` is reference metadata,
not proof that a plugin or hooks are installed.

## Skill dispatch

| Skill | Use it for |
|---|---|
| `preflight` | Ruff, formatting, strict mypy, unit tests, smoke checks, and staging hygiene before a commit or push |
| `verify-no-model` | Regex/dynamic/full-fake pipeline verification without a HuggingFace download |
| `add-entity-pattern` | Adding a new static regex entity and synchronizing enum, pattern, tests, docs, and smoke coverage |
| `optimize` | A small, measured improvement to regex, stemmer, state, or dynamic-expansion hot paths |
| `insight-preflight` | Ruff, formatting, strict mypy, unit tests, smoke checks, and staging hygiene before a commit or push |
| `insight-verify-no-model` | Regex/dynamic/full-fake pipeline verification without a HuggingFace download |
| `insight-add-entity-pattern` | Adding a new static regex entity and synchronizing enum, pattern, tests, docs, and smoke coverage |
| `insight-optimize` | A small, measured improvement to regex, stemmer, state, or dynamic-expansion hot paths |

The skills reuse the existing `.claude/skills/` checklists where those are more detailed. The Codex versions are the entrypoints and add Windows-friendly commands plus explicit stop conditions.

Expand Down
7 changes: 4 additions & 3 deletions .codex/skills/Insight-Optimizer/SKILL.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
---
name: optimize
name: insight-optimizer
description: Make 1-3 measured, minimal optimization to Insight_Extractor's regex, dynamic-stemmer, state, or keyword-expansion paths without changing public behavior or lazy model loading.
---

Expand All @@ -8,7 +8,7 @@ description: Make 1-3 measured, minimal optimization to Insight_Extractor's rege
This is a measured optimization workflow, not a broad refactor.
Target repo: `CGFixIT/Insight_Extractor` (Python 3.12+, Pydantic v2, lazy SentenceTransformer).
Source of truth order: live source → `CLAUDE.md` → `docs/SPEC.md` → `README.md`.
Always start from `main`. Trace every caller of a shared helper before touching it.
Start a new change from current `origin/main` on a `codex/` branch; preserve an existing task branch. Trace every caller of a shared helper before touching it.

## Hard Constraints (never violate)

Expand All @@ -18,7 +18,7 @@ Always start from `main`. Trace every caller of a shared helper before touching
- Do **not** remove or weaken lazy model loading (`model` / `tokenizer` properties).
- Do **not** break exact match order, deduplication, span fidelity, or exception paths.
- Preserve the keyword-mutation sequence after any change:
`stemmer.set_keywords(...)` → `registry.regenerate_dynamic_patterns(...)` → `_recompute_keyword_embeddings()` → `_auto_categorize_keywords()`.
Follow the current mutation path in `update_thread_keywords`; `load_state` rebuilds regex runtime and marks embeddings dirty without loading a model. Do not make state loading eager.
- Line length ≤ 100. Prefer `pathlib.Path` + `encoding="utf-8"`. No f-strings in log calls.
- Use `# ponytail:` comments for deliberate ceilings.

Expand Down Expand Up @@ -94,3 +94,4 @@ Do not manufacture a cache, configuration knob, or “money-mode” behavior to

# ponytail: TfidfVectorizer rebuilt every update_thread_keywords call;
# rolling corpus only after multi-document expansion is measured
```
5 changes: 4 additions & 1 deletion .codex/skills/preflight/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,12 +13,15 @@ Run from the repository root in PowerShell:
python -m ruff check src/ tests/
python -m ruff format --check src/ tests/
python -m mypy src/insight_extractor
$env:HF_HUB_OFFLINE = "1"
$env:TRANSFORMERS_OFFLINE = "1"
$env:HF_HUB_DISABLE_TELEMETRY = "1"
python -m pytest tests/unit/ -v --tb=short
git diff --check
git status --short
```

For changes to `config.py`, `constants.py`, packaging, or import boundaries, also run the model-free smoke check from the detailed checklist. Do not run the CLI from the repository root because it writes `insights_extracted.md` and `insight_extractor_state.json`.
For changes to `config.py`, `constants.py`, packaging, or import boundaries, also run the lightweight import check from the detailed checklist. That check is not the full CI smoke: inspect the `smoke-test` job in `.github/workflows/ci.yml`, which runs production CLI orchestration with fake model/tokenizer boundaries and verifies the report and state. Do not run the CLI from the repository root because it writes `insights_extracted.md` and `insight_extractor_state.json`.

Before staging, reject generated artifacts:

Expand Down
6 changes: 4 additions & 2 deletions .codex/skills/verify-no-model/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,9 @@ Rules:

- Run from a disposable directory, never the repository root.
- Regex and dynamic keyword paths must not access `.model` or `.tokenizer`.
- For full-pipeline tests, inject `extractor._model` and `extractor._tokenizer` before calling `extract()` or `load_state()`.
- `load_state()` must stay model-free; when a saved model name changes it clears both cached model and tokenizer. For subsequent full-pipeline checks, inject fake `_model` and `_tokenizer` after loading state.
- Model-free means no model download, not no ML dependencies: importing `extractor.py` imports `sentence_transformers`.
- Set `HF_HUB_OFFLINE=1` and `TRANSFORMERS_OFFLINE=1` for offline verification.
- Do not use `seed_keywords=[]` to mean “no keywords”; the implementation treats falsy input as the default seed bank.
- Assert concrete values such as a known CVE, IP, or entity, not merely non-empty output.
- Confirm the Markdown report and JSON state round-trip in the disposable directory, then confirm the repository has no generated artifacts.
Expand All @@ -29,4 +31,4 @@ try {
}
```

Escalate to `tests/integration/` only when changing model loading, real tokenization, embedding behavior, dependency compatibility, or model-path error handling. Use `[run-integration]` or workflow dispatch for that gate, and report it explicitly if skipped.
Escalate to `tests/integration/` only when changing model loading, real tokenization, embedding behavior, dependency compatibility, or model-path error handling. Inspect the optional suite for stale mocks before relying on it. Workflow dispatch enables its CI job; `[run-integration]` only works on qualifying push head commits, not ordinary PR events. Report failures or skipped model verification explicitly.
8 changes: 7 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -338,7 +338,7 @@ Insight_Extractor/
├── requirements.txt # Runtime deps with transformers compatibility note
├── constraints.txt # Pinned known-good versions
├── README.md # This file
├── SPEC.md # Full technical specification
├── docs/SPEC.md # Technical design reference; source is behavioral authority
├── src/
│ └── insight_extractor/
│ ├── __init__.py # Package entry point with lazy imports
Expand Down Expand Up @@ -366,6 +366,12 @@ Insight_Extractor/

---

## Codex setup

Start at [AGENTS.md](AGENTS.md) and [the Codex workflow map](.codex/README.md).
Repository skills are discoverable under `.agents/skills/` with `insight-` names.
See [the setup review](docs/CODEX_SETUP_REVIEW.md) for verified scope and follow-ups.

## Development Setup

```cmd
Expand Down
100 changes: 100 additions & 0 deletions docs/CODEX_SETUP_REVIEW.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,100 @@
# Codex setup review

Reviewed 2026-09-06 against origin/main at
`8097406` (before this documentation-only setup change).

## Architecture and scope

The packaged application is `src/insight_extractor`, installed with hatchling and
exposed through `insight-extract` and `python -m insight_extractor`.
The review covered package modules, unit and integration tests, README, SPEC,
dependency manifests, CI/security workflows, and every `.codex` file.
`docs/insight_extractor.py` is a separate legacy script with working-directory
inputs; its Porter-like logic is not the packaged stemmer implementation.

- `__init__` lazily exposes classes; lightweight models/config/constants/utilities
remain usable without importing torch. `extractor.py` itself imports ML libraries.
- The stemmer generates regexes, compiles growing keyword banks incrementally,
resolves original keywords, and maintains match spans. Registry regeneration is
explicit; it is not an observer of arbitrary stemmer mutations.
- `extract()` expands keywords with TF-IDF, prepares embeddings, then combines
static/dynamic regex, semantic hits, sentence scores, and keyword statistics.
- Reports escape Markdown and HTML. State loading restores settings and invalidates
embeddings lazily; changing the saved model name also clears model/tokenizer objects.

## Setup corrections

- Added four `insight-` skill entrypoints under `.agents/skills`, linking to existing
`.codex/skills` workflows. This preserves maintained references and avoids collisions
with generic personal skills. Repository discovery is documented in the
[official Codex skills guide](https://developers.openai.com/codex/skills/).
- Gave the extended optimizer a distinct name and closed its unfinished code fence.
- Corrected offline verification around state restoration and clarified that the
lightweight regex check does not reproduce the production CLI smoke job.
- Documented optional integration trigger limitations and stale test fixtures.
- Kept constraints on the editable development install in Codex onboarding.

The root AGENTS entrypoint already delegates to `.codex/AGENTS.md`. No project
config file or additional plugin installation is needed for these workflows.
Existing Ponytail JSON is reference metadata, not an installed plugin declaration.
Open a new task if its skill catalog does not show the new entrypoints.

## Follow-up findings

1. **Tokenizer termination and coverage:** `SentenceTokenizer.chunk_text` advances
by `max_tokens - overlap` without validating either argument. A long input with
equal budget/overlap never advances; `tokenize_sentences` fixes overlap at 50,
so small caller budgets also need review. Existing tokenizer tests cover normal
chunking but do not cover invalid budget/overlap combinations.
Add a bounded fake-tokenizer regression, then define validation and test normal,
zero, negative, and oversized overlap without downloading a model.
2. **Integration fixtures have drifted:** both integration files patch
`insight_extractor.tokenizer.AutoTokenizer`, which is only imported under
`TYPE_CHECKING` at module scope. `test_e2e.py` additionally uses unsupported
constructor flags and obsolete result fields. Some extraction occurs after
patch contexts exit. Repair fixtures against current signatures, use correctly
sized deterministic vectors, and distinguish offline orchestration from any
deliberate real-model test. A passing required CI gate does not certify this suite.
3. **State validation deserves a separate fix:** syntactically valid JSON with a
non-object root or invalid enum values can escape `StateLoadError`; field-by-field
assignment can leave partially changed state. Validate before mutation and test
failure atomicity while retaining lazy loading and the legacy category mapping.
4. **Documentation debt:** README advertises Porter/lemmatization and character-error
fuzzy matching, while the packaged stemmer uses suffix expansion and substring
regex matching. It also references a missing LICENSE. Packaging URLs use a hyphen
instead of the actual repository underscore. Historical CLAUDE/SPEC details and
security-audit conclusions require rechecking against current source/advisories.
This review is not a fresh dependency vulnerability audit.

## Suggested next skills

- `insight-tokenizer-check`: reproduce chunk-budget/overlap failures with an injected
tokenizer, enforce termination and token preservation, and verify lazy imports.
- `insight-state-contract`: validate state round trips, malformed schema handling,
failure atomicity, legacy categories, relative-path behavior, and model invalidation.

These are proposals for separate changes, not newly implemented runtime behavior.

## Local validation

The editable package was installed in `.venv` with system packages disabled and
`-c constraints.txt`. The available interpreter is Windows Python 3.12.0rc3;
this is not evidence of a clean stable-Python 3.13 installation or real-model inference.

- `pip check`: no broken requirements.
- Ruff lint and format: passed.
- Strict mypy: passed, 10 source files.
- Offline unit tests: 118 passed.
- Optional E2E diagnostic (`pytest tests/integration/test_e2e.py -x -q --tb=short`):
fails at fixture setup with `AttributeError: module 'insight_extractor.tokenizer'
has no attribute 'AutoTokenizer'`; stopped after the first failure. No real-model
download was attempted. The optional suite remains a separate repair task.
- Production CI smoke Python body, extracted from `ci.yml` and run in an ignored
scratch directory: passed (console, escaped Markdown, keyword expansion, state).
- Four new skill entrypoints: metadata validation passed; every workflow link resolves.
- Lightweight root import resolves to this checkout and does not import torch.
- `git diff --check`: passed.

The local Codex CLI is 0.153.4. Its sandbox diagnostic uses a separate sandbox home;
missing credentials and `TERM=dumb` there are not evidence of broken desktop auth.
The GitHub connector successfully resolved this repository and its access permissions.
Loading