Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 3 additions & 3 deletions .codex/AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,7 +20,7 @@ Important paths:

- `src/insight_extractor/` — package code;
- `tests/unit/` — model-free unit tests;
- `tests/integration/` — optional legacy tests; inspect mocks and API compatibility before running;
- `tests/integration/` — optional orchestration tests with injected fake ML boundaries;
- `.github/workflows/ci.yml` — required lint, type, unit, and smoke gates;
- `requirements.txt`, `constraints.txt`, `pyproject.toml` — dependency declarations and pins;
- `.codex/` — Codex instructions and workflow references;
Expand Down Expand Up @@ -48,8 +48,8 @@ python -m pytest tests/unit/ -v --tb=short
The CI smoke path uses fake BERT boundaries and production CLI orchestration.
Optional integration CI runs on workflow dispatch or a qualifying push head commit
containing `[run-integration]`; a PR commit message alone does not enable it.
The current integration files contain stale mocks/API calls and do not establish
real-model compatibility. Run with offline guards during diagnosis and report failures.
Those tests inject fake model/tokenizer objects and exercise the live public API
without downloading BERT weights.

Historical manual corrections: the live constraints already include the accelerate compatibility fix; `load_state()` invalidates embeddings lazily rather than recomputing them. Check current source and manifests before following historical sequences in `CLAUDE.md`.

Expand Down
19 changes: 8 additions & 11 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -276,13 +276,14 @@ jobs:
if-no-files-found: error

# ---------------------------------------------------------------------------
# 5. Integration tests (optional — heavy; skipped unless manually triggered
# or on a schedule, to avoid downloading BERT weights on every push)
# 5. Integration tests (optional orchestration suite; skipped unless manually
# triggered or [run-integration] appears in the push head commit message.
# Uses injected fake model/tokenizer boundaries — no BERT download.)
# ---------------------------------------------------------------------------
integration-tests:
name: Integration Tests (full BERT pipeline)
name: Integration Tests (orchestration, offline ML doubles)
runs-on: ubuntu-latest
# Run on manual dispatch, schedule, or when [run-integration] in commit msg
# Run on manual dispatch or when [run-integration] in push head commit msg
if: |
github.event_name == 'workflow_dispatch' ||
contains(github.event.head_commit.message, '[run-integration]')
Expand All @@ -299,14 +300,10 @@ jobs:
- name: Install package + dev deps
run: pip install -e ".[dev]"

- name: Cache HuggingFace model weights
uses: actions/cache@v4
with:
path: ~/.cache/huggingface
key: hf-models-all-MiniLM-L6-v2-${{ runner.os }}
restore-keys: hf-models-

- name: Run integration tests
env:
HF_HUB_OFFLINE: "1"
TRANSFORMERS_OFFLINE: "1"
run: |
pytest tests/integration/ \
-v \
Expand Down
13 changes: 7 additions & 6 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -80,10 +80,10 @@ Notes:
- `mypy` runs in `strict = true` mode. There is no lenient fallback.
- Line length is **100** (`[tool.ruff] line-length = 100`), not 88 or 120.
- Enabled ruff rule families: E, F, I, N, W, UP, B, C4, SIM.
- Integration tests (`tests/integration/`) download the real BERT model. CI runs them
- Integration tests (`tests/integration/`) exercise full extractor orchestration with
injected fake model/tokenizer boundaries (no weight download). CI runs them
**only** on `workflow_dispatch` or when the commit message contains
`[run-integration]`. Do not run them casually; do not move model-dependent tests
into `tests/unit/`.
`[run-integration]`. Do not move network/model-dependent checks into `tests/unit/`.

### Running the CLI

Expand Down Expand Up @@ -230,9 +230,10 @@ extractor._model = FakeModel() # .encode() -> deterministic np.ndarray
extractor._tokenizer = FakeTokenizer() # .tokenize_sentences() -> list[str]
```

Anything that genuinely needs the real model goes in `tests/integration/` and runs via
`[run-integration]`. Constructing `InsightExtractor` or `SentenceTokenizer` is safe in
unit tests (loading is lazy); *touching* `.model`/`.tokenizer` properties is not.
Anything that genuinely needs the real model should stay out of required CI and use an
explicit live-weight harness if added later. Constructing `InsightExtractor` or
`SentenceTokenizer` is safe in unit/integration tests (loading is lazy); *touching*
`.model`/`.tokenizer` properties without injection is not.

### 4.7 `\b` word-boundary regex traps (commit c3fef9b)

Expand Down
10 changes: 5 additions & 5 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -397,11 +397,11 @@ Required CI (`ci-pass` in `.github/workflows/ci.yml`) on every push/PR:

Optional integration CI (`pytest tests/integration/`) runs only on `workflow_dispatch`
or when a **push head-commit** message contains `[run-integration]`. A pull-request
commit message alone does **not** enable that job. Those files currently patch
`insight_extractor.tokenizer.AutoTokenizer` (imported only under `TYPE_CHECKING`) and
use stale constructor flags; they do **not** establish real-model compatibility and
are not expected to pass on `main`. Required CI going green does not certify this
suite.
commit message alone does **not** enable that job. The suite exercises full
extractor orchestration against the live public API with injected fake
model/tokenizer boundaries (no HuggingFace download). It does **not** certify
real BERT weight compatibility; required CI going green still does not imply a
live-model run.

```bash
# Required local gates (same as CI, minus the smoke job)
Expand Down
13 changes: 6 additions & 7 deletions docs/SPEC.md
Original file line number Diff line number Diff line change
Expand Up @@ -459,10 +459,9 @@ boundaries. `tests/unit/` must not download models.

Optional integration CI (`pytest tests/integration/`) runs only on `workflow_dispatch`
or a qualifying **push** commit message containing `[run-integration]`. A PR commit
message alone does not enable it. The current integration files use stale mocks
(including patching `AutoTokenizer` at a `TYPE_CHECKING`-only import) and do not
certify the real-model path. Required CI passing does not mean this suite passes on
`main`.
message alone does not enable it. The suite covers full extractor orchestration
against the live public API with injected fake model/tokenizer boundaries and does
not download BERT weights. Required CI passing still does not imply a live-model run.

### conftest.py
- Fixtures for `InsightExtractor`, `DynamicKeywordStemmer`, `SentenceTokenizer`
Expand Down Expand Up @@ -496,6 +495,6 @@ certify the real-model path. Required CI passing does not mean this suite passes
- Test markdown output generation

### test_extractor.py / test_e2e.py (integration — optional, not a required gate)
- Intended: full pipeline / real-model checks
- Current files: stale mocks and constructor flags; do not treat a green required
CI run as evidence this suite passed
- Full-pipeline orchestration against live `ExtractResult` / constructor APIs
- Inject fake `_model` / `_tokenizer` (no network, no weight download)
- Does not certify real BERT inference; keep true weight checks separate if needed
36 changes: 36 additions & 0 deletions tests/integration/conftest.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,36 @@
"""Shared fixtures for optional integration tests (no model download)."""

from __future__ import annotations

from pathlib import Path

import pytest

from insight_extractor.extractor import InsightExtractor
from tests.integration.fakes import attach_fakes


@pytest.fixture
def integration_extractor(temp_dir: Path) -> InsightExtractor:
"""Extractor with seed keywords and offline ML boundaries."""
extractor = InsightExtractor(
seed_keywords=["ransomware", "CVE", "exploit", "malware", "BERT", "Conti"],
output_dir=temp_dir,
top_k=10,
similarity_threshold=0.0,
enable_dynamic_regex=True,
)
return attach_fakes(extractor)


@pytest.fixture
def integration_extractor_no_dynamic(temp_dir: Path) -> InsightExtractor:
"""Extractor with dynamic keyword regex disabled."""
extractor = InsightExtractor(
seed_keywords=["ransomware", "CVE"],
output_dir=temp_dir,
top_k=5,
similarity_threshold=0.0,
enable_dynamic_regex=False,
)
return attach_fakes(extractor)
38 changes: 38 additions & 0 deletions tests/integration/fakes.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,38 @@
"""Offline model/tokenizer doubles for optional integration tests."""

from __future__ import annotations

from typing import Any

import numpy as np

from insight_extractor.extractor import InsightExtractor


class FakeModel:
"""Deterministic encode stub sized for cosine-similarity orchestration."""

def encode(
self,
texts: str | list[str],
*_args: Any,
**_kwargs: Any,
) -> np.ndarray[Any, np.dtype[np.float64]]:
items = [texts] if isinstance(texts, str) else list(texts)
rows = [[float(index), 1.0, 0.5, 0.25] for index, _ in enumerate(items, start=1)]
return np.array(rows, dtype=np.float64)


class FakeTokenizer:
"""Sentence splitter that never loads HuggingFace weights."""

def tokenize_sentences(self, text: str, *, max_tokens: int = 512) -> list[str]:
del max_tokens
return [part.strip() for part in text.split(".") if len(part.strip()) > 10]


def attach_fakes(extractor: InsightExtractor) -> InsightExtractor:
"""Inject model/tokenizer doubles so lazy loaders never run."""
extractor._model = FakeModel()
extractor._tokenizer = FakeTokenizer()
return extractor
Loading
Loading