From 164ea79adf71efe2a61fe9807af0946f8f03e5b9 Mon Sep 17 00:00:00 2001 From: "copilot-swe-agent[bot]" <198982749+Copilot@users.noreply.github.com> Date: Sat, 12 Sep 2026 16:51:02 +0000 Subject: [PATCH 1/3] Initial plan From 67dae39a442e29dc249d3bb7dab38b3a4f9a6b38 Mon Sep 17 00:00:00 2001 From: "copilot-swe-agent[bot]" <198982749+Copilot@users.noreply.github.com> Date: Sat, 12 Sep 2026 16:58:33 +0000 Subject: [PATCH 2/3] fix: repair optional integration suite for live extractor API Co-authored-by: cgfixit <17553614+cgfixit@users.noreply.github.com> --- .codex/AGENTS.md | 6 +- README.md | 10 +- docs/SPEC.md | 13 +- tests/integration/conftest.py | 36 +++++ tests/integration/fakes.py | 38 +++++ tests/integration/test_e2e.py | 235 ++++++++-------------------- tests/integration/test_extractor.py | 202 +++++++----------------- 7 files changed, 211 insertions(+), 329 deletions(-) create mode 100644 tests/integration/conftest.py create mode 100644 tests/integration/fakes.py diff --git a/.codex/AGENTS.md b/.codex/AGENTS.md index 7f345fd..fecc488 100644 --- a/.codex/AGENTS.md +++ b/.codex/AGENTS.md @@ -20,7 +20,7 @@ Important paths: - `src/insight_extractor/` — package code; - `tests/unit/` — model-free unit tests; -- `tests/integration/` — optional legacy tests; inspect mocks and API compatibility before running; +- `tests/integration/` — optional orchestration tests with injected fake ML boundaries; - `.github/workflows/ci.yml` — required lint, type, unit, and smoke gates; - `requirements.txt`, `constraints.txt`, `pyproject.toml` — dependency declarations and pins; - `.codex/` — Codex instructions and workflow references; @@ -48,8 +48,8 @@ python -m pytest tests/unit/ -v --tb=short The CI smoke path uses fake BERT boundaries and production CLI orchestration. Optional integration CI runs on workflow dispatch or a qualifying push head commit containing `[run-integration]`; a PR commit message alone does not enable it. -The current integration files contain stale mocks/API calls and do not establish -real-model compatibility. Run with offline guards during diagnosis and report failures. +Those tests inject fake model/tokenizer objects and exercise the live public API +without downloading BERT weights. Historical manual corrections: the live constraints already include the accelerate compatibility fix; `load_state()` invalidates embeddings lazily rather than recomputing them. Check current source and manifests before following historical sequences in `CLAUDE.md`. diff --git a/README.md b/README.md index 241d8c2..19f9a60 100644 --- a/README.md +++ b/README.md @@ -397,11 +397,11 @@ Required CI (`ci-pass` in `.github/workflows/ci.yml`) on every push/PR: Optional integration CI (`pytest tests/integration/`) runs only on `workflow_dispatch` or when a **push head-commit** message contains `[run-integration]`. A pull-request -commit message alone does **not** enable that job. Those files currently patch -`insight_extractor.tokenizer.AutoTokenizer` (imported only under `TYPE_CHECKING`) and -use stale constructor flags; they do **not** establish real-model compatibility and -are not expected to pass on `main`. Required CI going green does not certify this -suite. +commit message alone does **not** enable that job. The suite exercises full +extractor orchestration against the live public API with injected fake +model/tokenizer boundaries (no HuggingFace download). It does **not** certify +real BERT weight compatibility; required CI going green still does not imply a +live-model run. ```bash # Required local gates (same as CI, minus the smoke job) diff --git a/docs/SPEC.md b/docs/SPEC.md index 6ad8d54..e976169 100644 --- a/docs/SPEC.md +++ b/docs/SPEC.md @@ -459,10 +459,9 @@ boundaries. `tests/unit/` must not download models. Optional integration CI (`pytest tests/integration/`) runs only on `workflow_dispatch` or a qualifying **push** commit message containing `[run-integration]`. A PR commit -message alone does not enable it. The current integration files use stale mocks -(including patching `AutoTokenizer` at a `TYPE_CHECKING`-only import) and do not -certify the real-model path. Required CI passing does not mean this suite passes on -`main`. +message alone does not enable it. The suite covers full extractor orchestration +against the live public API with injected fake model/tokenizer boundaries and does +not download BERT weights. Required CI passing still does not imply a live-model run. ### conftest.py - Fixtures for `InsightExtractor`, `DynamicKeywordStemmer`, `SentenceTokenizer` @@ -496,6 +495,6 @@ certify the real-model path. Required CI passing does not mean this suite passes - Test markdown output generation ### test_extractor.py / test_e2e.py (integration — optional, not a required gate) -- Intended: full pipeline / real-model checks -- Current files: stale mocks and constructor flags; do not treat a green required - CI run as evidence this suite passed +- Full-pipeline orchestration against live `ExtractResult` / constructor APIs +- Inject fake `_model` / `_tokenizer` (no network, no weight download) +- Does not certify real BERT inference; keep true weight checks separate if needed diff --git a/tests/integration/conftest.py b/tests/integration/conftest.py new file mode 100644 index 0000000..ed6fa37 --- /dev/null +++ b/tests/integration/conftest.py @@ -0,0 +1,36 @@ +"""Shared fixtures for optional integration tests (no model download).""" + +from __future__ import annotations + +from pathlib import Path + +import pytest + +from insight_extractor.extractor import InsightExtractor +from tests.integration.fakes import attach_fakes + + +@pytest.fixture +def integration_extractor(temp_dir: Path) -> InsightExtractor: + """Extractor with seed keywords and offline ML boundaries.""" + extractor = InsightExtractor( + seed_keywords=["ransomware", "CVE", "exploit", "malware", "BERT", "Conti"], + output_dir=temp_dir, + top_k=10, + similarity_threshold=0.0, + enable_dynamic_regex=True, + ) + return attach_fakes(extractor) + + +@pytest.fixture +def integration_extractor_no_dynamic(temp_dir: Path) -> InsightExtractor: + """Extractor with dynamic keyword regex disabled.""" + extractor = InsightExtractor( + seed_keywords=["ransomware", "CVE"], + output_dir=temp_dir, + top_k=5, + similarity_threshold=0.0, + enable_dynamic_regex=False, + ) + return attach_fakes(extractor) diff --git a/tests/integration/fakes.py b/tests/integration/fakes.py new file mode 100644 index 0000000..1fd9040 --- /dev/null +++ b/tests/integration/fakes.py @@ -0,0 +1,38 @@ +"""Offline model/tokenizer doubles for optional integration tests.""" + +from __future__ import annotations + +from typing import Any + +import numpy as np + +from insight_extractor.extractor import InsightExtractor + + +class FakeModel: + """Deterministic encode stub sized for cosine-similarity orchestration.""" + + def encode( + self, + texts: str | list[str], + *_args: Any, + **_kwargs: Any, + ) -> np.ndarray[Any, np.dtype[np.float64]]: + items = [texts] if isinstance(texts, str) else list(texts) + rows = [[float(index), 1.0, 0.5, 0.25] for index, _ in enumerate(items, start=1)] + return np.array(rows, dtype=np.float64) + + +class FakeTokenizer: + """Sentence splitter that never loads HuggingFace weights.""" + + def tokenize_sentences(self, text: str, *, max_tokens: int = 512) -> list[str]: + del max_tokens + return [part.strip() for part in text.split(".") if len(part.strip()) > 10] + + +def attach_fakes(extractor: InsightExtractor) -> InsightExtractor: + """Inject model/tokenizer doubles so lazy loaders never run.""" + extractor._model = FakeModel() + extractor._tokenizer = FakeTokenizer() + return extractor diff --git a/tests/integration/test_e2e.py b/tests/integration/test_e2e.py index 4e0c429..8d0037f 100644 --- a/tests/integration/test_e2e.py +++ b/tests/integration/test_e2e.py @@ -1,126 +1,59 @@ -"""End-to-end tests with mocked BERT and full pipeline.""" +"""End-to-end pipeline checks with offline model/tokenizer doubles.""" from __future__ import annotations from pathlib import Path -from unittest.mock import MagicMock, patch - -import numpy as np -import pytest from insight_extractor.extractor import InsightExtractor -from insight_extractor.models import ExtractResult - -# ── Fixtures ───────────────────────────────────────────────────────────────── - - -@pytest.fixture -def mock_bert() -> MagicMock: - """Return a mock SentenceTransformer producing 384-dim embeddings.""" - mock = MagicMock() - # Return deterministic vectors so similarity calculations are stable - rng = np.random.default_rng(seed=42) - mock.encode = MagicMock( - side_effect=lambda texts, **kw: ( - rng.random((len(texts), 384)).astype(np.float32) - if isinstance(texts, list) - else rng.random((1, 384)).astype(np.float32) - ) - ) - return mock - - -@pytest.fixture -def mock_tokenizer() -> MagicMock: - mock = MagicMock() - mock.encode = MagicMock( - side_effect=lambda text, **kw: list(range(max(1, len(text.split()) * 2))) - ) - mock.decode = MagicMock(side_effect=lambda tokens, **kw: " ".join(["word"] * len(tokens))) - return mock - - -@pytest.fixture -def e2e_extractor(mock_bert: MagicMock, mock_tokenizer: MagicMock) -> InsightExtractor: - """Fully configured InsightExtractor with all dependencies mocked.""" - with ( - patch( - "insight_extractor.extractor.SentenceTransformer", - return_value=mock_bert, - ), - patch( - "insight_extractor.tokenizer.AutoTokenizer.from_pretrained", - return_value=mock_tokenizer, - ), - ): - ext = InsightExtractor( - model_name="all-MiniLM-L6-v2", - config_path=None, - seed_keywords=[ - "ransomware", - "CVE", - "exploit", - "malware", - "phishing", - "BERT", - "Conti", - ], - top_k=10, - similarity_threshold=0.3, - enable_dynamic=True, - enable_semantic=True, - enable_regex=True, - ) - return ext - - -# ── End-to-end tests ───────────────────────────────────────────────────────── +from insight_extractor.models import ExtractResult, KeywordStats, SemanticHit, SentenceScore +from tests.integration.fakes import attach_fakes class TestFullPipeline: """Run the complete extraction pipeline end-to-end.""" - def test_full_pipeline(self, e2e_extractor: InsightExtractor, sample_text: str) -> None: + def test_full_pipeline(self, integration_extractor: InsightExtractor, sample_text: str) -> None: """extract() on sample_text produces a valid ExtractResult.""" - result = e2e_extractor.extract(sample_text) + result = integration_extractor.extract(sample_text) - # Top-level type assert isinstance(result, ExtractResult) - - # All expected fields are present - assert result.text_hash != "" - assert isinstance(result.text_hash, str) - assert isinstance(result.keywords, list) - assert isinstance(result.keyword_stats, list) - assert isinstance(result.regex_matches, list) - assert isinstance(result.semantic_matches, list) - assert isinstance(result.sentence_scores, list) + assert result.input_hash != "" + assert isinstance(result.input_hash, str) + assert result.word_count > 0 + assert isinstance(result.regex_entities, dict) + assert isinstance(result.dynamic_keyword_matches, dict) + assert isinstance(result.semantic_keywords, list) + assert isinstance(result.key_sentences, list) + assert isinstance(result.newly_expanded_keywords, list) + assert isinstance(result.total_tracked_keywords, int) + assert isinstance(result.keyword_stats, KeywordStats) assert isinstance(result.timestamp, str) + assert "T" in result.timestamp or result.timestamp.endswith("Z") - # Timestamp is a valid ISO string - assert "T" in result.timestamp or "Z" in result.timestamp - - # At least some keywords were identified - assert len(result.keywords) > 0 - - # Each keyword stat is populated - for stat in result.keyword_stats: - assert stat.keyword - assert stat.count >= 1 + assert "CVE_ID" in result.regex_entities + assert result.dynamic_keyword_matches + assert result.semantic_keywords + assert all(isinstance(hit, SemanticHit) for hit in result.semantic_keywords) + if result.key_sentences: + assert isinstance(result.key_sentences[0], SentenceScore) + assert result.key_sentences[0].sentence + assert 0.0 <= result.key_sentences[0].score <= 1.0 - # Sentence scores are populated when text has multiple sentences - if len(result.sentence_scores) > 0: - assert result.sentence_scores[0].sentence - assert 0.0 <= result.sentence_scores[0].score <= 1.0 + assert result.keyword_stats.total_keywords == result.total_tracked_keywords + assert result.total_tracked_keywords == len(integration_extractor.thread_keywords) def test_dynamic_keyword_matches_present( - self, e2e_extractor: InsightExtractor, sample_text: str + self, integration_extractor: InsightExtractor, sample_text: str ) -> None: - """Keywords present in the text appear in regex_matches.""" - result = e2e_extractor.extract(sample_text) - matched_keywords = {m.keyword for m in result.regex_matches} - # At least one seed keyword should have been matched in the text - assert matched_keywords, "Expected at least one regex match" + """Seed keywords present in the text appear under DYNAMIC_KEYWORD matches.""" + result = integration_extractor.extract(sample_text) + assert result.dynamic_keyword_matches, "Expected at least one dynamic keyword match" + assert "DYNAMIC_KEYWORD" in result.dynamic_keyword_matches + matched = " ".join(result.dynamic_keyword_matches["DYNAMIC_KEYWORD"]).lower() + assert any( + keyword in matched + for keyword in ("ransomware", "cve", "bert", "conti", "exploit", "malware") + ) class TestMarkdownOutput: @@ -128,28 +61,20 @@ class TestMarkdownOutput: def test_markdown_output( self, - e2e_extractor: InsightExtractor, - temp_dir: Path, + integration_extractor: InsightExtractor, sample_text: str, ) -> None: - result = e2e_extractor.extract(sample_text) - md_path = temp_dir / "e2e_report.md" - e2e_extractor.save_results_to_markdown(result, md_path) + result = integration_extractor.extract(sample_text) + md_path = integration_extractor.save_results_to_markdown(result, "e2e_report.md") assert md_path.exists() content = md_path.read_text(encoding="utf-8") - - # Should contain a top-level heading assert content.startswith("# ") - - # Should mention the text hash - assert result.text_hash in content - - # Should contain expected markdown sections - assert "##" in content - - # Should list at least one keyword - assert len(result.keywords) == 0 or any(kw in content for kw in result.keywords[:3]) + assert result.input_hash in content + assert "## Regex Entities" in content + assert "## Dynamic Keyword Matches" in content + assert "## Semantic Keywords" in content + assert "## Key Sentences" in content class TestStatePersistence: @@ -157,72 +82,46 @@ class TestStatePersistence: def test_state_persistence( self, - e2e_extractor: InsightExtractor, + integration_extractor: InsightExtractor, temp_dir: Path, sample_text: str, ) -> None: - # Run extraction to populate state - e2e_extractor.extract(sample_text) - pre_keywords = set(e2e_extractor.top_keywords(n=50)) + integration_extractor.extract(sample_text) + pre_keywords = set(integration_extractor.top_keywords(n=50)) assert pre_keywords, "Expected some keywords before saving" - # Save state state_path = temp_dir / "e2e_state.json" - e2e_extractor.save_state(state_path) + integration_extractor.save_state(state_path) assert state_path.exists() - # Load into a fresh extractor with the same mock setup - mock_bert = MagicMock() - rng = np.random.default_rng(seed=99) - mock_bert.encode = MagicMock( - side_effect=lambda texts, **kw: rng.random( - (len(texts) if isinstance(texts, list) else 1, 384) - ).astype(np.float32) - ) - mock_tok = MagicMock() - mock_tok.encode = MagicMock( - side_effect=lambda text, **kw: list(range(max(1, len(text.split()) * 2))) - ) - mock_tok.decode = MagicMock( - side_effect=lambda tokens, **kw: " ".join(["word"] * len(tokens)) + fresh = InsightExtractor( + seed_keywords=["ransomware", "CVE"], + output_dir=temp_dir, + top_k=10, + similarity_threshold=0.0, ) - - with ( - patch( - "insight_extractor.extractor.SentenceTransformer", - return_value=mock_bert, - ), - patch( - "insight_extractor.tokenizer.AutoTokenizer.from_pretrained", - return_value=mock_tok, - ), - ): - fresh = InsightExtractor( - model_name="all-MiniLM-L6-v2", - config_path=None, - seed_keywords=[], - top_k=10, - similarity_threshold=0.3, - ) - fresh.load_state(state_path) + attach_fakes(fresh) + assert fresh.load_state(state_path) is True post_keywords = set(fresh.top_keywords(n=50)) assert post_keywords == pre_keywords def test_multiple_extractions_accumulate( - self, e2e_extractor: InsightExtractor, sample_text: str, short_text: str + self, + integration_extractor: InsightExtractor, + sample_text: str, + short_text: str, ) -> None: - """Running extract multiple times accumulates keyword frequency.""" - r1 = e2e_extractor.extract(sample_text) - r2 = e2e_extractor.extract(short_text) + """Running extract multiple times keeps a queryable keyword bank.""" + r1 = integration_extractor.extract(sample_text) + r2 = integration_extractor.extract(short_text) assert isinstance(r1, ExtractResult) assert isinstance(r2, ExtractResult) + assert r1.word_count > 0 + assert r2.word_count > 0 + assert r2.total_tracked_keywords >= r1.total_tracked_keywords - # Both runs should return keywords - assert len(r1.keywords) >= 0 - assert len(r2.keywords) >= 0 - - # Top keywords should still be queryable - top = e2e_extractor.top_keywords(n=5) + top = integration_extractor.top_keywords(n=5) assert isinstance(top, list) + assert top diff --git a/tests/integration/test_extractor.py b/tests/integration/test_extractor.py index 79e36f5..3ed6b42 100644 --- a/tests/integration/test_extractor.py +++ b/tests/integration/test_extractor.py @@ -1,128 +1,54 @@ -"""Integration tests for InsightExtractor with mocked ML dependencies.""" +"""Integration coverage for InsightExtractor orchestration (offline ML doubles).""" from __future__ import annotations from pathlib import Path -from unittest.mock import MagicMock, patch - -import numpy as np -import pytest from insight_extractor.extractor import InsightExtractor from insight_extractor.models import ExtractResult, KeywordStats - -# ── Fixtures ───────────────────────────────────────────────────────────────── - - -@pytest.fixture -def mock_model() -> MagicMock: - """Return a mock SentenceTransformer-like model.""" - mock = MagicMock() - mock.encode = MagicMock(return_value=np.random.rand(10, 384).astype(np.float32)) - return mock - - -@pytest.fixture -def mock_tokenizer() -> MagicMock: - """Return a mock HuggingFace tokenizer.""" - mock = MagicMock() - mock.encode = MagicMock( - side_effect=lambda text, **kw: list(range(max(1, len(text.split()) * 2))) - ) - mock.decode = MagicMock(side_effect=lambda tokens, **kw: " ".join(["word"] * len(tokens))) - return mock - - -@pytest.fixture -def extractor(mock_model: MagicMock, mock_tokenizer: MagicMock) -> InsightExtractor: - """Return an InsightExtractor with mocked heavy dependencies.""" - with ( - patch( - "insight_extractor.extractor.SentenceTransformer", - return_value=mock_model, - ), - patch( - "insight_extractor.tokenizer.AutoTokenizer.from_pretrained", - return_value=mock_tokenizer, - ), - ): - ext = InsightExtractor( - model_name="all-MiniLM-L6-v2", - config_path=None, - seed_keywords=["ransomware", "CVE"], - top_k=5, - similarity_threshold=0.5, - enable_dynamic_regex=True, - ) - return ext - - -@pytest.fixture -def extractor_no_dynamic(mock_model: MagicMock, mock_tokenizer: MagicMock) -> InsightExtractor: - """Return an extractor with dynamic regex disabled.""" - with ( - patch( - "insight_extractor.extractor.SentenceTransformer", - return_value=mock_model, - ), - patch( - "insight_extractor.tokenizer.AutoTokenizer.from_pretrained", - return_value=mock_tokenizer, - ), - ): - ext = InsightExtractor( - model_name="all-MiniLM-L6-v2", - config_path=None, - seed_keywords=["ransomware", "CVE"], - top_k=5, - similarity_threshold=0.5, - enable_dynamic_regex=False, - ) - return ext - - -# ── Tests ──────────────────────────────────────────────────────────────────── +from tests.integration.fakes import attach_fakes class TestExtract: - """Core extraction behaviour.""" + """Core extraction behaviour against the live ExtractResult contract.""" def test_extract_returns_extract_result( - self, extractor: InsightExtractor, sample_text: str + self, integration_extractor: InsightExtractor, sample_text: str ) -> None: - result = extractor.extract(sample_text) + result = integration_extractor.extract(sample_text) assert isinstance(result, ExtractResult) - def test_extract_has_timestamp(self, extractor: InsightExtractor, sample_text: str) -> None: - result = extractor.extract(sample_text) + def test_extract_has_timestamp( + self, integration_extractor: InsightExtractor, sample_text: str + ) -> None: + result = integration_extractor.extract(sample_text) assert isinstance(result.timestamp, str) assert result.timestamp != "" + assert "T" in result.timestamp or result.timestamp.endswith("Z") def test_extract_regex_only( - self, extractor_no_dynamic: InsightExtractor, sample_text: str + self, integration_extractor_no_dynamic: InsightExtractor, sample_text: str ) -> None: """With dynamic regex disabled, static regex entities should still be found.""" - result = extractor_no_dynamic.extract(sample_text) + result = integration_extractor_no_dynamic.extract(sample_text) assert isinstance(result, ExtractResult) - # keyword_stats is a KeywordStats object, not a list assert isinstance(result.keyword_stats, KeywordStats) assert result.dynamic_keyword_matches == {} + assert "CVE_ID" in result.regex_entities def test_keyword_expansion_after_extract( - self, extractor: InsightExtractor, sample_text: str + self, integration_extractor: InsightExtractor, sample_text: str ) -> None: """After extraction, top_keywords returns list of (keyword, count) tuples.""" - extractor.extract(sample_text) - top = extractor.top_keywords(n=20) + integration_extractor.extract(sample_text) + top = integration_extractor.top_keywords(n=20) assert isinstance(top, list) - # Each element is a (keyword, count) tuple for item in top: assert isinstance(item, tuple) assert len(item) == 2 kw, count = item assert isinstance(kw, str) assert isinstance(count, int) - # At least one seed keyword should appear kw_names = [item[0] for item in top] assert any(k in kw_names for k in ["ransomware", "CVE"]) @@ -132,42 +58,27 @@ class TestPersistence: def test_save_load_state( self, - extractor: InsightExtractor, + integration_extractor: InsightExtractor, temp_dir: Path, sample_text: str, ) -> None: - extractor.extract(sample_text) + integration_extractor.extract(sample_text) state_path = temp_dir / "state.json" - extractor.save_state(state_path) + integration_extractor.save_state(state_path) assert state_path.exists() - # Load into a fresh extractor - with ( - patch( - "insight_extractor.extractor.SentenceTransformer", - return_value=extractor._model, - ), - patch( - "insight_extractor.tokenizer.AutoTokenizer.from_pretrained", - return_value=MagicMock(), - ), - ): - fresh = InsightExtractor( - model_name="all-MiniLM-L6-v2", - config_path=None, - seed_keywords=[], - top_k=5, - similarity_threshold=0.5, - ) - fresh.load_state(state_path) + fresh = InsightExtractor( + seed_keywords=["ransomware", "CVE"], + output_dir=temp_dir, + top_k=5, + similarity_threshold=0.0, + ) + attach_fakes(fresh) + assert fresh.load_state(state_path) is True loaded_top = fresh.top_keywords(n=20) - original_top = extractor.top_keywords(n=20) - assert isinstance(loaded_top, list) - # Keyword names should be preserved across save/load - loaded_kws = {item[0] for item in loaded_top} - original_kws = {item[0] for item in original_top} - assert loaded_kws == original_kws + original_top = integration_extractor.top_keywords(n=20) + assert {item[0] for item in loaded_top} == {item[0] for item in original_top} class TestMarkdownOutput: @@ -175,30 +86,26 @@ class TestMarkdownOutput: def test_save_results_to_markdown( self, - extractor: InsightExtractor, - temp_dir: Path, + integration_extractor: InsightExtractor, sample_text: str, ) -> None: - extractor_with_dir = InsightExtractor( - seed_keywords=["ransomware", "CVE"], - output_dir=temp_dir, - ) - result = extractor_with_dir.extract(sample_text, update_keywords=False) - md_path = extractor_with_dir.save_results_to_markdown(result, "report.md") + result = integration_extractor.extract(sample_text, update_keywords=False) + md_path = integration_extractor.save_results_to_markdown(result, "report.md") assert md_path.exists() content = md_path.read_text(encoding="utf-8") assert "# Insight Extraction Results" in content - assert len(content) > 0 + assert "## Regex Entities" in content + assert result.input_hash in content class TestTopKeywords: """Frequency tracking.""" def test_top_keywords_returns_tuples( - self, extractor: InsightExtractor, sample_text: str + self, integration_extractor: InsightExtractor, sample_text: str ) -> None: - extractor.extract(sample_text) - top = extractor.top_keywords(n=3) + integration_extractor.extract(sample_text) + top = integration_extractor.top_keywords(n=3) assert len(top) <= 3 assert isinstance(top, list) for item in top: @@ -207,41 +114,44 @@ def test_top_keywords_returns_tuples( assert isinstance(kw, str) assert isinstance(count, int) - def test_top_keywords_empty(self, extractor: InsightExtractor) -> None: - """Before extraction, top_keywords returns seed keywords.""" - top = extractor.top_keywords(n=5) + def test_top_keywords_empty(self, integration_extractor: InsightExtractor) -> None: + """Before extraction, top_keywords still returns seed keyword frequencies.""" + top = integration_extractor.top_keywords(n=5) assert isinstance(top, list) + assert top class TestKeywordStats: """KeywordStats retrieval.""" - def test_get_keyword_stats(self, extractor: InsightExtractor) -> None: - stats = extractor.get_keyword_stats() + def test_get_keyword_stats(self, integration_extractor: InsightExtractor) -> None: + stats = integration_extractor.get_keyword_stats() assert isinstance(stats, KeywordStats) - assert isinstance(stats.total_keywords, int) assert stats.total_keywords >= 0 assert isinstance(stats.category_counts, dict) assert isinstance(stats.stem_mode, str) def test_keyword_stats_from_extract_result( - self, extractor: InsightExtractor, sample_text: str + self, integration_extractor: InsightExtractor, sample_text: str ) -> None: - result = extractor.extract(sample_text) + result = integration_extractor.extract(sample_text) assert isinstance(result.keyword_stats, KeywordStats) - assert result.keyword_stats.total_keywords == len(extractor.thread_keywords) + assert result.keyword_stats.total_keywords == len(integration_extractor.thread_keywords) class TestInitCustomSeeds: """Custom seed keywords are used at init time.""" - def test_init_custom_seeds(self) -> None: - ext = InsightExtractor( - model_name="all-MiniLM-L6-v2", - seed_keywords=["custom_seed_1", "custom_seed_2"], + def test_init_custom_seeds(self, temp_dir: Path) -> None: + ext = attach_fakes( + InsightExtractor( + seed_keywords=["custom_seed_1", "custom_seed_2"], + output_dir=temp_dir, + ) ) kw_names = [kw for kw, _ in ext.top_keywords(n=10)] - assert "custom_seed_1" in kw_names or "custom_seed_2" in kw_names + assert "custom_seed_1" in kw_names + assert "custom_seed_2" in kw_names class TestDisabledDynamicRegex: @@ -249,9 +159,9 @@ class TestDisabledDynamicRegex: def test_disabled_dynamic_regex( self, - extractor_no_dynamic: InsightExtractor, + integration_extractor_no_dynamic: InsightExtractor, sample_text: str, ) -> None: - result = extractor_no_dynamic.extract(sample_text) + result = integration_extractor_no_dynamic.extract(sample_text) assert isinstance(result, ExtractResult) assert result.dynamic_keyword_matches == {} From 71ac403f187c674a65844074f83c7e5a3e782a99 Mon Sep 17 00:00:00 2001 From: "copilot-swe-agent[bot]" <198982749+Copilot@users.noreply.github.com> Date: Sat, 12 Sep 2026 17:02:04 +0000 Subject: [PATCH 3/3] docs: align CI job name and manuals with offline integration suite Co-authored-by: cgfixit <17553614+cgfixit@users.noreply.github.com> --- .github/workflows/ci.yml | 19 ++++++++----------- CLAUDE.md | 13 +++++++------ 2 files changed, 15 insertions(+), 17 deletions(-) diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index 115132f..4b6722c 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -276,13 +276,14 @@ jobs: if-no-files-found: error # --------------------------------------------------------------------------- - # 5. Integration tests (optional — heavy; skipped unless manually triggered - # or on a schedule, to avoid downloading BERT weights on every push) + # 5. Integration tests (optional orchestration suite; skipped unless manually + # triggered or [run-integration] appears in the push head commit message. + # Uses injected fake model/tokenizer boundaries — no BERT download.) # --------------------------------------------------------------------------- integration-tests: - name: Integration Tests (full BERT pipeline) + name: Integration Tests (orchestration, offline ML doubles) runs-on: ubuntu-latest - # Run on manual dispatch, schedule, or when [run-integration] in commit msg + # Run on manual dispatch or when [run-integration] in push head commit msg if: | github.event_name == 'workflow_dispatch' || contains(github.event.head_commit.message, '[run-integration]') @@ -299,14 +300,10 @@ jobs: - name: Install package + dev deps run: pip install -e ".[dev]" - - name: Cache HuggingFace model weights - uses: actions/cache@v4 - with: - path: ~/.cache/huggingface - key: hf-models-all-MiniLM-L6-v2-${{ runner.os }} - restore-keys: hf-models- - - name: Run integration tests + env: + HF_HUB_OFFLINE: "1" + TRANSFORMERS_OFFLINE: "1" run: | pytest tests/integration/ \ -v \ diff --git a/CLAUDE.md b/CLAUDE.md index 4d097d6..1f2a267 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -80,10 +80,10 @@ Notes: - `mypy` runs in `strict = true` mode. There is no lenient fallback. - Line length is **100** (`[tool.ruff] line-length = 100`), not 88 or 120. - Enabled ruff rule families: E, F, I, N, W, UP, B, C4, SIM. -- Integration tests (`tests/integration/`) download the real BERT model. CI runs them +- Integration tests (`tests/integration/`) exercise full extractor orchestration with + injected fake model/tokenizer boundaries (no weight download). CI runs them **only** on `workflow_dispatch` or when the commit message contains - `[run-integration]`. Do not run them casually; do not move model-dependent tests - into `tests/unit/`. + `[run-integration]`. Do not move network/model-dependent checks into `tests/unit/`. ### Running the CLI @@ -230,9 +230,10 @@ extractor._model = FakeModel() # .encode() -> deterministic np.ndarray extractor._tokenizer = FakeTokenizer() # .tokenize_sentences() -> list[str] ``` -Anything that genuinely needs the real model goes in `tests/integration/` and runs via -`[run-integration]`. Constructing `InsightExtractor` or `SentenceTokenizer` is safe in -unit tests (loading is lazy); *touching* `.model`/`.tokenizer` properties is not. +Anything that genuinely needs the real model should stay out of required CI and use an +explicit live-weight harness if added later. Constructing `InsightExtractor` or +`SentenceTokenizer` is safe in unit/integration tests (loading is lazy); *touching* +`.model`/`.tokenizer` properties without injection is not. ### 4.7 `\b` word-boundary regex traps (commit c3fef9b)