docs: add test coverage analysis and Hebbian architecture review - #11
Merged
Merged
Conversation
Comprehensive analysis of test coverage (69% overall, 1044 tests) identifying priority gaps in cli.py (41%), brain.py (78%), server.py (77%), and embeddings.py (78%). Includes Hebbian/learning architecture review confirming BCM, LTD, STC, and spreading activation are correctly implemented, with actionable recommendations for session priming, encoded-with synapses, and reconsolidation. https://claude.ai/code/session_01DG8xJRDX8tCzoRbZBKicsg
Prevents pytest-cov artifacts from appearing as untracked files. https://claude.ai/code/session_01DG8xJRDX8tCzoRbZBKicsg
New test files: - test_coverage_m1_session_prime: proves session prime boost (0.05) is below min_activation (0.1), preventing primed atoms from propagating through spreading activation - test_coverage_brain: status(), update_task(), get_stale_memories() - test_coverage_server_tasks: create_task, update_task, list_tasks, stale_memories, stats MCP tools - test_coverage_embeddings: error paths (ConnectionError, ResponseError, batch mismatch), ensure_ollama_ready, search_similar no-vec - test_coverage_cli_hooks: session-start, prompt-submit, post-tool, pre-tool-use, pre-compact, format helpers, synapse weight coherence Updated analysis doc: corrected M3 finding (encoded-with synapses ARE created in brain.py:264-296), updated coverage numbers. https://claude.ai/code/session_01DG8xJRDX8tCzoRbZBKicsg
SQLite's DEFAULT CURRENT_TIMESTAMP has second-level granularity, so both atoms often get the same created_at. detect_supersedes requires strict ordering (atom.created_at > candidate.created_at), causing the test to fail nondeterministically. Replace the ineffective asyncio.sleep(0.01) with an explicit UPDATE that shifts the second atom's created_at forward by 1 second. https://claude.ai/code/session_01DG8xJRDX8tCzoRbZBKicsg
When hooks auto-capture tool failures as antipatterns, they previously relied solely on keyword heuristics in extract_antipattern_fields() to infer severity. Error messages like "permission denied" have no explicit severity keywords, leaving severity=None and making the antipattern less likely to influence the LLM's behavior on recall. New _infer_error_severity() maps error signatures to severity: - permission denied, access denied, rm -rf, DROP TABLE → "high" - Python tracebacks, command not found, module not found → "medium" - Generic errors → "medium" (safe default) Both _hook_post_tool and _hook_post_tool_failure now pass this severity to brain.remember() for antipattern atoms. The brain's remember() still calls extract_antipattern_fields() as a fallback when severity=None, so explicit values from _infer_error_severity take precedence. 16 new tests prove the pipeline end-to-end. https://claude.ai/code/session_01DG8xJRDX8tCzoRbZBKicsg
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Comprehensive analysis of test coverage (69% overall, 1044 tests) identifying
priority gaps in cli.py (41%), brain.py (78%), server.py (77%), and
embeddings.py (78%). Includes Hebbian/learning architecture review confirming
BCM, LTD, STC, and spreading activation are correctly implemented, with
actionable recommendations for session priming, encoded-with synapses, and
reconsolidation.
https://claude.ai/code/session_01DG8xJRDX8tCzoRbZBKicsg