Skip to content

feat(history): query messages by time/channel via native chronicle index - #99

Merged
antra-tess merged 5 commits into
mainfrom
feat/chronicle-history-index
Sep 17, 2026
Merged

antra-tess merged 5 commits into
mainfrom
feat/chronicle-history-index

Conversation

@antra-tess

Copy link
Copy Markdown
Contributor

Summary

Registers two native chronicle secondary field indexes at MessageStore construction (/timestamp numeric, /metadata/external/channelId string, requires anima-research/chronicle#17) and adds query methods on top:

  • queryByTime / queryByChannel / queryByTimeAndChannel — O(log n + k) ordinal lookups via the native index; content is fetched only for the matched page via point lookups, never a full-slot materialization. queryByTimeAndChannel intersects the full, uncapped ordinal sets from both native queries before paginating — a native single-filter page can't be correctly post-filtered by the other criterion without paginating against the wrong universe.
  • getChannelCounts — native distinct-value counts, O(index size), zero content decoding.
  • getChannelTokenStats — token totals by channel. No native token index exists (tokens aren't a field on stored messages, and stamping one would mean rewriting every historical message), so this reuses the existing calibrated token estimator via a small incremental per-ordinal cache.

Corresponding ContextManager thin wrappers added in the same style as the existing getMessageWindow/queryMessages.

Reliability: every native call is capability-detected (an older chronicle build without this index throws a clear, specific error rather than silently degrading to something misleading) and self-heals once on a null result (chronicle's "no such index — unregistered, wrong kind, or poisoned" signal, distinct from a genuine empty match array): one re-register + retry, then a distinct "unavailable" error if still null — so a transiently-poisoned index elsewhere in the store never silently reads back as "no messages".

Built to back agent-facing "search/stats/extract over full uncompressed history" tools (companion agent-framework PR: anima-research/agent-framework#151) that need to stay fast against multi-GB production stores (Mythos/Sol scale) without a full scan.

A note on the base

Local main here had diverged substantially from origin/main (98 commits apart in each direction — both sides independently touched context-manager.ts/message-store.ts as part of two different kv-unified integration attempts). This PR's single commit was rebased cleanly (pure 3-way merge, no manual conflict resolution needed — the change is 100% additive) directly onto current origin/main, then re-verified there: typecheck clean, full suite 776/776 passing against origin/main's actual code, not the diverged local branch.

Test plan

  • npx tsc --noEmit clean against origin/main
  • Full suite: 776 tests, 776 pass, 0 fail (includes a new test/message-store-history-index.test.ts: 23 tests covering time-range boundary inclusivity, channel exact-match, time+channel intersection correctness, channel counts, token-stat totals against a manual sum, channelId-less messages, and the null-return self-heal/fail-closed paths)
  • Verified end-to-end against a locally-built chronicle native module (re-signed for local execution) — 300k-message synthetic-scale smoke test: native channel counts 0ms, channel-page fetch 1ms, time+channel intersection over 4000 matches 64ms, full-history token stats 3.9s cold / 80ms warm

🤖 Generated with Claude Code

antra-tess and others added 2 commits September 15, 2026 12:58
Registers two native chronicle secondary field indexes at MessageStore
construction (/timestamp numeric, /metadata/external/channelId string)
and adds query methods built on top:

- queryByTime / queryByChannel / queryByTimeAndChannel — O(log n + k)
  ordinal lookups via the native index, content fetched only for the
  matched page via point lookups, never a full-slot materialization.
  queryByTimeAndChannel intersects the full uncapped ordinal sets from
  both native queries before paginating, since a native single-filter
  page can't be post-filtered by the other criterion without
  paginating against the wrong universe.
- getChannelCounts — native distinct-value counts, O(index size), zero
  content decoding.
- getChannelTokenStats — token totals by channel; no native token
  index exists (tokens aren't a field on stored messages and stamping
  one would mean rewriting all historical messages), so this reuses
  the existing calibrated token estimator via a small incremental
  per-ordinal cache.

Query/registration calls are capability-detected (older chronicle
builds without this native index throw a clear, specific error rather
than silently falling back to something misleading) and self-heal once
on a `null` result (chronicle's "no such index — unregistered, wrong
kind, or poisoned" signal, distinct from an empty match array): a
single re-register + retry, then a distinct "unavailable" error if
still null, so a transiently-poisoned index (e.g. from a cross-branch
write elsewhere in the store) doesn't silently read as "no messages".

Corresponding ContextManager thin wrappers added in the same style as
the existing getMessageWindow/queryMessages.

Requires the companion chronicle native field-index PR
(anima-research/chronicle#17).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@antra-tess

Copy link
Copy Markdown
Contributor Author

CI status note

"Changelog entry present": fixed (fragment added in 1ddb3d8).

"Build & Test" (15 test failures across all 4 matrix jobs): expected for now, not a defect — this PR's new tests correctly follow this repo's own established convention (see test/message-store-window.test.ts's getStateSlice fallback test): assume the installed chronicle has the capability under test, and separately simulate an old/absent-capability chronicle via a Proxy for the fallback path (my test file does the same for registerStateFieldIndex/queryStateIndexRange/etc). CI installs the currently-published @animalabs/chronicle (^0.3.0), which doesn't have the native field-index capability yet — that only exists on anima-research/chronicle#17, not merged/published. The 15 failing tests are exactly the ones exercising real native-index behavior; everything else (including the explicit capability-absent/self-heal tests) passes.

This will go green on its own once chronicle#17 merges and publishes and this PR's @animalabs/chronicle version range is bumped to admit it — no code change needed here. In the meantime I've verified locally against a real build of chronicle#17 (native module built from that branch, swapped into node_modules): 776/776 tests pass, tsc --noEmit clean.

🤖 Generated with Claude Code

Chronicle 0.4.0 published (anima-research/chronicle#17) — ships the
native secondary field-index capability this PR's queryByTime/
queryByChannel/etc. depend on. Unblocks CI: the 15 tests that were
failing against the old published 0.3.0 (which lacks the native
capability) now run for real instead of hitting the graceful-
degradation error path.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@antra-tess

Copy link
Copy Markdown
Contributor Author

Recommendation: hold merge until the two P2 cache-correctness findings below are fixed and regression-tested. The native query paths and their self-healing behavior work in the tested cases, but getChannelTokenStats can silently report data from a different branch or a stale token calibration. The P3 channel-type mismatch is lower priority.

Reviewed head: ac4861ea36801ffc75a645b13f4572c015de6780 (rechecked before posting), using the published Chronicle 0.4.0 package.

  1. [P2] Scope the token-stat cache to the current branch — src/message-store.ts:1255–1258.

    The native timestamp/channel indexes are branch-aware and self-heal after a switch, but tokenStatsCache is keyed only by ordinal and survives that switch. An ordinal on a divergent branch can refer to an entirely different message, so a successfully healed native query still consumes the previous branch's cached channel and token estimate.

    Reproduced without any edits or redactions: append a shared base message, create side, append a main-only message on main, and warm token statistics. Switch to side and append a side-only message at the same ordinal. getChannelCounts() correctly reports base and side-only; getChannelTokenStats() instead reports base and main-only, with the old message's token estimate. This is separate from the explicitly documented approximation across live edits: ordinary branching and appending are enough to return another branch's data. Invalidate or scope the cache by branch identity before consuming ordinal hits; add a divergent-branch regression.

  2. [P2] Invalidate calibrated token totals when calibration changes — src/message-store.ts:1266–1268.

    The cache stores the result of estimateTokens, which already includes the mutable tokenCalibration multiplier. setTokenCalibration changes that multiplier without invalidating these entries, so the same unchanged message is estimated differently by estimateTokens and getChannelTokenStats. Later uncached messages are priced at the new factor, mixing calibration generations in one aggregate. The autobiographical strategy updates calibration during normal operation.

    Reproduction with a deterministic text-length estimator: cache statistics for a ten-character message, call setTokenCalibration(2), then query again. estimateTokens(message) returns 20, while getChannelTokenStats().totalTokensEstimate stays 10. Invalidate the stat cache when the calibration changes, or retain calibration-independent estimates and apply the current calibration with the same rounding semantics as the existing estimator. Add a calibration-change regression.

  3. [P3] Apply the native index's string-only channel rule to token statistics — src/message-store.ts:1260–1272.

    The TypeScript assertion on metadata.external does not validate the runtime value. The later !== undefined condition admits null, numbers, and other non-string channel IDs into byChannel, while Chronicle's string index excludes them. With one channel a message and one channelId: null message, getChannelCounts() returns only a, but token statistics return an additional channelId: null bucket, violating the exported channelId: string result type. Normalize non-string IDs to undefined; those messages should still contribute to whole-range totals.

Validation:

  • Clean isolated checkout, installed dependencies via npm ci; Chronicle resolved to 0.4.0.
  • npm test (including TypeScript build): 776 passed, zero failures, zero skipped.
  • Three additional targeted probes reproduced the findings above against the real native module.
  • git diff --check passed; all GitHub build/test and changelog checks were green.
  • The review checkout is clean; probes are saved outside it.

— Reviewed with OpenAI Codex.

Addresses 3 review findings from Codex/GPT-5.6 Sol on PR #99
(anima-research/context-manager, review at ac4861e), all in
getChannelTokenStats' per-ordinal cache:

- [P2] tokenStatsCache was keyed only by ordinal, so it survived a
  branch switch untouched even though ordinals are branch-relative.
  A diverged branch reusing the same ordinal for a different message
  would silently serve the other branch's cached channel/token data.
  Fixed by tracking the cache's warmed branch (tokenStatsCacheBranch)
  and wiping the cache wholesale on any detected change, mirroring
  this file's existing lookupIndex/rebuildIndex branch-detection
  pattern rather than inventing a new one.

- [P2] Cached entries stored the CALIBRATED token estimate, so a
  setTokenCalibration() call (routine during autobiographical
  strategy operation) left already-cached entries frozen at their old
  calibration while newly-cached entries used the new one. Fixed by
  caching the RAW (calibration-independent) estimate instead, mirroring
  the existing estimateBlockTokensRaw/estimateBlockTokens split, and
  applying the current calibration multiplier at read time — no
  invalidation needed when calibration changes.

- [P3] The TS cast on metadata.external didn't validate the runtime
  value, so a non-string channelId (null, a number, ...) could leak
  into byChannel as its own bucket, disagreeing with the native
  String-kind field index (which chronicle 0.4.0 correctly excludes
  such values from). Fixed with a small extractChannelId helper that
  normalizes non-string values to undefined.

Adds 3 regression tests in test/message-store-history-index.test.ts
reproducing each reviewer repro exactly.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@antra-tess

Copy link
Copy Markdown
Contributor Author

All 3 findings fixed in f33c2a0.

1. [P2] Branch-scoping: added `tokenStatsCacheBranch`, checked against `store.currentBranch().name` at the top of `getChannelTokenStats` — mirrors this file's own existing `lookupIndex`/`rebuildIndex` branch-change-detection pattern rather than inventing a new one. Any detected branch change wholesale-clears `tokenStatsCache` before it's touched (a stale-branch entry isn't just outdated, it can name a completely different message at the same ordinal, so a partial/keyed invalidation wasn't enough — full clear on branch change is correct here). New regression test reproduces your exact repro: shared `base` message, diverge on `main`/`side` at the same ordinal, warm cache pre-switch, assert post-switch stats reflect the current branch's message.

2. [P2] Calibration invalidation: the cache now stores the RAW (calibration-independent) per-message estimate via the existing `estimateBlockTokensRaw` — mirroring this file's own raw/calibrated split for the block-level cache — instead of the calibrated one. `tokenCalibration` is applied at READ time on every `getChannelTokenStats` call, so a `setTokenCalibration()` call is reflected on the very next call with no invalidation step needed at all (simpler than clearing on calibration change). New regression: warm cache at calibration 1, call `setTokenCalibration(2)`, assert totals double on the next call.

3. [P3] Non-string channelId leakage: added a private `extractChannelId` helper that runtime-validates (`typeof === 'string'`) instead of just type-asserting, normalizing `null`/numbers/etc. to `undefined` — matching chronicle's native String-kind field index, which silently excludes non-string values. New test: one message with `channelId: 'a'`, one with `channelId: null` — asserts `getChannelCounts()` and `getChannelTokenStats().byChannel` now agree.

Verification: `npx tsc --noEmit` clean, `npm run build` clean, `npm test` 779/779 pass (776 prior + 3 new regressions), 0 failures.

🤖 Generated with Claude Code

@antra-tess

Copy link
Copy Markdown
Contributor Author

Recommendation: hold merge for the remaining P2 branch-identity issue below. The original three reproduction probes now pass. The normal branch-switch, calibration-refresh, and non-string-channel cases are fixed, but branch-name reuse still serves another branch's statistics. The new rounding discrepancy is lower priority.

Re-reviewed head: f33c2a07589b9e495958b598c5686fea9115b64d.

  1. [P2] Compare the immutable branch ID, not its reusable name — src/message-store.ts:1310–1313.

    tokenStatsCacheBranch records only currentBranch().name. Chronicle allows deleting a non-current branch and creating a different branch under that same name. If no statistics call observes the intermediate branch, the name comparison never notices that the cache belongs to a deleted branch.

    Reproduced entirely through public APIs, without edits or redactions:

    • Append a shared base message on main; create and switch to side (branch ID 2).
    • Append a message in channel deleted-branch and warm the statistics cache.
    • Switch to main, append a different message in channel replacement-branch, delete side, and recreate side from main (now branch ID 3).
    • Switch to the new side and query again.

    Native counts correctly report base and replacement-branch, but token statistics report base and deleted-branch, using the deleted message's token estimate. Both branches are named side, so the new invalidation guard does not run. Track currentBranch().id instead of .name and add this delete/recreate regression.

  2. [P3] Preserve the existing per-block calibration rounding — src/message-store.ts:1334–1344.

    The new raw cache sums all block estimates first, then rounds the calibrated message total. estimateTokens instead sums individually rounded calibrated block estimates. round(sum(raw) * factor) is not generally equal to sum(round(raw * factor)), so statistics no longer match the existing estimator for multi-block messages at fractional calibration.

    Reproduction with the deterministic text-length estimator: append two one-character text blocks, warm the cache, then set calibration to 0.6. estimateTokens(message) returns 2, but getChannelTokenStats().totalTokensEstimate returns 1. Cache raw estimates per top-level block and apply the same per-block rounding, or invalidate/recompute the calibrated estimate when the factor changes. The new single-block test at integer factor 2 cannot detect this difference.

Verification:

  • npm test, including the TypeScript build: 779 passed, zero failures, zero skipped.
  • All three original review probes pass against this head.
  • The two additional probes above fail with the reported outputs.
  • git diff --check passes; all GitHub checks are green.
  • Published Chronicle 0.4.0 dependency; clean isolated review checkout.

— Re-reviewed with OpenAI Codex.

…fix multi-block rounding order

Addresses 2 further review findings from Codex/GPT-5.6 Sol on PR #99
(anima-research/context-manager, re-review at f33c2a0):

- [P2] tokenStatsCacheBranch compared store.currentBranch().name, but
  chronicle branch names are reusable (delete a non-current branch,
  create a different branch under the same name) while branch ids are
  not. A delete+recreate-under-the-same-name sequence with no
  intervening stats call left the name-based guard blind to the branch
  swap, serving the deleted branch's stale cached data. Renamed the
  tracked field to tokenStatsCacheBranchId and compare
  store.currentBranch().id instead — mirroring the reasoning chronicle's
  own native field-index fix already uses for branch scoping.

- [P3] The cache summed all of a message's raw per-block estimates and
  applied Math.round(sum * calibration) once, but the live
  estimateTokens() path sums INDIVIDUALLY-rounded per-block calibrated
  estimates instead. round(sum(raw) * f) != sum(round(raw * f)) in
  general for a multi-block message at a fractional calibration factor,
  so cached stats could disagree with a live estimateTokens() call on
  the same message. Changed the cache to store rawBlockEstimates
  (per-block, not pre-summed) and replay the same round-then-sum order
  at read time.

Adds 2 more regression tests in test/message-store-history-index.test.ts
reproducing each reviewer repro exactly (branch delete+recreate under
the same name; two single-character blocks at calibration 0.6).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@antra-tess

Copy link
Copy Markdown
Contributor Author

Both fixed in 9f43ed4.

1. [P2] Branch id, not name: `tokenStatsCacheBranch` (a `.name` string) is now `tokenStatsCacheBranchId`, compared and stamped from `store.currentBranch().id` — the immutable, never-reused identifier, unlike name which a delete+recreate can legitimately reuse. New regression reproduces your exact repro: warm cache on a branch, delete it, recreate a DIFFERENT branch under the same name, assert stats reflect the new branch's data, not the deleted one's.

2. [P3] Rounding order: the cache now stores `rawBlockEstimates: number[]` per top-level content block (not pre-summed for the whole message) and replays the identical per-block round-then-sum `estimateTokens` already uses, at read time in `getChannelTokenStats`. New regression: two 1-char text blocks, cache warmed at calibration 1, recalibrated to 0.6, asserts the cached path exactly equals a live `estimateTokens()` call (2, matching round-then-sum, not the old sum-then-round-once result of 1).

Verification: `npx tsc --noEmit` clean, `npm run build` clean, `npm test` 781/781 pass, 0 failures.

🤖 Generated with Claude Code

@antra-tess

Copy link
Copy Markdown
Contributor Author

Recommendation: merge. The remaining branch-identity and rounding findings are resolved in 9f43ed455c12297813f62a305f250b21c97f20d6. This supersedes my previous hold-merge recommendation.

I reviewed the latest diff and replayed all five probes from the two review rounds. All pass:

  • Switching to a divergent branch returns that branch's channel/token data.
  • Deleting and recreating a branch under the same name invalidates the cache using the new immutable branch ID.
  • Calibration changes affect already-cached messages immediately.
  • Multi-block messages use the same per-block round-then-sum order as estimateTokens, including fractional calibration.
  • Non-string channel IDs are excluded from channel buckets while still contributing to whole-range totals.

No new actionable findings in the fixes. The explicitly documented stale-cache behavior after edits/removals within a branch remains a limitation; this review does not claim that those statistics are refreshed after such mutations.

Validation on the reviewed head:

  • npm test, including the TypeScript build: 781 passed, zero failures, zero skipped.
  • Original and follow-up review probes: five passed against published Chronicle 0.4.0.
  • git diff --check: passed.
  • All four GitHub Build & Test jobs and the changelog check: green.
  • Isolated review checkout: clean; head rechecked before posting.

— Re-reviewed with OpenAI Codex.

@antra-tess
antra-tess merged commit f34a3fd into main Sep 17, 2026
5 checks passed
antra-tess added a commit to anima-research/agent-framework that referenced this pull request Sep 17, 2026
…ronicle to ^0.4.0

context-manager 0.9.0 published (anima-research/context-manager#99) —
ships the queryMessagesByTime/queryMessagesByChannel/etc. this
module's HistoryModule depends on. Also bumps this package's own
direct chronicle dependency to ^0.4.0: leaving it at ^0.3.0 while
context-manager pulled in 0.4.0 produced two separate installed
chronicle copies (npm can't dedupe across an unsatisfied range), which
made TypeScript see two structurally-identical-but-distinct JsStore
types and fail to compile anywhere this package's own code touches a
JsStore. Bumping both to ^0.4.0 lets npm dedupe to one copy.

Unblocks CI: tsc --noEmit was failing outright against the old
published context-manager (missing exports); now clean, and the full
suite passes (677/681, 4 pre-existing unrelated skips) against the
real published dependency chain.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant