Skip to content

Fix matching and lexer performance - #153

Merged
chrisjenkinson merged 5 commits into
masterfrom
claude/matching-and-tokens
Sep 25, 2026
Merged

chrisjenkinson merged 5 commits into
masterfrom
claude/matching-and-tokens

Conversation

@chrisjenkinson

@chrisjenkinson chrisjenkinson commented Sep 25, 2026 •

Copy link
Copy Markdown
Owner

This fixes several problems in matching and speeds up lexing:

  • SimpleTextMatcher matches all remaining text, including newlines. Before, it stopped at the first \n, and the next position then failed with NoTokenFoundException.
  • RegexFinder throws RegexFailedException when preg_match() fails, instead of treating the failure as no match.
  • The all key contract is documented on MatcherInterface and enforced. A match whose all is missing, isn't a string, or isn't at the start of the remaining text throws InvalidMatchedTextException, naming the matcher. This catches matchers that forget the A anchor, which before could silently produce wrong tokens or a misleading ambiguity error.
  • Cursor tracks a byte offset and TokenStream a current index, so lexing and consuming no longer take quadratic time. Tokenising a 609 KB, 16,000-line document went from 11.2 s to 0.2 s.

@chrisjenkinson
chrisjenkinson force-pushed the claude/matching-and-tokens branch from 6f81d4d to fe86448 Compare September 25, 2026 09:30
@chrisjenkinson chrisjenkinson changed the title Fix matching, token positions and lexer performance Fix matching and lexer performance Sep 25, 2026
@chrisjenkinson
chrisjenkinson merged commit 62be614 into master Sep 25, 2026
4 checks passed
@chrisjenkinson
chrisjenkinson deleted the claude/matching-and-tokens branch September 25, 2026 09:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant