Skip to content

F03: the C parser can see four tokens - #80

Merged
tamnd merged 1 commit into
mainfrom
f03-four-tokens
Sep 6, 2026
Merged

tamnd merged 1 commit into
mainfrom
f03-four-tokens

Conversation

@tamnd

@tamnd tamnd commented Sep 6, 2026

Copy link
Copy Markdown
Owner

Third lesson of M3 (#22), and the first one about the parser rather than about what feeds it.

The idea

GCC's C parser is hand written recursive descent whose entire memory of your program is four token slots and the symbol table. Everything surprising about C error messages falls out of that: why one missing semicolon produces eight different sentences, why the caret is usually on the line above the mistake and sometimes is not, why three mistakes come out as two errors, and why A * b; cannot be parsed without asking what A is.

The readout

The parser has no dump flag, so the readout is the diagnostic itself. gxray/cparse.py reads -fdiagnostics-format=sarif-stderr back into spans, fix-it hints and related locations, and transcribes the two tables that decide a message:

  • get_missing_token_insertion_kind, seven token types, which says where a hint goes
  • the thirteen c_parse_error branches, which say how the sentence ends

tests/test_cparse.py parses both out of releases/gcc-16.2.0 and compares them against the transcription in both directions, so a GCC that grows a case fails the build rather than making a paragraph quietly false.

What is here

corpora/diag/f03.json 15 programs, 22 diagnostics
lessons/f03-four-tokens/ the notebook (71 cells, 35 of them code) and the boss fight
gxray/cparse.py the SARIF reader and the two tables
tests/test_cparse.py 56 tests
blueprints/BP-CPARSE.md the specification, status partial
GLOSSARY.md six new terms, 105 total
tier0 f03-cparse, kind offline, 3 new cached responses

Two honest limits, named rather than papered over

A message ending in one quoted character could have come from the character constant branch or from %qE on a one letter identifier, and the text cannot tell you which. The notebook says so, suffix_for returns nothing for it, and question 2 of the boss fight asks about it.

The fourth lookahead slot exists only for c_parser_peek_conflict_marker, which needs to see seven identical characters at the start of a line. C itself is parsed in three.

Gates

All green: pytest -q (1620 passed), bpc check|pages --check|coverage, refcheck check (663 citations), prosecheck (50 files), nbbuild check|claims|verify|run|index (20 lessons ran), tier0 check|coverage|offline|online (33 experiments, 25 cache hits, 0 misses), dumpparse check, matrix table|digests --check, ruff check, ruff format --check.

The third lesson of M3, and the first one about the parser rather than
about what feeds it.

The idea is that GCC's C parser is hand written recursive descent whose
entire memory of your program is four token slots and the symbol table.
Everything surprising about C error messages falls out of that: why one
missing semicolon produces eight different sentences, why the caret is
usually on the line above the mistake and sometimes is not, why three
mistakes come out as two errors, and why `A * b;` cannot be parsed
without asking what `A` is.

The parser has no dump flag, so the readout is the diagnostic itself.
`gxray/cparse.py` reads `-fdiagnostics-format=sarif-stderr` back into
spans, fix-it hints and related locations, and transcribes the two
tables that decide a message: `get_missing_token_insertion_kind`, which
says where a hint goes, and the thirteen `c_parse_error` branches, which
say how the sentence ends. `tests/test_cparse.py` parses both out of the
pinned tree and compares them in both directions, so a GCC that grows a
case fails the build rather than making a paragraph quietly false.

  corpora/diag/f03.json     15 programs, 22 diagnostics
  lessons/f03-four-tokens/  the notebook, 71 cells, and the boss fight
  gxray/cparse.py           the SARIF reader and the two tables
  tests/test_cparse.py      56 tests
  blueprints/BP-CPARSE.md   the specification, status partial
  GLOSSARY.md               six new terms, 105 total
  tier0                     f03-cparse, offline, 3 new cached responses

Two honest limits are named rather than papered over. A message ending
in one quoted character could have come from either the character
constant branch or from `%qE` on a one letter identifier, and the text
cannot tell you which; the notebook says so and the boss fight asks
about it. And the fourth lookahead slot exists only for
`c_parser_peek_conflict_marker`, so C itself is parsed in three.
@tamnd
tamnd merged commit 883d0f1 into main Sep 6, 2026
9 checks passed
@tamnd
tamnd deleted the f03-four-tokens branch September 6, 2026 11:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant