F03: the C parser can see four tokens - #80
Merged
Merged
Conversation
The third lesson of M3, and the first one about the parser rather than about what feeds it. The idea is that GCC's C parser is hand written recursive descent whose entire memory of your program is four token slots and the symbol table. Everything surprising about C error messages falls out of that: why one missing semicolon produces eight different sentences, why the caret is usually on the line above the mistake and sometimes is not, why three mistakes come out as two errors, and why `A * b;` cannot be parsed without asking what `A` is. The parser has no dump flag, so the readout is the diagnostic itself. `gxray/cparse.py` reads `-fdiagnostics-format=sarif-stderr` back into spans, fix-it hints and related locations, and transcribes the two tables that decide a message: `get_missing_token_insertion_kind`, which says where a hint goes, and the thirteen `c_parse_error` branches, which say how the sentence ends. `tests/test_cparse.py` parses both out of the pinned tree and compares them in both directions, so a GCC that grows a case fails the build rather than making a paragraph quietly false. corpora/diag/f03.json 15 programs, 22 diagnostics lessons/f03-four-tokens/ the notebook, 71 cells, and the boss fight gxray/cparse.py the SARIF reader and the two tables tests/test_cparse.py 56 tests blueprints/BP-CPARSE.md the specification, status partial GLOSSARY.md six new terms, 105 total tier0 f03-cparse, offline, 3 new cached responses Two honest limits are named rather than papered over. A message ending in one quoted character could have come from either the character constant branch or from `%qE` on a one letter identifier, and the text cannot tell you which; the notebook says so and the boss fight asks about it. And the fourth lookahead slot exists only for `c_parser_peek_conflict_marker`, so C itself is parsed in three.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Third lesson of M3 (#22), and the first one about the parser rather than about what feeds it.
The idea
GCC's C parser is hand written recursive descent whose entire memory of your program is four token slots and the symbol table. Everything surprising about C error messages falls out of that: why one missing semicolon produces eight different sentences, why the caret is usually on the line above the mistake and sometimes is not, why three mistakes come out as two errors, and why
A * b;cannot be parsed without asking whatAis.The readout
The parser has no dump flag, so the readout is the diagnostic itself.
gxray/cparse.pyreads-fdiagnostics-format=sarif-stderrback into spans, fix-it hints and related locations, and transcribes the two tables that decide a message:get_missing_token_insertion_kind, seven token types, which says where a hint goesc_parse_errorbranches, which say how the sentence endstests/test_cparse.pyparses both out ofreleases/gcc-16.2.0and compares them against the transcription in both directions, so a GCC that grows a case fails the build rather than making a paragraph quietly false.What is here
corpora/diag/f03.jsonlessons/f03-four-tokens/gxray/cparse.pytests/test_cparse.pyblueprints/BP-CPARSE.mdpartialGLOSSARY.mdf03-cparse, kindoffline, 3 new cached responsesTwo honest limits, named rather than papered over
A message ending in one quoted character could have come from the character constant branch or from
%qEon a one letter identifier, and the text cannot tell you which. The notebook says so,suffix_forreturns nothing for it, and question 2 of the boss fight asks about it.The fourth lookahead slot exists only for
c_parser_peek_conflict_marker, which needs to see seven identical characters at the start of a line. C itself is parsed in three.Gates
All green:
pytest -q(1620 passed),bpc check|pages --check|coverage,refcheck check(663 citations),prosecheck(50 files),nbbuild check|claims|verify|run|index(20 lessons ran),tier0 check|coverage|offline|online(33 experiments, 25 cache hits, 0 misses),dumpparse check,matrix table|digests --check,ruff check,ruff format --check.