Skip to content

F02, the preprocessor is not a text editor - #79

Merged
tamnd merged 1 commit into
mainfrom
f02-tokens-not-text
Sep 6, 2026
Merged

tamnd merged 1 commit into
mainfrom
f02-tokens-not-text

Conversation

@tamnd

@tamnd tamnd commented Sep 6, 2026

Copy link
Copy Markdown
Owner

The second lesson of M3, and the first blueprint of the front end.

The lesson

One demonstration carries it. #define EMPTY, then a+EMPTY+b, and gcc -E -P prints a+ +b. Text substitution gives a++b. The space is in no input file. It is there because the empty expansion left a padding token and cpp_avoid_paste was asked whether + may be printed against +.

Thirty eight token pairs were probed by putting them next to each other with nothing between them. Twenty nine came back separated, nine did not, and the nine are exactly the ones that do not lex as a single token.

The rest of the lesson is the same point from four other directions:

  • 443 macros before a line of your program is read, against 408 on the other target. 373 names shared, and 50 of the shared names defined to different text, which is the set #ifdef cannot see. -std=c23 changes one name and no values, because the default was already gnu23.
  • 19 of the 20 built-ins are absent from -dM, which claims to print every macro. __FILE__ and __LINE__ are among the missing. The one that shows up is __STDC__, because cpp_init_builtins defines it a second time the ordinary way.
  • Four headers differing by one line. A comment after the #endif costs nothing. One declaration costs the multiple include optimization. -H names neither, because report_missing_guard only reports a file whose stack_count is 1, so the advice arrives for the header that has not cost anything yet.
  • One #include <stdio.h> opens 42 files on aarch64 Darwin and 31 on x86-64 Linux, and the two sets have no path in common.

What is new

  • gxray/cpp.py, a reader for recorded preprocessor output. It never expands a macro: expanding one needs the include path, the target's macros, the date and a hash table with the state of every other macro in it, and a lesson that faked those would teach a fiction.
  • corpora/cpp/f02.json, five macro tables, six expansions, thirty eight probes and four include traces, on two targets.
  • tests/test_cpp.py, including a check that compares the table of paste-avoiding pairs against the case labels of cpp_avoid_paste in the pinned tree. A GCC that grows a case fails the build rather than making a paragraph quietly false.
  • blueprints/BP-CPP.md, stub to partial.

BP-CPP

The 32-byte token and its 13 flags, the 90 token types and the four names in the table that are positions rather than types, the hash node, the macro, the 20 built-ins. Then cpp_get_token_1, argument prescan and why the two step stringify idiom exists, pasting as spelling-then-re-lexing and why ## on two macros that expand to + gives PLUSPLUS with no diagnostic, the two flags behind blue paint, the nine places that maintain the include guard state without ever scanning a file, and the printer's two independent reasons to print a space.

Section 9 has the boundary the whole lesson is about: two compilers that share no header file and disagree about 50 macro values agree exactly on what #define foo foo + 1 does.

Also

The glossary's token entry said twenty four bytes. libcpp's own comment at libcpp/include/cpplib.h:261 says thirty two on a 64-bit host.

Gates

pytest 1560 passed, nbbuild run 19 lessons, refcheck, prosecheck, bpc check, bpc pages --check, tier0 check/offline/online, dumpparse, ruff, glossary, all green. matrix digests --check still reports the two missing boot digests, which is the M2 gap tracked in #6 and unrelated to this branch.

The second lesson of M3. F01 showed the language the driver's decisions
are written in. This one goes inside the first program the driver runs.

The lesson is built on one demonstration: `#define EMPTY` and then
`a+EMPTY+b`, which prints `a+ +b`. Text substitution gives `a++b`. The
space is in no input file, and it is there because an empty expansion
leaves a padding token and cpp_avoid_paste was asked whether `+` may be
printed against `+`. Thirty eight token pairs were probed the same way.
Twenty nine came back separated and nine did not, and the nine are the
ones that do not lex as a single token.

gxray.cpp reads recorded preprocessor output: macro tables, line
markers, -H traces, expansion demonstrations. It never expands a macro,
because expanding one needs the include path, the target's macros, the
date and a hash table with the state of every other macro in it.

The corpus holds five macro tables, six expansions, the thirty eight
probes and four include traces, on aarch64 Darwin and x86-64 Linux. 443
macros against 408, 373 names shared, and 50 of the shared names defined
to different text, which is the set #ifdef cannot see. One #include
<stdio.h> opens 42 files on one machine and 31 on the other, and the two
sets have no path in common.

Four headers differing by one line measure the multiple include
optimization: a comment after the #endif costs nothing, one declaration
costs the whole thing. -H names neither of them, because
report_missing_guard only reports a file whose stack_count is 1, so the
advice arrives for the header that has not cost anything yet.

tests/test_cpp.py compares the table of paste-avoiding pairs against the
case labels of cpp_avoid_paste in the pinned tree, so a GCC that grows a
case fails the build rather than making a paragraph quietly false.

BP-CPP goes from stub to partial: the 32-byte token and its 13 flags,
the 90 token types and the four positions in the table, the hash node,
the macro, the 20 built-ins of which -dM prints one; cpp_get_token_1,
argument prescan, pasting by re-lexing, the two flags behind blue paint,
the nine places that maintain the include guard state, and the printer's
two independent reasons to print a space.

The glossary's `token` entry said twenty four bytes. libcpp's own
comment says thirty two on a 64-bit host.
@tamnd
tamnd merged commit 35e4803 into main Sep 6, 2026
9 checks passed
@tamnd
tamnd deleted the f02-tokens-not-text branch September 6, 2026 10:25
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant