Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
54 changes: 53 additions & 1 deletion GLOSSARY.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,7 @@ This file is generated from `gxray/glossary.py`. Edit that and run `just build-g

## Index

[DejaGnu](#dejagnu) | [GENERIC](#generic) | [GIMPLE](#gimple) | [IRA](#ira) | [LRA](#lra) | [RTL](#rtl) | [RTX](#rtx) | [SSA](#ssa) | [SSA name](#ssa-name) | [TODO flags](#todo-flags) | [allocno](#allocno) | [alternative](#alternative) | [assembler directive](#assembler-directive) | [back end](#back-end) | [basic block](#basic-block) | [blue paint](#blue-paint) | [bootstrap](#bootstrap) | [bubbling](#bubbling) | [build config](#build-config) | [cc1](#cc1) | [checking build](#checking-build) | [collect2](#collect2) | [compiler table](#compiler-table) | [constraint](#constraint) | [control flow graph](#control-flow-graph) | [cross compiler](#cross-compiler) | [current function](#current-function) | [debug counter](#debug-counter) | [default definition](#default-definition) | [define_insn](#define_insn) | [definition](#definition) | [directive](#directive) | [dominance](#dominance) | [driver](#driver) | [dump file](#dump-file) | [edge](#edge) | [effective target](#effective-target) | [excess errors](#excess-errors) | [expand](#expand) | [final](#final) | [front end](#front-end) | [garbage collector](#garbage-collector) | [gate](#gate) | [generated file](#generated-file) | [gengtype](#gengtype) | [gimplification](#gimplification) | [hard register](#hard-register) | [immediate dominator](#immediate-dominator) | [include guard](#include-guard) | [inferior call](#inferior-call) | [insn](#insn) | [interference](#interference) | [line marker](#line-marker) | [live range](#live-range) | [loop](#loop) | [machine description](#machine-description) | [machine mode](#machine-mode) | [middle end](#middle-end) | [mode iterator](#mode-iterator) | [optimization level](#optimization-level) | [out of SSA](#out-of-ssa) | [out of tree build](#out-of-tree-build) | [output template](#output-template) | [param](#param) | [pass](#pass) | [pass manager](#pass-manager) | [pass positioning](#pass-positioning) | [phi node](#phi-node) | [plugin](#plugin) | [plugin ABI](#plugin-abi) | [plugin event](#plugin-event) | [poly_int](#poly_int) | [port](#port) | [preprocessor](#preprocessor) | [pretty printer](#pretty-printer) | [pseudo register](#pseudo-register) | [pseudo-event](#pseudo-event) | [register allocation](#register-allocation) | [register class](#register-class) | [register pressure](#register-pressure) | [section](#section) | [spec](#spec) | [spec function](#spec-function) | [specs file](#specs-file) | [spill](#spill) | [stage comparison](#stage-comparison) | [stamp file](#stamp-file) | [sum file](#sum-file) | [target hook](#target-hook) | [target triple](#target-triple) | [temporary](#temporary) | [three address form](#three-address-form) | [token](#token) | [token pasting](#token-pasting) | [torture options](#torture-options) | [translation unit](#translation-unit) | [tree](#tree) | [use](#use) | [wide_int](#wide_int)
[DejaGnu](#dejagnu) | [GENERIC](#generic) | [GIMPLE](#gimple) | [IRA](#ira) | [LRA](#lra) | [RTL](#rtl) | [RTX](#rtx) | [SSA](#ssa) | [SSA name](#ssa-name) | [TODO flags](#todo-flags) | [allocno](#allocno) | [alternative](#alternative) | [assembler directive](#assembler-directive) | [back end](#back-end) | [basic block](#basic-block) | [blue paint](#blue-paint) | [bootstrap](#bootstrap) | [bubbling](#bubbling) | [build config](#build-config) | [cc1](#cc1) | [checking build](#checking-build) | [collect2](#collect2) | [compiler table](#compiler-table) | [constraint](#constraint) | [control flow graph](#control-flow-graph) | [cross compiler](#cross-compiler) | [current function](#current-function) | [debug counter](#debug-counter) | [default definition](#default-definition) | [define_insn](#define_insn) | [definition](#definition) | [diagnostic](#diagnostic) | [directive](#directive) | [dominance](#dominance) | [driver](#driver) | [dump file](#dump-file) | [edge](#edge) | [effective target](#effective-target) | [error recovery](#error-recovery) | [excess errors](#excess-errors) | [expand](#expand) | [final](#final) | [fix-it hint](#fix-it-hint) | [front end](#front-end) | [garbage collector](#garbage-collector) | [gate](#gate) | [generated file](#generated-file) | [gengtype](#gengtype) | [gimplification](#gimplification) | [hard register](#hard-register) | [immediate dominator](#immediate-dominator) | [include guard](#include-guard) | [inferior call](#inferior-call) | [insn](#insn) | [interference](#interference) | [line marker](#line-marker) | [live range](#live-range) | [lookahead](#lookahead) | [loop](#loop) | [machine description](#machine-description) | [machine mode](#machine-mode) | [middle end](#middle-end) | [mode iterator](#mode-iterator) | [optimization level](#optimization-level) | [out of SSA](#out-of-ssa) | [out of tree build](#out-of-tree-build) | [output template](#output-template) | [param](#param) | [parser](#parser) | [pass](#pass) | [pass manager](#pass-manager) | [pass positioning](#pass-positioning) | [phi node](#phi-node) | [plugin](#plugin) | [plugin ABI](#plugin-abi) | [plugin event](#plugin-event) | [poly_int](#poly_int) | [port](#port) | [preprocessor](#preprocessor) | [pretty printer](#pretty-printer) | [pseudo register](#pseudo-register) | [pseudo-event](#pseudo-event) | [register allocation](#register-allocation) | [register class](#register-class) | [register pressure](#register-pressure) | [section](#section) | [spec](#spec) | [spec function](#spec-function) | [specs file](#specs-file) | [spill](#spill) | [stage comparison](#stage-comparison) | [stamp file](#stamp-file) | [sum file](#sum-file) | [target hook](#target-hook) | [target triple](#target-triple) | [temporary](#temporary) | [three address form](#three-address-form) | [token](#token) | [token pasting](#token-pasting) | [torture options](#torture-options) | [translation unit](#translation-unit) | [tree](#tree) | [typedef name](#typedef-name) | [use](#use) | [wide_int](#wide_int)

## Reading the source

Expand Down Expand Up @@ -282,6 +282,58 @@ While a macro is being expanded, its name is disabled; any occurrence of it in t

Also written `NO_EXPAND`, painted blue, self-reference. Taught in F02. See also [preprocessor](#preprocessor), [token](#token). In the source: [`libcpp/macro.cc:1590@releases/gcc-16.2.0`](https://github.com/gcc-mirror/gcc/blob/releases/gcc-16.2.0/libcpp/macro.cc#L1590).

## In the parser

The program that reads C, and the four token slots everything surprising about it comes out of. F03 is the lesson.

### parser

**The recursive descent code that turns a token stream into trees. It can see four tokens.**

GCC's C parser is written by hand rather than generated, and all of its memory of your program is a `c_parser` struct: four token slots, a few flags, and the symbol table it shares with the rest of the front end. There is no dump flag for it, because it has no output of its own to print. What it produces is GENERIC, and what you can watch it do is complain. Nearly everything surprising about a C error message follows from the size of that buffer and from the fact that the symbol table has to answer a question before the parse can continue.

Also written `c_parser`, recursive descent, `cc1`. Taught in F03. See also [lookahead](#lookahead), [typedef name](#typedef-name), [GENERIC](#generic). In the source: [`gcc/c/c-parser.cc:191@releases/gcc-16.2.0`](https://github.com/gcc-mirror/gcc/blob/releases/gcc-16.2.0/gcc/c/c-parser.cc#L191).

### lookahead

**How far ahead the parser can look before deciding what it is reading. In C, four tokens.**

`c_parser_peek_token` gives the next one, `c_parser_peek_2nd_token` the one after it, and `c_parser_peek_nth_token` reaches as far as the fourth. The buffer behind all three is `c_token tokens_buf[4]` and nothing widens it. Most of the peeking in the C parser is one token deep, and the deepest constant peek in the whole file is there to recognise a version control conflict marker, which is not a C construct at all. When a grammar needs to see further than four, the parser does not get more; it commits, and then recovers.

Also written peek, `tokens_buf`, LL(k). Taught in F03. See also [parser](#parser), [token](#token), [error recovery](#error-recovery). In the source: [`gcc/c/c-parser.cc:572@releases/gcc-16.2.0`](https://github.com/gcc-mirror/gcc/blob/releases/gcc-16.2.0/gcc/c/c-parser.cc#L572).

### typedef name

**An identifier that names a type, and the reason C cannot be parsed without a symbol table.**

`A * b;` declares `b` as a pointer if `A` is a typedef name and multiplies two variables if it is not, and no amount of looking at tokens will tell you which. The parser asks the symbol table instead, and the answer is written into the token as `CPP_KEYWORD` or `CPP_NAME` the first time that token is looked at. That moment can come too early: a token peeked while one scope was open and used after it closed is carrying a stale answer, which is what `c_parser_maybe_reclassify_token` exists to undo.

Also written lexer hack, `CPP_KEYWORD`, `c_parser_maybe_reclassify_token`. Taught in F03. See also [parser](#parser), [token](#token), [lookahead](#lookahead). In the source: [`gcc/c/c-parser.cc:2326@releases/gcc-16.2.0`](https://github.com/gcc-mirror/gcc/blob/releases/gcc-16.2.0/gcc/c/c-parser.cc#L2326).

### diagnostic

**One complaint, with a message, a severity, a place, and often a suggested repair.**

A diagnostic is not a line of text. It is a structure with a primary location, any number of secondary ones, and any number of fix-it hints, and the text on your terminal is one rendering of it. `-fdiagnostics-format=sarif-stderr` is another, and it is the one to reach for when you want to read the structure rather than the prose. For the parser the message is finished by `c_parse_error`, which chooses one of thirteen endings from the type of the token the parser is looking at, which is why one missing semicolon can produce eight different sentences depending on what comes after it.

Also written `-fdiagnostics-format`, SARIF, `c_parse_error`. Taught in F03. See also [fix-it hint](#fix-it-hint), [parser](#parser), [pretty printer](#pretty-printer). In the source: [`gcc/c-family/c-common.cc:7004@releases/gcc-16.2.0`](https://github.com/gcc-mirror/gcc/blob/releases/gcc-16.2.0/gcc/c-family/c-common.cc#L7004).

### fix-it hint

**A machine applicable edit hung off a diagnostic, saying insert or delete or replace this text here.**

GCC does not suggest a repair for every mistake. For a missing token it suggests one only for the seven token types `get_missing_token_insertion_kind` knows, and the hint is what decides where the caret goes: two of the seven are inserted before the token that upset the parser and five after the token before it, and in the second case the caret moves back to the end of that previous token. That is why an error about a semicolon points at the line above the one you were reading. `-fdiagnostics-parseable-fixits` prints the hints in a form an editor can apply.

Also written `fixit_hint`, `rich_location`, `-fdiagnostics-parseable-fixits`. Taught in F03. See also [diagnostic](#diagnostic), [parser](#parser). In the source: [`libcpp/include/rich-location.h:620@releases/gcc-16.2.0`](https://github.com/gcc-mirror/gcc/blob/releases/gcc-16.2.0/libcpp/include/rich-location.h#L620).

### error recovery

**What the parser does after an error so that it can carry on and find the next one.**

A parser that stopped at the first mistake would make you compile a file once per typo, so after complaining it throws tokens away until it reaches one it can start again from, usually a semicolon or a closing brace at the right nesting depth. This is why three missing semicolons can come out as two errors rather than three, and why an error near the end of a file is sometimes a consequence of one near the top rather than a mistake of its own. The habit to build is to fix the first error and compile again.

Also written resynchronise, `c_parser_skip_until_found`. Taught in F03. See also [parser](#parser), [diagnostic](#diagnostic), [lookahead](#lookahead). In the source: [`gcc/c/c-parser.cc:1353@releases/gcc-16.2.0`](https://github.com/gcc-mirror/gcc/blob/releases/gcc-16.2.0/gcc/c/c-parser.cc#L1353).

## The four shapes a function takes

The same function, written down four different ways on its way to assembly. T02, T03 and T07 are the lessons.
Expand Down
3 changes: 2 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -58,8 +58,9 @@ Lessons do not invent a fresh example each time either. There are three programs
| B05 | [Sixty lines of C++, and you are inside the compiler](https://github.com/tamnd/gcc-internals/blob/main/lessons/b05-the-plugin/b05.ipynb) | GCC's plugin mechanism from the outside in: the three things a plugin has to have, the three ways it is refused, what an event actually is, a GIMPLE pass of your own inserted after ssa with its own dump file, switching one of GCC's passes off from outside and watching the assembly move, and why the thing you are writing against is not an API | M2 | [![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/tamnd/gcc-internals/blob/main/lessons/b05-the-plugin/b05.ipynb) |
| F01 | [The driver is an interpreter](https://github.com/tamnd/gcc-internals/blob/main/lessons/f01-the-spec-language/f01.ipynb) | That `gcc -dumpspecs` prints a program, in a language with conditionals and function calls, which the driver interprets to decide what to run; how to read it; and four text files that change what your compiler does without patching or rebuilding it | M3 | [![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/tamnd/gcc-internals/blob/main/lessons/f01-the-spec-language/f01.ipynb) |
| F02 | [The preprocessor is not a text editor](https://github.com/tamnd/gcc-internals/blob/main/lessons/f02-tokens-not-text/f02.ipynb) | That the preprocessor lexes your file into tokens before it does anything else, and that its printed output is a rendering of those tokens rather than the tokens themselves; the space GCC inserts that is in no input file; the four hundred macros you did not write; and why one #include of stdio.h opens thirty eight files | M3 | [![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/tamnd/gcc-internals/blob/main/lessons/f02-tokens-not-text/f02.ipynb) |
| F03 | [The C parser can see four tokens](https://github.com/tamnd/gcc-internals/blob/main/lessons/f03-four-tokens/f03.ipynb) | That GCC's C parser is hand written recursive descent whose whole memory is four token slots and the symbol table; why one missing semicolon produces eight different messages; why the caret is on the line above the mistake; why three mistakes come out as two errors; and why `A * b;` needs a symbol table to read | M3 | [![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/tamnd/gcc-internals/blob/main/lessons/f03-four-tokens/f03.ipynb) |

19 of 96 written.
20 of 96 written.
<!-- nbbuild:end index -->

This table is generated from the lessons, by the same command that builds them, so it cannot list a lesson that does not exist or miss one that does. T05 is the pilot and it is deliberately in the middle of Part I rather than at the start, because everything the course promises has to be true of a hard lesson before it is worth writing the easy ones.
Expand Down
Loading
Loading