Skip to content

The pages say what they are about, and the card's size is read off the card - #81

Merged
ChelseaKR merged 3 commits into
mainfrom
seo/the-pages-say-what-they-are-about
Sep 18, 2026
Merged

ChelseaKR merged 3 commits into
mainfrom
seo/the-pages-say-what-they-are-about

Conversation

@ChelseaKR

@ChelseaKR ChelseaKR commented Sep 13, 2026 •

Copy link
Copy Markdown
Owner

Every page of the documentation site stated what it was to a person and nothing
at all to a machine. Parsing the five live pages at
https://chelseakr.github.io/gauntlet/ for <script type="application/ld+json">
returns zero blocks on all five. After this change it returns one on each.

What the page now says

One @graph per page, four nodes: the site, the page, the share card, and the
software the page is about.

WebSite             -- the site, named as og:site_name names it
WebPage             -- this page: its <title>, its description, its canonical
ImageObject         -- the share card, at the size read off the card itself
SoftwareApplication -- what the page is about, from the packaging metadata

Nothing in it is typed. Every value is read back out of the same constant, tag
or file that renders the visible head, so the claim a crawler reads and the
claim a person reads cannot disagree:

node property read from
WebPage.name the page's <title>
WebPage.description the page's <meta name="description">
WebPage.url the existing canonical builder
WebSite.name the constant og:site_name already renders
ImageObject.url, .caption the constants og:image and og:image:alt already render
ImageObject.width, .height the committed PNG's own IHDR chunk
SoftwareApplication.alternateName, .description, .url, .codeRepository the installed distribution's metadata, which is pyproject.toml speaking through the install
every inLanguage the constant that renders <html lang>

Read from the installed metadata rather than from pyproject.toml for the
reason ASSETS is a package path: gauntlet site has to work from a wheel,
where there is no repository to read a source file out of. The test reads the
source table instead, so the two agree only when the environment was synced
from this tree and a stale install is a failing comparison rather than an older
sentence on a published page.

What it deliberately does not say

No Dataset, no distribution, no DCAT. A dataset descriptor is not a
description, it is an invitation: it exists so that dataset search engines and
state open-data catalogs harvest what it names and list it as a dataset of
record, and a catalog listing is far easier to acquire than to withdraw.
Whether an evaluation pack, a JSON pack or a reviewer document should solicit
that is an open question with an owner's name on it. test_it_solicits_no_dataset_harvest
forbids the vocabulary permanently so the difference stays a decision somebody
makes rather than a line somebody adds, and a companion test runs each
forbidden term through the same scanner to prove the scanner bites.

No softwareVersion. These pages are built from main, which carries the
version being prepared, while the index carries the last one released -- the two
have differed for most of this project's life, including right now. The field
would announce a release that does not exist yet, to consumers that read
structured data and to nobody in a position to see it was wrong.

The check

tests/test_site.py gains a section that parses all five built pages with a
parser matching on the script element and its type attribute, never on the
string. That distinction is not fastidiousness: an audit of this portfolio
scored a sibling project as carrying structured data because the string
application/ld+json occurred on its page, and the one occurrence was the
accept attribute of a file picker.

Coverage is 5 pages examined / 5 examinable, and
test_there_are_pages_to_examine asserts the examinable set equals what the
build wrote and is not empty, so the section cannot pass having read nothing.
It runs inside make verify, which is what CI runs.

Each assertion class was checked by sabotaging the generator, rebuilding, and
confirming the gate goes red: a missing block, a page node that stopped reading
the <title>, an @id reference to a node the graph does not define, card
dimensions that stopped being read off the PNG, a software node that stopped
reading the packaging metadata, a Dataset node with a distribution, an
executable <script>, an empty property, and a dropped node. All nine reddened;
each was reverted and the tree hash checked back to the byte.

Two things fixed on the way

The share card's size is read off the card. og:image:width and
og:image:height were the literals 1200 and 630. They happened to be right, and
nothing would have said so if the card were re-rendered at another size: every
page would have gone on announcing the old one with every check green. A build
now refuses a card it cannot read rather than stating a size it guessed, and
refuses a missing card before it renders anything rather than after.

The no-script check now counts what it was named for. It asserted
scripts == 0, which the data block above would end. A script element whose
type is not a script type is a data block: the HTML spec never prepares it, it
never executes, and script-src has no say over it, so "static pages, no
runtime, nothing for a CSP to have to allow" is unchanged. What replaces the old
count is stricter rather than looser -- it goes red for an inline script, a
src, a type="module" and a type="text/javascript" alike, none of which the
old assertion distinguished because none of them could occur. The three places
that stated "no script at all" in prose (.htmlvalidate.mjs, tools/a11y.mjs,
package.json) now say "no executable script" and name the data block, because
a promise quietly narrowed is worse than a promise restated. This is the one
judgement call in the change and it is the thing to disagree with if any part of
it is wrong.

make verify is green (1253 passed, 96.20% coverage), and npx html-validate
and node tools/a11y.mjs are clean over the five built pages.

Prepared with AI assistance; reviewed before submission.

…e card

Every page of the documentation site stated what it was to a person and
nothing at all to a machine: parsing the five live pages for
`<script type="application/ld+json">` returned zero blocks on all five.

Each page now carries one `@graph` of four nodes -- the site, the page, the
share card, and the software the page is about. Nothing in it is typed. The
page node's name is the `<title>`, its description is the
`<meta name="description">`, its url is the canonical; the site node's name
is what `og:site_name` already renders; the card node's dimensions are read
out of the committed PNG's own IHDR chunk; and the software node's name,
sentence and links are the installed distribution's metadata, which is
`pyproject.toml` speaking through the install rather than a paraphrase of it.
Read from the installed metadata rather than the file because `gauntlet site`
has to work from a wheel; the test reads the source table instead, so a stale
environment is a failing comparison rather than an older sentence published.

There is no `Dataset` node, no `distribution`, and no DCAT vocabulary, and
a test forbids all of it permanently. A dataset descriptor is not a
description, it is an invitation: it exists so dataset search engines and
state open-data catalogs harvest what it names, and a catalog listing is far
easier to acquire than to withdraw. Whether an evaluation pack should solicit
that is an open question with an owner's name on it, and the test is there so
it stays a decision somebody makes rather than a line somebody adds. There is
no `softwareVersion` either: these pages are built from `main`, which carries
the version being prepared and not the one on the index, so the field would
announce a release that does not exist yet.

The check parses every built page with a parser that matches on the script
element and its `type` attribute, never on the string, and holds every value
against the tag, the file or the `pyproject.toml` table it came from, so a
node that stopped being derived fails while it still says something
plausible. It names the examinable set so it cannot pass having read nothing,
and it runs inside `make verify`.

Two things fixed on the way:

`og:image:width` and `og:image:height` were the literals 1200 and 630. They
happened to be right, and nothing would have said so if the card were
re-rendered at another size. Both are now read off the card, a build refuses
a card it cannot read rather than stating a size it guessed, and a missing
card refuses before anything is rendered rather than after.

The no-script check asserted `scripts == 0`, which the data block would end.
A script element whose type is not a script type is never prepared and never
executed, so "static pages, no runtime, nothing for a CSP to have to allow"
is unchanged; the count that replaces it is stricter rather than looser,
because it reddens for an inline script, a `src`, a `type="module"` and a
`type="text/javascript"` alike. The three places that said "no script at all"
in prose now say "no executable script" and name the data block, because a
promise quietly narrowed is worse than a promise restated.
…t-they-are-about

# Conflicts:
#	CHANGELOG.md
Resolves the overlap with GA4 (#83): each page now carries the GA4 loader and the ld+json data block, and the script checks in tests/test_site.py and tests/test_analytics.py count executable scripts so the data block is not mistaken for a second loader.
@ChelseaKR
ChelseaKR merged commit 666a9cb into main Sep 18, 2026
8 checks passed
@ChelseaKR
ChelseaKR deleted the seo/the-pages-say-what-they-are-about branch September 18, 2026 18:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant