Skip to content

docs: every published page said what it was to a person and nothing to a crawler - #185

Merged
ChelseaKR merged 1 commit into
mainfrom
seo/pages-say-what-they-are-about
Sep 18, 2026
Merged

ChelseaKR merged 1 commit into
mainfrom
seo/pages-say-what-they-are-about

Conversation

@ChelseaKR

Copy link
Copy Markdown
Owner

What was wrong

All 53 indexable pages on sprout.chelseakr.com carry a unique title, a unique description, an absolute self-referencing canonical and a complete card. Every one of those describes the page to a person. Not one of them stated, in any vocabulary a machine reads, what the page is or what the thing on it is.

Measured across the portfolio by parsing every application/ld+json element rather than counting a string: 10 of 25 live sites carry structured data, 15 do not. This was one of the 15.

What it now emits

One block per page — WebSite, WebPage, BreadcrumbList, and Sprout itself as the SoftwareApplication the page is about. The @id values are stable across pages, so a crawler that reads two of them sees one site and one piece of software rather than fifty of each.

Two sources, because the site has two.

  • docs_hooks/structured_data.py writes the block for the 52 mkdocs pages, from page.title, the description page_description already derived from the page's own opening paragraph, page.canonical_url, the nav's ancestors, and mkdocs.yml. Nothing is written for the block — it is handed only what the page already has, so it cannot say something the page does not. The breadcrumb comes from the nav rather than a list, so moving a page in the nav moves its trail instead of stranding it.
  • web-static/public/index.html is copied over mkdocs' index at the site root and has no render to derive from, so its block is written once, in the file. What keeps it honest is the gate, which holds every value in it against that page's own <title>, <meta name=description> and canonical over the deployed tree. Change the title without changing the node and the build fails. That is the same protection the generated pages get, by a different route, and the PR says so rather than pretending the page is generated.

The gate

sprout site-check gains the check, over the deployed tree, wired into the command CI already runs:

exactly one block per page · valid JSON · @context of https://schema.org · only node types that describe a page · no empty property · no @id pointing at a node the graph does not define · a WebPage whose url and description are the page's own and whose name is how its <title> begins · a breadcrumb that is numbered in order and ends at this page by way of addresses the build actually wrote.

53 pages examined / 53 examinable. The two HTML files not examined are 404.html and the noindex eval report, both excluded by the existing _indexable — correctly, since neither is offered for indexing.

It reads the element and its type attribute, never the string application/ld+json. The portfolio audit that asked for this work scored a sibling project as carrying structured data because that string appeared in its HTML, and the single occurrence turned out to be the accept attribute of a file picker. Two tests hold both directions: a file picker is not structured data, and a node written with unusual spacing is still found.

What it refuses

No Dataset node, no DCAT, and the gate rejects both.

A dataset descriptor is not a description, it is an invitation. It exists so that dataset search engines and open-data catalogs harvest the thing it names and list it as a dataset of record, and a catalog listing is far easier to acquire than to withdraw. Whether this project should solicit that over its corpus is an open question with an owner's name on it, parked in chalkline #88. Saying "this page is about a piece of software" asks for none of it. The refusal is a test, so the difference stays a decision somebody makes rather than a line somebody adds during an unrelated change.

A real defect the gate found on its first run

Titles and descriptions reach the hook as HTML: mkdocs keeps a title as it was written in the markdown, and page_description escapes what it writes because mkdocs-material interpolates the value straight into an attribute. JSON carries text, not markup.

Compared raw, 17 pages failed for being correct — and the only way to pass would have been to put &amp; into a node where it means nothing. Both sides are decoded now, and test_an_ampersand_in_a_title_is_not_a_disagreement pins it so the gate cannot regress into failing right answers.

Negative controls

Seven, each with the mutation proved to have landed by a changed git hash-object before the red was believed, each restored to a byte-identical tree (aad175872c931dc73b5588b2eb87a7b17c2f9c5a before and after all seven), __pycache__ cleared between runs, and all seven re-run after the final refactor so they are controls on the code being merged rather than on an earlier draft.

# sabotage fired what went red
1 the hook stops injecting yes 52 pages reported as stating nothing
2 the title emitted still escaped yes the one page with an ampersand
3 description taken from the site config yes 52 pages
4 a Dataset node added to the graph yes harvest refusal on all 52
5 @context set to http://schema.org yes 52 pages
6 the trail ends at a page the build did not write yes 104 reports, both halves
7 the hand-written home page's node renamed yes that page

All seven fired. Controls 2–6 redden the unit tests as well as the site check. Control 1 reddens only the site check, because the unit tests exercise graph() directly and the sabotage is in on_post_page — the two layers catch different things, which is why both are here rather than one standing in for the other.

Also

23 new tests in tests/test_site_meta.py, in the file's existing style — each breaks one property of a known-good tree and asserts the gate notices. ruff check, ruff format --check and mypy clean over src tests docs_hooks; mkdocs build --strict clean; site-check green.

The block is inert data, not a script: nothing is fetched, nothing is reported, and no analytics, beacon, pixel or cookie is involved anywhere in this change.

Prepared with AI assistance; reviewed before submission.

@ChelseaKR

Copy link
Copy Markdown
Owner Author

Merge-order note, so whoever merges second is not surprised.

This PR and the in-flight #182 ("The sitemap said all 53 pages changed today, every day") both touch three files:

  • mkdocs.yml — both register a new hook in the same hooks: list
  • src/sprout/site_meta.py — both extend the same gate module
  • tests/test_site_meta.py — both add tests to the same file

Neither change contradicts the other: #182 dates the sitemap from content, this one states what each page is. They are additive in every one of the three places, so the conflict is textual rather than semantic — both hooks belong in the list, both checks belong in check_site, both sets of tests belong in the file.

Whichever lands first, the second will need a rebase, and the resolution is "keep both" in all three files. Recorded here rather than discovered at merge time.

Prepared with AI assistance; reviewed before submission.

@ChelseaKR
ChelseaKR force-pushed the seo/pages-say-what-they-are-about branch from a443203 to 700200d Compare September 18, 2026 19:28
…o a crawler

All 53 indexable pages carried a unique title, a unique description, an
absolute self-referencing canonical and a full card. Not one of them stated,
in any vocabulary a machine reads, what the page is or what the thing on it
is. Measured across the portfolio by parsing rather than grepping, 10 of 25
live sites carry structured data; this was one of the 15 that carry none.

Each page now emits one `application/ld+json` block: the site, the page, the
nav trail to it, and Sprout itself as the software the page is about. The
`@id` values are stable across pages, so a crawler reading two of them sees
one site and one piece of software rather than fifty of each.

Two sources, because the site has two. `docs_hooks/structured_data.py` writes
the block for the 52 mkdocs pages out of `page.title`, the description
`page_description` already derived from the page's own opening paragraph,
`page.canonical_url`, the nav's `ancestors`, and `mkdocs.yml`. Nothing is
written for the block; it is given only what the page already has, so it
cannot say something the page does not. `web-static/public/index.html` is
copied over mkdocs' index at the root and has no render to derive from, so its
block is written once — and `sprout site-check` holds every value in it against
that page's own title, description and canonical, which is the same protection
by a different route.

`sprout site-check` gains the check, over the deployed tree: exactly one block
per page, valid JSON, `@context` of `https://schema.org`, only node types that
describe a page, no empty property, no `@id` pointing at a node the graph does
not define, a `WebPage` whose url and description are the page's own and whose
name is how the `<title>` begins, and a breadcrumb that ends at this page by
way of addresses the build actually wrote. **53 pages examined of 53
examinable** — `404.html` and the `noindex` eval report are excluded by the
existing `_indexable`, correctly.

It reads the element and its `type` attribute, never the string
`application/ld+json`. The audit that asked for this work scored a sibling
project as carrying structured data because that string appeared on its page,
and the occurrence was the `accept` attribute of a file picker. Two tests hold
both directions: a file picker is not structured data, and a node written with
unusual spacing is still found.

**No `Dataset` node, no DCAT, and the gate refuses them.** A dataset
descriptor is not a description, it is an invitation: it exists so dataset
search engines and open-data catalogs harvest what it names and list it as a
dataset of record, and a listing is far easier to acquire than to withdraw.
Whether this project should solicit that over its corpus is an open question
with an owner's name on it. Saying "this page is about a piece of software"
asks for none of it. The refusal is a test so the difference stays a decision
somebody makes rather than a line somebody adds.

**A real defect the gate found on its first run over the real site.** Titles
and descriptions reach the hook as HTML — mkdocs keeps a title as written in
the markdown, and `page_description` escapes what it writes because
mkdocs-material interpolates it into an attribute. JSON carries text, not
markup. Compared raw, 17 pages failed for being correct, and the only way to
pass would have been to put `&amp;` into a node where it means nothing. Both
sides are decoded now, and a regression test pins it.

Seven negative controls, each with the mutation proved to have landed by a
changed `git hash-object` before the red was believed, each restored to a
byte-identical tree (`aad175872c931dc73b5588b2eb87a7b17c2f9c5a` before and
after all seven), `__pycache__` cleared between runs, and all seven re-run
after the final refactor:

  1. the hook stops injecting            → 52 pages reported as stating nothing
  2. the title emitted still escaped     → the ampersand page reported
  3. description taken from site config  → 52 pages reported
  4. a `Dataset` node added              → harvest refusal on all 52
  5. `@context` set to `http://`         → 52 pages reported
  6. the trail ends at a page not built  → 104 reports, both halves
  7. the hand-written home page renamed  → that page reported

All seven fired. Controls 2-6 redden the unit tests as well as the site
check; control 1 reddens only the site check, because the unit tests exercise
`graph()` directly and the sabotage is in `on_post_page` — the two layers
catch different things, which is why both are here.

The block is inert data, not a script: nothing is fetched, nothing is
reported, and no analytics, beacon, pixel or cookie is involved.
@ChelseaKR
ChelseaKR force-pushed the seo/pages-say-what-they-are-about branch from 700200d to d467770 Compare September 18, 2026 19:33
@ChelseaKR
ChelseaKR merged commit a0c42dc into main Sep 18, 2026
15 checks passed
@ChelseaKR
ChelseaKR deleted the seo/pages-say-what-they-are-about branch September 18, 2026 19:36
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant