Repository navigation
docs: every published page said what it was to a person and nothing to a crawler - #185
Merged
Merged
Conversation
Owner
Author
|
Merge-order note, so whoever merges second is not surprised. This PR and the in-flight #182 ("The sitemap said all 53 pages changed today, every day") both touch three files:
Neither change contradicts the other: #182 dates the sitemap from content, this one states what each page is. They are additive in every one of the three places, so the conflict is textual rather than semantic — both hooks belong in the list, both checks belong in Whichever lands first, the second will need a rebase, and the resolution is "keep both" in all three files. Recorded here rather than discovered at merge time. Prepared with AI assistance; reviewed before submission. |
ChelseaKR
force-pushed
the
seo/pages-say-what-they-are-about
branch
from
September 18, 2026 19:28
a443203 to
700200d
Compare
…o a crawler All 53 indexable pages carried a unique title, a unique description, an absolute self-referencing canonical and a full card. Not one of them stated, in any vocabulary a machine reads, what the page is or what the thing on it is. Measured across the portfolio by parsing rather than grepping, 10 of 25 live sites carry structured data; this was one of the 15 that carry none. Each page now emits one `application/ld+json` block: the site, the page, the nav trail to it, and Sprout itself as the software the page is about. The `@id` values are stable across pages, so a crawler reading two of them sees one site and one piece of software rather than fifty of each. Two sources, because the site has two. `docs_hooks/structured_data.py` writes the block for the 52 mkdocs pages out of `page.title`, the description `page_description` already derived from the page's own opening paragraph, `page.canonical_url`, the nav's `ancestors`, and `mkdocs.yml`. Nothing is written for the block; it is given only what the page already has, so it cannot say something the page does not. `web-static/public/index.html` is copied over mkdocs' index at the root and has no render to derive from, so its block is written once — and `sprout site-check` holds every value in it against that page's own title, description and canonical, which is the same protection by a different route. `sprout site-check` gains the check, over the deployed tree: exactly one block per page, valid JSON, `@context` of `https://schema.org`, only node types that describe a page, no empty property, no `@id` pointing at a node the graph does not define, a `WebPage` whose url and description are the page's own and whose name is how the `<title>` begins, and a breadcrumb that ends at this page by way of addresses the build actually wrote. **53 pages examined of 53 examinable** — `404.html` and the `noindex` eval report are excluded by the existing `_indexable`, correctly. It reads the element and its `type` attribute, never the string `application/ld+json`. The audit that asked for this work scored a sibling project as carrying structured data because that string appeared on its page, and the occurrence was the `accept` attribute of a file picker. Two tests hold both directions: a file picker is not structured data, and a node written with unusual spacing is still found. **No `Dataset` node, no DCAT, and the gate refuses them.** A dataset descriptor is not a description, it is an invitation: it exists so dataset search engines and open-data catalogs harvest what it names and list it as a dataset of record, and a listing is far easier to acquire than to withdraw. Whether this project should solicit that over its corpus is an open question with an owner's name on it. Saying "this page is about a piece of software" asks for none of it. The refusal is a test so the difference stays a decision somebody makes rather than a line somebody adds. **A real defect the gate found on its first run over the real site.** Titles and descriptions reach the hook as HTML — mkdocs keeps a title as written in the markdown, and `page_description` escapes what it writes because mkdocs-material interpolates it into an attribute. JSON carries text, not markup. Compared raw, 17 pages failed for being correct, and the only way to pass would have been to put `&` into a node where it means nothing. Both sides are decoded now, and a regression test pins it. Seven negative controls, each with the mutation proved to have landed by a changed `git hash-object` before the red was believed, each restored to a byte-identical tree (`aad175872c931dc73b5588b2eb87a7b17c2f9c5a` before and after all seven), `__pycache__` cleared between runs, and all seven re-run after the final refactor: 1. the hook stops injecting → 52 pages reported as stating nothing 2. the title emitted still escaped → the ampersand page reported 3. description taken from site config → 52 pages reported 4. a `Dataset` node added → harvest refusal on all 52 5. `@context` set to `http://` → 52 pages reported 6. the trail ends at a page not built → 104 reports, both halves 7. the hand-written home page renamed → that page reported All seven fired. Controls 2-6 redden the unit tests as well as the site check; control 1 reddens only the site check, because the unit tests exercise `graph()` directly and the sabotage is in `on_post_page` — the two layers catch different things, which is why both are here. The block is inert data, not a script: nothing is fetched, nothing is reported, and no analytics, beacon, pixel or cookie is involved.
ChelseaKR
force-pushed
the
seo/pages-say-what-they-are-about
branch
from
September 18, 2026 19:33
700200d to
d467770
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What was wrong
All 53 indexable pages on
sprout.chelseakr.comcarry a unique title, a unique description, an absolute self-referencing canonical and a complete card. Every one of those describes the page to a person. Not one of them stated, in any vocabulary a machine reads, what the page is or what the thing on it is.Measured across the portfolio by parsing every
application/ld+jsonelement rather than counting a string: 10 of 25 live sites carry structured data, 15 do not. This was one of the 15.What it now emits
One block per page —
WebSite,WebPage,BreadcrumbList, and Sprout itself as theSoftwareApplicationthe page isabout. The@idvalues are stable across pages, so a crawler that reads two of them sees one site and one piece of software rather than fifty of each.Two sources, because the site has two.
docs_hooks/structured_data.pywrites the block for the 52 mkdocs pages, frompage.title, the descriptionpage_descriptionalready derived from the page's own opening paragraph,page.canonical_url, the nav'sancestors, andmkdocs.yml. Nothing is written for the block — it is handed only what the page already has, so it cannot say something the page does not. The breadcrumb comes from the nav rather than a list, so moving a page in the nav moves its trail instead of stranding it.web-static/public/index.htmlis copied over mkdocs' index at the site root and has no render to derive from, so its block is written once, in the file. What keeps it honest is the gate, which holds every value in it against that page's own<title>,<meta name=description>and canonical over the deployed tree. Change the title without changing the node and the build fails. That is the same protection the generated pages get, by a different route, and the PR says so rather than pretending the page is generated.The gate
sprout site-checkgains the check, over the deployed tree, wired into the command CI already runs:exactly one block per page · valid JSON ·
@contextofhttps://schema.org· only node types that describe a page · no empty property · no@idpointing at a node the graph does not define · aWebPagewhose url and description are the page's own and whose name is how its<title>begins · a breadcrumb that is numbered in order and ends at this page by way of addresses the build actually wrote.53 pages examined / 53 examinable. The two HTML files not examined are
404.htmland thenoindexeval report, both excluded by the existing_indexable— correctly, since neither is offered for indexing.It reads the element and its
typeattribute, never the stringapplication/ld+json. The portfolio audit that asked for this work scored a sibling project as carrying structured data because that string appeared in its HTML, and the single occurrence turned out to be theacceptattribute of a file picker. Two tests hold both directions: a file picker is not structured data, and a node written with unusual spacing is still found.What it refuses
No
Datasetnode, no DCAT, and the gate rejects both.A dataset descriptor is not a description, it is an invitation. It exists so that dataset search engines and open-data catalogs harvest the thing it names and list it as a dataset of record, and a catalog listing is far easier to acquire than to withdraw. Whether this project should solicit that over its corpus is an open question with an owner's name on it, parked in
chalkline#88. Saying "this page is about a piece of software" asks for none of it. The refusal is a test, so the difference stays a decision somebody makes rather than a line somebody adds during an unrelated change.A real defect the gate found on its first run
Titles and descriptions reach the hook as HTML: mkdocs keeps a title as it was written in the markdown, and
page_descriptionescapes what it writes because mkdocs-material interpolates the value straight into an attribute. JSON carries text, not markup.Compared raw, 17 pages failed for being correct — and the only way to pass would have been to put
&into a node where it means nothing. Both sides are decoded now, andtest_an_ampersand_in_a_title_is_not_a_disagreementpins it so the gate cannot regress into failing right answers.Negative controls
Seven, each with the mutation proved to have landed by a changed
git hash-objectbefore the red was believed, each restored to a byte-identical tree (aad175872c931dc73b5588b2eb87a7b17c2f9c5abefore and after all seven),__pycache__cleared between runs, and all seven re-run after the final refactor so they are controls on the code being merged rather than on an earlier draft.Datasetnode added to the graph@contextset tohttp://schema.orgAll seven fired. Controls 2–6 redden the unit tests as well as the site check. Control 1 reddens only the site check, because the unit tests exercise
graph()directly and the sabotage is inon_post_page— the two layers catch different things, which is why both are here rather than one standing in for the other.Also
23 new tests in
tests/test_site_meta.py, in the file's existing style — each breaks one property of a known-good tree and asserts the gate notices.ruff check,ruff format --checkandmypyclean oversrc tests docs_hooks;mkdocs build --strictclean;site-checkgreen.The block is inert data, not a script: nothing is fetched, nothing is reported, and no analytics, beacon, pixel or cookie is involved anywhere in this change.
Prepared with AI assistance; reviewed before submission.