Discoverability layer: richer paper metadata, real sitemap dates, llms.txt, IndexNow - #4
Merged
studiofarzulla merged 1 commit intoSep 20, 2026
Conversation
…tes, llms.txt, IndexNow Paper pages (build-papers.js): - ScholarlyArticle JSON-LD now built as an object: @id = DOI URL, author @id = ORCID URL, publisher/affiliation @id = https://dissensus.ai/#organization, plus sameAs (arXiv, Zenodo, SSRN, PhilPapers), keywords, abstract, dateModified, license, isPartOf (series / journal when accepted), creativeWorkStatus, encoding (the PDF), mainEntityOfPage, identifier list. - Highwire: citation_author_orcid, citation_author_institution, citation_technical_report_institution, citation_arxiv_id; DC.subject; <link rel=alternate> to the PDF. - New optional papers.json fields: `keywords` (seeded for 17 papers from each PDF's own embedded Keywords field) and `findings` (rendered as a Key findings list; left empty). Site-level: - sitemap.xml: lastmod from paper date/updateDate and news article:published_time; static pages carry none. Every URL used to be stamped with the build date. - research.html: ItemList JSON-LD of every paper keyed by DOI. - News posts: BlogPosting JSON-LD derived from their own meta tags (idempotent sync). - index.html / about.html: shared @id nodes for the Organization and founder. - llms.txt generated from papers.json; robots.txt gains a Content-Signal line and a note on where crawler blocking actually happens; indexnow.js + key file for IndexNow pings. - README: brand name per CLAUDE.md; CLAUDE.md documents the new fields and the lastmod rule. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KkqSm9RbUPUpYZ5TLa7HuF
Deploying with
|
| Status | Name | Latest Commit | Preview URL | Updated (UTC) |
|---|---|---|---|---|
| ✅ Deployment successful! View logs |
dissensus-ai | db0b077 | Commit Preview URL Branch Preview URL |
Sep 19 2026, 11:05 PM |
There was a problem hiding this comment.
Copilot review overview
🔵 Needs a closer look
Correct the superseded DOI for semantic-first-vision and regenerate the derived outputs.
Review effort: Lite
Findings: None
What changed in this PR
Adds richer scholarly metadata, structured data, accurate sitemap dates, llms.txt, crawler signals, and IndexNow support.
Changes:
- Enriches paper and news JSON-LD and metadata.
- Generates content-aware sitemaps and
llms.txt. - Adds shared entity identifiers, crawler signals, and IndexNow tooling.
- Requires correcting the superseded DOI used for
semantic-first-vision.
| File | Summary |
|---|---|
README.md |
Documents discoverability workflow. |
public/terms.html |
Updates stylesheet cache version. |
public/sitemap.xml |
Uses content dates. |
public/robots.txt |
Adds crawler content signals. |
public/research.html |
Adds publication ItemList JSON-LD. |
public/projects.html |
Updates stylesheet cache version. |
public/privacy.html |
Updates stylesheet cache version. |
public/papers/whitepaper-factor-analysis.html |
Adds enriched paper metadata. |
public/papers/trauma-training-data.html |
Adds enriched paper metadata. |
public/papers/temporal-bitmap-interpretation.html |
Adds enriched paper metadata. |
public/papers/substrate-independent-friendship.html |
Adds enriched paper metadata. |
public/papers/stakes-without-voice.html |
Adds enriched paper metadata. |
public/papers/sentiment-abm.html |
Adds enriched paper metadata. |
public/papers/semantic-first-vision.html |
Adds enriched metadata; DOI requires correction. |
public/papers/replicator-optimization-mechanism.html |
Adds enriched paper metadata. |
public/papers/marl-coordination.html |
Adds enriched paper metadata. |
public/papers/market-reaction-asymmetry.html |
Adds enriched paper metadata. |
public/papers/identity-thesis.html |
Adds enriched paper metadata. |
public/papers/hedging-paradox.html |
Adds enriched paper metadata. |
public/papers/genre-mimicry.html |
Adds enriched paper metadata. |
public/papers/first-crisis-governance.html |
Adds enriched paper metadata. |
public/papers/consent-to-consideration.html |
Adds enriched paper metadata. |
public/papers/consensual-sovereignty.html |
Adds enriched paper metadata. |
public/papers/consciousness-monograph.html |
Adds enriched paper metadata. |
public/papers/cbdc-privacy.html |
Adds enriched paper metadata. |
public/papers/axiom-of-consent.html |
Adds enriched paper metadata. |
public/papers/autonomous-red-team.html |
Adds enriched paper metadata. |
public/papers/asri.html |
Adds enriched paper metadata. |
public/papers/alpha-asymmetry-fx.html |
Adds enriched paper metadata. |
public/news/trident.html |
Adds BlogPosting JSON-LD. |
public/news/temporal-bitmap.html |
Adds BlogPosting JSON-LD. |
public/news/research-update-september-2026.html |
Adds BlogPosting JSON-LD. |
public/news/lab-update-aug2026.html |
Adds BlogPosting JSON-LD. |
public/news/incorporation.html |
Adds BlogPosting JSON-LD. |
public/news/iaseai-affiliate.html |
Adds BlogPosting JSON-LD. |
public/news/digital-finance-accept.html |
Adds BlogPosting JSON-LD. |
public/news.html |
Updates stylesheet cache version. |
public/llms.txt |
Provides generated paper catalogue. |
public/join.html |
Updates stylesheet cache version. |
public/index.html |
Adds shared entity identifiers. |
public/css/site.css |
Styles optional findings. |
public/about.html |
Adds organization and founder metadata. |
public/404.html |
Updates stylesheet cache version. |
public/31dde54ee5c4516b4c1d7949cd7eecc7.txt |
Publishes IndexNow verification key. |
papers.json |
Adds paper keywords and metadata. |
indexnow.js |
Submits sitemap URLs to IndexNow. |
CLAUDE.md |
Documents metadata and sitemap rules. |
build-papers.js |
Generates metadata, JSON-LD, sitemap, and llms.txt. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Makes every paper landing page a better retrieval object for Google Scholar, the citation-graph indexers (OpenAlex, Semantic Scholar) and LLM search crawlers, without changing any visible content.
Paper pages (
build-papers.js)ScholarlyArticleJSON-LD is now built as an object rather than a string template.@idis the DOI URL, the author@idis the ORCID URL, publisher/affiliation@idishttps://dissensus.ai/#organization. AddssameAs(arXiv, Zenodo, SSRN, PhilPapers),keywords,abstract,dateModified,license,isPartOf(series, plus the journal once accepted),creativeWorkStatus,encoding(the PDF),mainEntityOfPageand an identifier list (DOI, arXiv, paper number).citation_author_orcid,citation_author_institution,citation_technical_report_institution,citation_arxiv_id. Dublin Core gainsDC.subject. A<link rel="alternate" type="application/pdf">points at the PDF.papers.jsonfields:keywords(seeded for 17 papers from each PDF's own embedded Keywords field, so nothing is invented) andfindings(rendered as a "Key findings" list under the abstract).findingsis empty everywhere on purpose: several papers carryupdateNotes withdrawing earlier claims, so those lists should be written from the current manuscripts.Site level
sitemap.xml:lastmodnow comes from each paper'sdate/updateDateand each news post'sarticle:published_time; the static pages carry none. Every URL used to be stamped with the build date, which is the one pattern Google documents as untrustworthy.research.html: anItemListJSON-LD of every paper keyed by DOI, inside the generated block.BlogPostingJSON-LD block derived from each post's existing meta tags (marker-delimited, idempotent).index.html/about.html: shared@idnodes for the Organization and founder. farzulla.com uses the same@ids, so the two sites now describe one author, one lab, one set of works.public/llms.txtgenerated frompapers.json.robots.txt:Content-Signal: search=yes, ai-input=yes(no statement on training) and a note that crawler blocking, where it happens, is the Cloudflare zone policy, not this file.indexnow.js+ public key file:node indexnow.jsafter a deploy submits the sitemap URLs to the IndexNow engines.lastmodrule.Verification
node build-papers.jsruns clean; a second run produces an identical sitemap and exactly one JSON-LD block per news post.public/parse; everyScholarlyArticlecarries the new properties.sitemap.xmlis well-formed, 37 URLs, 29 with a real date, no.htmltwins.Referencesheading (Scholar's inclusion requirements).Not in this PR (needs the account owner)
Semantic Scholar currently lists only the 9 arXiv-hosted papers (two of them duplicated); the 13 Zenodo/SSRN/PhilPapers-only papers need the author-page "add paper" flow. OpenAlex profile claim, Zenodo related-identifier links, Bing Webmaster Tools and the Cloudflare AI-bot policy are all dashboard actions.
🤖 Generated with Claude Code
https://claude.ai/code/session_01KkqSm9RbUPUpYZ5TLa7HuF
Generated by Claude Code