docs: the note now carries what the launch posts promise - #20
Merged
Merged
Conversation
Every launch post ends by pointing here, and the note was written six days before half of what they describe existed. Three gaps. Markdig was never named. The note said the words were "recovered with a Markdown parser", which is the vaguest possible phrasing of the strongest claim in the project. A .NET reader wants the library, and wants to know why it has to be one we did not write: our own reader would share our escaping assumptions, forget to interpret what the serialiser forgot to escape, and agree the document was fine. Shown now with the welded-emphasis case, where Markdig reads two words back as one token. The serialiser had one paragraph. It now has a section, because that is where the difficult problems are: two escaping passes, why code spans are deliberately not escaped, and the rule that drops formatting from a span with no letters or digits in it. Rendered faithfully, one document in the corpus would have lost 3,440 of its 4,664 words to an italic full stop. --report did not exist when this was written. It does now, and it is the only channel there is for documents we will never be allowed to see. Also corrects the performance section. It described the chained-predicate bug as the finding; the larger one is that any predicate in a match pattern is expensive and its content is irrelevant, which is a different defect, filed separately as phoenixmldb-xslt#95. The old text also promised the engine fix was probably done by now. It is not. Attribution fixed throughout: "one real document lost 3,440 words" had no subject and read as though docmd lost them. A naive serialiser would have; the coverage check is why it never shipped. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018wZEgtuzaZPswykiGEBxbz
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Reopening — #15 was closed without
merging on 28 Sep and the author did not intend to. Rebased onto current
main, no conflicts;nothing has touched the note since.
Every launch post ends by pointing at this note, and it was written six days before half of what
they describe existed.
Markdig was never named
The note said the words were "recovered with a Markdown parser" — the vaguest possible phrasing
of the strongest claim in the project, to an audience that knows Markdig. It now says which
library, and why it has to be one we did not write:
Our own reader would share our escaping assumptions — forget to interpret exactly what the
serialiser forgot to escape — and both halves would agree the document was fine.
The serialiser had one paragraph
Now a section, because that is where the difficult problems are: two escaping passes and why
line-leading characters need the second, why code spans are deliberately not escaped, and the
rule that drops delimiters from a span containing no letters or digits.
--reportdid not exist when this was writtenIt does now, and it is the only channel there is for documents we will never be allowed to see.
The performance section was out of date and slightly wrong
It described the chained-predicate bug as the finding. The larger one is that any predicate in
a match pattern was expensive and its content irrelevant — filed as
#95, since fixed. The old text
also promised the engine fix was probably done by the time you read it, which was untrue when
written and is now true for the wrong reason.
Chained predicates remain quadratic — 21.5 s at n=1500 on 2.4.1 — so that half of the story
stands.
Attribution
"One real document lost 3,440 of its 4,664 words" had no subject and reads as though docmd
lost them. A naive serialiser would have; the coverage check is why it never shipped.
Note goes from 1,836 to 2,821 words.
🤖 Generated with Claude Code
https://claude.ai/code/session_018wZEgtuzaZPswykiGEBxbz