fix(html): keep head metadata out of markdown without dropping stray content - #2483
Open
Manohar Paturi (ManoharPaturi) wants to merge 1 commit into
Open
Conversation
…content A browser relocates stray content into <body> and keeps document metadata in <head>, but BeautifulSoup's html.parser leaves elements where they were written. The converter converted only the <body> element when one existed, which silently dropped content placed before or after it, and converted the whole document when no <body> existed, which leaked <title> (and other <head>) text into the markdown body. Remove the metadata containers (<head>, and a document <title> that is not inside SVG/MathML) instead of selecting the <body> element, and convert everything that remains. The document title is now captured before removal.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #2482.
The converter now extracts
<head>(capturing the document title first, keeping SVG<title>elements where they are) and converts everything else, instead of converting only<body>when present and the whole document otherwise. Stray content before/after<body>survives, head metadata does not leak.3 new tests fail on main (title leak, dropped pre/post content, SVG title kept) and pass here; suite otherwise unchanged (891 passed; the 5 failures are pre-existing env issues identical on main).