Skip to content

fix(documents): collapse whitespace runs inside a text node (#1370) - #1386

Merged
dgunning merged 2 commits into
dgunning:mainfrom
wolfgang-aura:mailman/issue-1370
Sep 30, 2026
Merged

dgunning merged 2 commits into
dgunning:mainfrom
wolfgang-aura:mailman/issue-1370

Conversation

@wolfgang-aura

Copy link
Copy Markdown
Contributor

Closes #1370.

HTMLParser kept the newlines of hard-wrapped <P> text. The issue's own snippet returned 'Our stock has been quoted since August 19,\n2004. Prior to that time there was no market.', where a browser shows one space.

DocumentBuilder._collapse_edges in edgar/documents/strategies/document_builder.py collapsed whitespace at the edges of a text node (the 5.45.0 word-gluing fix) but left runs inside the node alone. The inline-text path near the end of the same file had the same gap.

Both now collapse runs of HTML ASCII whitespace (space, tab, CR, LF, FF) to one space, following the browser's whitespace model:

  • \xa0 is not touched. A non-breaking space is content.
  • Text inside <pre>, or under a white-space: pre / pre-wrap / pre-line / break-spaces style, keeps its whitespace. The check walks up from the element that owns the text.
  • <br> still produces a line break.
  • ParserConfig(preserve_whitespace=True) still skips all of this.

After, on the issue's snippet: 'Our stock has been quoted since August 19, 2004. Prior to that time there was no market.'

How this was tested

Fixtures: inline HTML strings in the tests. No network, no filing downloads.

  • New TestHardWrappedTextNode in tests/documents/test_html_parser_regressions.py has 4 tests: the issue's paragraph, a newline inside an inline <font>, <br> still breaking, and <pre> keeping its newlines.
  • pytest tests/documents tests/test_html*.py tests/test_text.py tests/core/test_markdown_characterization.py tests/core/test_headers_characterization.py -m "not (slow or network or performance or batch)" gave 784 passed at the base commit and 788 passed with the change. tests/test_html.py::test_render_paragraph and ::test_parse_and_identify_headings were deselected on both runs because they raise UnicodeDecodeError reading a fixture on this Windows machine at the base commit too.
  • The 11 test files that import edgar.documents.strategies or that the diff changes, with no marker filter: 208 passed, 0 failed.
  • scripts/check_offline_audit.py's fast-lane command on the changed test file: 61 passed.

I did not re-run the Google FY2004 10-K from the issue, so I can't confirm the reporter's count of about 500 newlines in Item 7 from here. Run on Windows 11 / Python 3.14.3; CI covers the rest of the matrix. Changelog fragment: changelog.d/1370.fixed.md.

AI disclosure

This change was drafted with AI assistance: Claude Opus 5.5 wrote the patch and a second Claude Opus 5.5 pass reviewed it, both running under Mailman. The test results above come from the harness running the commands itself, not from either agent's account of its own work. The account filing this pull request answers for every line of it.

wolfgang-aura and others added 2 commits September 27, 2026 15:38
Hard-wrapped <p> text kept its newlines because _collapse_edges only
collapsed a text node's edges. Runs of ASCII whitespace inside the node
now collapse to one space, except under <pre>, white-space: pre* or
preserve_whitespace=True. <br> and non-breaking spaces are unchanged.
…hitespace

Follow-up to the gh-1370 whitespace collapse, which failed 18 tests in CI.
Six were real losses; the rest were whitespace-only moves.

Losses, now fixed:
- Plain-text filings. Filing.html() wraps a .txt primary document in
  <html><body><div>, and Filing.parsed_items hands the raw .txt straight to the
  parser. Their line breaks are the structure ("Item  8.01  Other Events" on its
  own line), so collapsing them lost 8-K items (GMAC, BayCorp 2005; a 1999 8-K)
  and a 2001 20-F's Item 7. The parser now marks either input
  white-space: pre-wrap, which the collapse already honours. Input that is only
  comments still fails to parse.
- Block children of an inline wrapper. <font><p>..</p>\n<p>..</p></font> is
  separated only by the newline between the <p>s, and _get_element_text
  collapsed it, joining Merck's 2005 CORRESP letter into one line. That break
  is no longer collapsed; runs inside each part still are.
- 424B cover agents. With "per Note.\nBofA Securities, Inc. ("BofAS")" joined
  onto one line, _resolve_abbreviation's capture opened on "Note." A leading
  "Word. " sentence tail is trimmed; "J.P. Morgan" is unaffected.

Whitespace-only moves, re-pinned:
- dt1f1 era titles and wrapped headers: -1 and -2 chars. Word streams of all
  four fixtures are identical before and after.
- Filing text baselines: 4 of 5 filings moved, re-captured with
  scripts/recapture_text_baseline.py, which verified every changed line
  differs by whitespace only. Of the paragraph breaks removed, BofA's 424B2 lost
  367 of 393 mid-sentence; the rest sit inside one text node.

Checked: fast lane 7,736 passed, 0 failed; all 60 section network regression
files pass, one per call; new tests fail on the PR's previous head.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@dgunning

Copy link
Copy Markdown
Owner

Thanks @wolfgang-aura, the approach is right and I've pushed a follow-up commit (2fda300) so CI passes. The first run failed 18 tests. Six were real losses that the collapse exposed:

  • Plain-text filings. Filing.html() wraps a .txt primary document in <html><body><div>, and Filing.parsed_items passes raw .txt straight to the parser. In these, line breaks are the structure ("Item 8.01 Other Events" on its own line), so 8-K items were lost (GMAC and BayCorp 2005, a 1999 8-K), and so was a 2001 20-F's Item 7. The parser now marks both inputs as white-space: pre-wrap, which your _is_preformatted check already honours.
  • Block children of an inline wrapper. <font><p>..</p>\n<p>..</p></font> has only the newline between the <p>s separating them. The collapse in _get_element_text joined Merck's 2005 CORRESP letter into a single line. That break is now kept; runs inside each part still collapse.
  • 424B cover agents. Once "per Note.\nBofA Securities, Inc. (“BofAS”)" is joined onto one line, which is correct, _resolve_abbreviation's capture opened on "Note." It now trims a leading "Word. " sentence tail.

The other 12 only moved whitespace:

  • The dt1f1 length pins moved by 1–2 characters, and the word streams are identical.
  • Four of the five text baselines were re-captured with scripts/recapture_text_baseline.py, which checks that the changes are whitespace-only.

New tests cover each fix. Fast lane, regression and all 60 section network files are green.

🤖 Generated with Claude Code

@dgunning
dgunning merged commit da38745 into dgunning:main Sep 30, 2026
10 of 11 checks passed
dgunning added a commit that referenced this pull request Oct 1, 2026
…n there is no SIGNATURES heading (#1393)

* fix(documents): end an 8-K's last item at the signature statement when there is no SIGNATURES heading

test_eightk_with_no_signature_header failed in the post-merge network job
after #1386 (gh-1370 whitespace collapse). Bisected to da38745, but it is not
a regression: on c5bfcff Item 9.01 already contained the whole signature
block. The test passed only because the hard-wrapped "Pursuant to the\nrequirements"
did not match its substring. #1370 collapsed the wrap and exposed it.

Strategy 5c in pattern_section_extractor: for 8-K/8-K/A, when Strategy 5b
found no bare SIGNATURES line, the first ParagraphNode after the last Item
header whose whitespace-normalized text matches the prescribed opening ("Pursuant
to the requirements of the Securities Exchange Act of 1934 ... caused this
report|amendment to be signed") becomes the terminal header, labelled
SIGNATURES so later stages treat it exactly like a real one. "After the last
Item header" keeps the cover page's "Pursuant to Section 13 or 15(d)" line
from ever ending an item.

Measured on 90 8-Ks (12 fixtures + 78 sampled from 2010Q2/2018Q3/2024Q1),
with and without the change: 3 changed, each a pure truncation of the last
item by exactly its signature block, which moves to a `signatures` section
(80 of 90 already have one). Nothing else moved. 2 filings still carry the
block because it sits inside a table; filed as edgartools-6qmq.

The test now compares whitespace-normalized text and asserts where the item
ends, so it cannot pass by accident again; it fails without the fix.
Section network regression files (60, one per call): 559 passed, 0 failed.
tests/documents + company_reports + meta fast: 869 passed. 8-K test files
including network: 21 + 31 passed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci: the offline audit treats an all-network test file as nothing to audit

#1393 changed only tests/company_reports/test_eightK.py, whose 21 tests are
all network-marked. `-m fast` deselected every one, the output held no
"passed", and check_offline_audit.py took that for a run that produced no
result: exit 2, so the required test-fast job failed in 77 s, before the
fast suite ran. Any PR editing only network test files would hit it.

nothing_to_audit() recognizes "N deselected" with no passed/failed/error
count as a clean result. The "produced no result" guard still catches
collection errors and empty runs. The real case now prints "nothing to
audit" and exits 0; a fast file (test_vcr_interception_guard.py) still
audits normally.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
dgunning added a commit that referenced this pull request Oct 1, 2026
…ssage counts (#1401)

test_2010_20f_resolves_all_eight_wrapped_items_without_the_legacy_parser
has failed locally since #1361 (bf0441b, bisected from 66e3d48 where
the pins held): Item 5 107,457 -> 106,096, Item 6 58,425 -> 57,677, Item
11 7,504 -> 7,081. It skips in CI because its fixture lives in the
gitignored text_boundary_corpus, so nobody saw it; #1386 then took 3
whitespace characters off Item 5 (word stream identical).

The shrinkage is #1361 removing duplicates, not losing text. Against the
source's own item spans: Item 5's technology transfer (July 2009) and W2E
license (January 2010) paragraphs occur once and were printed twice;
Item 6's "247,900", "Company Headcount", "100,000" (4), "50,000" (11) and
"(2)" (15) were each printed once too often; Item 11's forward-contract
paragraph ("notional amount of $1,695") occurs once in Item 11 (its other
copy is in Item 19's notes) and was printed twice. After #1361 every count
matches the source.

Re-pinned, with count assertions for those passages so the next change is
judged on content rather than length; they fail on bf0441b~1. The other
40 tests on the ignored corpus pass.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
@dgunning dgunning mentioned this pull request Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Literal newlines inside a text node are preserved instead of collapsed to a space

2 participants