Skip to content

fix: make the generated wiki follow the pdf page by page - #124

Merged
Rl0007 merged 33 commits into
mainfrom
fix/nephrology-fidelity
Oct 9, 2026
Merged

Rl0007 merged 33 commits into
mainfrom
fix/nephrology-fidelity

Conversation

@Rl0007

@Rl0007 Rl0007 commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor
  1. Imported PDFs came out as a scattered wiki: sub-points split into separate pages, a section's intro never shown, page headers and footers left in, tables flattened, and charts replaced by made-up tables.
  2. Now the wiki reads like the PDF: the AI groups sub-sections into one page, intros, tables and charts appear where they sit in the PDF, and headers and footers are gone.

You can import a PDF from /wikify and open its generated wiki to see it.

Rl0007 added 19 commits October 7, 2026 01:50

@greptile-apps greptile-apps Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Your trial has ended. Reactivate Greptile to resume code reviews.

@greptile-apps greptile-apps Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Your trial has ended. Reactivate Greptile to resume code reviews.

@greptile-apps greptile-apps Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Your trial has ended. Reactivate Greptile to resume code reviews.

@greptile-apps greptile-apps Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Your trial has ended. Reactivate Greptile to resume code reviews.

@greptile-apps greptile-apps Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Your trial has ended. Reactivate Greptile to resume code reviews.

@greptile-apps greptile-apps Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Your trial has ended. Reactivate Greptile to resume code reviews.

@greptile-apps greptile-apps Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Your trial has ended. Reactivate Greptile to resume code reviews.

@Rl0007
Rl0007 merged commit aed2810 into main Oct 9, 2026
7 checks passed
@vishwajeet-13

Copy link
Copy Markdown
Collaborator

I tested this PR by importing 10 test PDFs (charts, long tables, headers/footers, scanned page, two-column paper, tiny memo, etc.) and generating wikis from them. The header/footer removal and sub-section grouping work well 👍. But I found these bugs:

Bugs seen in real imports

1. Small PDFs give an empty wiki — all text is lost

  • A 1-page memo with 2–3 short headings parses fine, but ends up with 0 sections and an empty wiki.
  • Why: in page_plan.py, short unnumbered sections are never made pages, and a short first section is treated as a "leading fragment". If every section is like that, no page is picked, and fold_sections returns [], so the text is thrown away.
  • Easy repro: plan_pages([Section("Preamble", 30 words), Section("Notes", 20 words)]) returns [].
  • Fix idea: if no page was picked, make the first section a page.
PDF Wiki

2. Charts show up twice, plus the whole PDF page as an image

  • The VLM still writes its own image tags like ![Figure 2…](image2.png), even though the prompt says not to.
  • place_figures adds the real crop too, then repair_broken_image_tags turns the fake image2.png into the full page render.
  • Result: full page image, then the crop, then the full page again, then the crop.
  • Fix idea: when we have crops for a page, place_figures should remove any image tag that isn't one of our crops.

3. Scanned pages: the good VLM output is thrown away

  • The new figure hint tells the VLM to write only a token for a chart, but the judge rubric (verify/judge.py, not changed) still says "a placeholder must score 1".
  • So the judge marks the VLM output as "lost content" (its note: "bar chart replaced with a placeholder… losing significant content"), and we keep the garbled OCR instead.
  • Result: chart numbers turn into junk, and the table loses its Count column and the "AV graft" row.
  • Fix idea: tell the judge that a cropped figure image counts as captured.
PDF Wiki

4. A table over 3+ pages is only joined at the first page break

  • In stitch_cross_page_tables, after page 2's rows move up to page 1, page 2 is empty. Then page 3 is compared with that empty page, so it never joins.
  • Each pair of pages joins fine on its own; only the chain of 3 fails.
  • Result for a 75-row table over 3 pages: the header row shows up again in the middle:
| Acyclovir #48 | Other | 100% | 50-75% | 25% |  |

| Drug | Class | CrCl >50 | CrCl 10-50 | CrCl <10 | Notes |
|---|---|---|---|---|---|
| Amoxicillin #49 | Antimicrobial | ...
  • Fix idea: keep joining into the last page that still has the table, not just k.

5. A figure lands in the wrong section, and its OCR text stays in

  • Two-column paper: the Results chart shows up under "1 Introduction".
  • "3 Results" keeps OCR junk from the chart (625, —e- Standard care, Bs7s, S550...), and its small table is flattened into plain lines.
  • Same OCR junk also appears after the 2-panel figure in the charts PDF and on the scanned page.
Introduction (chart should not be here) Results (junk + flat table)

6. Vector charts are lost

  • Charts drawn as vectors (not pictures) are never found as figures, so only the axis numbers are left (150 / 100 / 50 / 0 / Q1 Q2 Q3 Q4).
  • If this is out of scope, we should at least say so; the PR says charts are fixed.

Bugs found by reading the code

  1. Duplicate table rows: finalize.py doesn't save a page that became empty after stitching, so the old copy stays. The second clean_pages run in sectionize_document joins the same rows again → table rows 1,2,3,4,3,4.
  2. Crash: fold_sections raises KeyError when an unnumbered top-level section with children comes after a leading fragment (that section never gets an owner).
  3. Flowchart text deleted: drop_echoes_after_figures removes short lines after a figure whose words are not in the PDF text layer. That is exactly the flowchart text the new figure hint asks the VLM to write.
  4. Data table read as a contents page: toc_end_line counts numbered table rows ending in a number as contents lines, so a page with a numbered equipment table loses its heading (e.g. 5.3) into a "Contents" section.
  5. Lettered sub-steps merged: merge_continuation_rows runs on every table, not just at page breaks. Rows like | a | gather gloves | get merged into the row above.
  6. Unrelated figure dropped: drop_transcribed_figures deletes a chart if any table of 4+ lines comes right after it, even when the table is unrelated.
  7. Real headings stripped in long docs: the repeated-line limit is now capped at 10 pages, so a heading like "Purpose" that starts many SOPs gets removed as a running header.
  8. Preamble goes to the wrong page: leading text is put on the first leaf page deep in the tree (not the top), and that page's page_start is changed to page 1.

Not from this PR (already broken on main), just FYI

  • Any line with "Page N of M" in it is deleted, even a real sentence like "Page 4 of 6 of the consent form must be signed".
  • Top-level "Appendix A/B" headings get nested under the previous numbered section.
  • A long PDF with no headings becomes one huge page titled "Preamble".

@vishwajeet-13

Copy link
Copy Markdown
Collaborator

Fixes for the real bugs are in #129 (it targets this branch).

Corrections to my comment above, after re-checking:

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants