Skip to content

docs(paper): bring TECHNICAL.tex up to version 2.7.3 and add five figures - #86

Merged
ianktoo merged 7 commits into
masterfrom
docs/paper-updates
Sep 29, 2026
Merged

ianktoo merged 7 commits into
masterfrom
docs/paper-updates

Conversation

@ianktoo

@ianktoo ianktoo commented Sep 29, 2026

Copy link
Copy Markdown
Owner

Pushes the paper work that had been sitting on local master.

Two changes to docs/TECHNICAL.tex, already merged locally as docs/paper-v2.7.3 and docs/paper-figures:

  • Bring the paper up to version 2.7.3, matching the released code.
  • Add five figures.

Net effect against master is docs/TECHNICAL.tex alone, 424 insertions and 22 deletions. No other file changes, and no build artifacts are tracked.

Verified with pdflatex: compiles to 23 pages with no errors, only underfull hbox warnings and the usual rerun-for-cross-references notice.

The Zenodo record (v1.1.0, 23 Sep) predates the scrape path and the
v2.6.0 error handling work, so the note no longer described what the
code does. Diffed the paper against HEAD and closed the gaps.

Added a subsection, "The pipeline without the LLM" (sec:nollm), under
System overview. `dataforge scrape` had only been mentioned in passing,
in the agent guide list and as an MCP table row. It now covers: the same
politeness gate as a run (robots.txt, per-domain rate limit,
Crawl-delay); no key, session or database; page title, Markdown and
text; every HTML table extracted as a header row plus data rows, with
the three header forms real pages use, colspan and rowspan expanded so
the result is rectangular, nested tables kept separate, layout tables
skipped and spans capped; the --check rule-based checks (no text, under
fifty words, duplicate text); per-table CSV and JSON plus tables.jsonl
carrying source_url; and the scrape folder feeding back in as a run's
input so collection and generation need not be one decision.

Cost control now records the v2.6.0 change: a budget bounds what a
working run spends but not what a broken one wastes, so credential and
connection failures are exempt from retry, propagate out of the
per-chunk handler, and stop dispatch on the first one.

Two corrections in Limitations. The claim that tabular data is
"flattened or lost" stopped being true in v2.4.5, but only halfway: the
extractor serves the no-LLM path, while the dataset pipeline still
flattens a table through the Markdown conversion. Routing those rows
into chunking is now named as the reachable half of that gap. The
lineage paragraph gained the v2.6.0 recurrence, where a try written for
one failure class had quietly become the handler for every class.

Smaller edits: Processing says what conversion does to a table and
points at both paths; the MCP section describes scrape_page as a library
call beside explore_site rather than leaving it a bare table row;
Discovery cross-references the new subsection; the abstract gains a
sentence on the no-LLM path; System overview states that the note
describes version 2.7.3.

Verified structurally only: refs, cites, braces and environments all
balance, but there is no LaTeX toolchain on this machine and no CI job
for the .tex, so it has not been through pdflatex.

The DOI in the preamble still points at the v1.1.0 record. A new Zenodo
version mints a new DOI, so that is set at publish time; the concept DOI
in CITATION.cff stays as it is.
The note had four tables and no figures, and three of its arguments are
easier to see than to read: the stage structure, why overlapping the
stages matters, and why a random split leaks. All five figures are TikZ
or pgfplots in the document itself, so there are no binary assets and
they diff in git.

- Figure 1 (System overview): the six stages, what passes between them,
  and where the guarantees attach. Bands mark the three stages that
  overlap as concurrent pools and the two that cannot. Complements
  Table 1 by showing structure rather than repeating its columns.
- Figure 2 (Concurrency): batch vs streaming as a timeline, with the
  LLM idle span marked on the batch row. Drawn schematically and the
  caption says so, because the benchmark ran in streaming mode only and
  the saving is a design rationale, not a measured result.
- Figure 3 (Leak-aware splitting): page to chunks to samples, then the
  same samples under a random split and under a group-aware split. The
  random panel marks one page landing on both sides of the boundary.
  This is the paper's central claim and the one that most wanted a
  picture.
- Figure 4 (Empirical evaluation): per-site stage timings as stacked
  bars, from evals/results/pipeline_metrics_20260922T202407Z.json.
  Shows USCIS bound by its Crawl-delay and the Python tutorial bound by
  the LLM, which the surrounding prose argues.
- Figure 5 (Empirical evaluation): the quality funnel, 831 generated to
  811 approved, with each gate's rejections broken out. Carries the
  ordering claim that the free filters run before the judge that costs
  money.

Every figure is referenced from the prose. Colours were checked for
colour-vision-deficiency separation against a light surface, and every
coloured mark also carries a direct label or a letter, so the figures
survive greyscale printing.

Verified: compiles clean under tectonic (XeTeX), 24 pages, no errors,
and each figure was rendered and inspected. Not yet compiled under
pdfLaTeX, which is what Overleaf and arXiv run; pgfplots compat is
pinned to 1.16 rather than a newer value so it builds on older TeX Live
as well.
@ianktoo
ianktoo merged commit d562b49 into master Sep 29, 2026
21 checks passed
@ianktoo
ianktoo deleted the docs/paper-updates branch September 29, 2026 04:48
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant