Skip to content

docs(paper): bring TECHNICAL.tex up to version 2.7.3, and add five figures - #79

Open
ianktoo wants to merge 4 commits into
masterfrom
docs/paper-v2.7.3
Open

ianktoo wants to merge 4 commits into
masterfrom
docs/paper-v2.7.3

Conversation

@ianktoo

@ianktoo ianktoo commented Sep 25, 2026 •

Copy link
Copy Markdown
Owner

Two commits. The first brings the technical note's text up to date with
the code; the second gives the note its first figures.

The Zenodo record (v1.1.0, 23 Sep) predates the scrape path and the
v2.6.0 error handling work, so the note no longer described what the
code does. This diffs the paper against HEAD (v2.7.3) and closes the
gaps.

1. Text: up to version 2.7.3

Added: "The pipeline without the LLM" (sec:nollm)

New subsection under System overview. dataforge scrape had only been
mentioned in passing, in the agent guide list and as an MCP table row,
so the capability was effectively undocumented in the paper. It now
covers:

  • the same politeness gate as a run (robots.txt, per-domain rate limit,
    Crawl-delay), with no key, session or database
  • page title, Markdown and plain text, plus every HTML table extracted
    as a header row plus data rows: the three header forms real pages
    use, colspan and rowspan expanded so the result is rectangular,
    nested tables kept separate, layout tables skipped, spans capped
  • the --check rule-based checks (no text, under fifty words,
    duplicate text) and why they are the failures the quality stage would
    otherwise spend LLM calls to notice
  • per-table CSV and JSON plus tables.jsonl carrying source_url
  • the scrape folder feeding back in as a run's input, so collection and
    generation need not be a single decision

Cost control: the v2.6.0 change

A budget bounds what a working run spends but not what a broken one
wastes. Credential and connection failures are now exempt from retry,
propagate out of the per-chunk handler, and stop dispatch on the first
one. Forty chunks with a bad key now ends in about three seconds with
one message, where it previously made three calls per chunk.

Limitations: two corrections

  1. The claim that tabular data is "flattened or lost" stopped being
    true in v2.4.5, but only halfway. The extractor serves the no-LLM
    path, while the dataset pipeline still flattens a table through the
    Markdown conversion. Routing those rows into chunking is now named
    as the reachable half of that gap.
  2. The lineage paragraph gained the v2.6.0 recurrence, where a try
    written for one failure class had quietly become the handler for
    every class.

Smaller edits

  • Processing says what conversion does to a table and points at both paths
  • the MCP section describes scrape_page as a library call beside
    explore_site, rather than leaving it a bare table row
  • Discovery cross-references the new subsection
  • the abstract gains a sentence on the no-LLM path
  • System overview states that the note describes version 2.7.3

2. Figures

The note had four tables and no figures, and three of its arguments are
easier to see than to read. All five are TikZ or pgfplots in the
document itself, so there are no binary assets and they diff in git.

  • Figure 1 (System overview): the six stages, what passes between
    them, and where the guarantees attach. Bands mark the three stages
    that overlap as concurrent pools and the two that cannot. Complements
    Table 1 by showing structure rather than repeating its columns.
  • Figure 2 (Concurrency): batch vs streaming as a timeline, with
    the LLM idle span marked on the batch row. Drawn schematically and
    the caption says so, because the benchmark ran in streaming mode only
    and the saving is a design rationale, not a measured result.
  • Figure 3 (Leak-aware splitting): page to chunks to samples, then
    the same samples under a random split and under a group-aware split,
    with one page marked landing on both sides of the boundary. This is
    the paper's central claim and the one that most wanted a picture.
  • Figure 4 (Empirical evaluation): per-site stage timings as
    stacked bars, from
    evals/results/pipeline_metrics_20260922T202407Z.json. Shows USCIS
    bound by its Crawl-delay and the Python tutorial bound by the LLM,
    which the surrounding prose argues.
  • Figure 5 (Empirical evaluation): the quality funnel, 831
    generated to 811 approved, with each gate's rejections broken out.
    Carries the ordering claim that the free filters run before the judge
    that costs money.

Every figure is referenced from the prose. Colours were checked for
colour-vision-deficiency separation against a light surface, and every
coloured mark also carries a direct label or a letter, so the figures
survive greyscale printing.

Verification

Compiles clean under tectonic (XeTeX), 24 pages, no errors. Refs,
cites, braces and environments all balance. Every figure was rendered
and inspected, which caught two label collisions the compiler was
perfectly happy with.

Not yet compiled under pdfLaTeX, which is what Overleaf and arXiv run.
To keep that build uneventful: pgfplots compat is pinned to 1.16 rather
than a newer value so it builds on older TeX Live, nothing uses
\includegraphics, and a local style that shadowed TikZ's built-in
thin was renamed. Worth an Overleaf pass before the next Zenodo
upload.

Not in this PR

The DOI in the preamble still points at the v1.1.0 record. A new Zenodo
version mints a new DOI, so that is set at publish time. The concept
DOI in CITATION.cff stays as it is.

The v2.7.1 to v2.7.3 changes are CLI path-resolution fixes and a retag,
judged below the paper's altitude and covered by the version statement
instead.

The Zenodo record (v1.1.0, 23 Sep) predates the scrape path and the
v2.6.0 error handling work, so the note no longer described what the
code does. Diffed the paper against HEAD and closed the gaps.

Added a subsection, "The pipeline without the LLM" (sec:nollm), under
System overview. `dataforge scrape` had only been mentioned in passing,
in the agent guide list and as an MCP table row. It now covers: the same
politeness gate as a run (robots.txt, per-domain rate limit,
Crawl-delay); no key, session or database; page title, Markdown and
text; every HTML table extracted as a header row plus data rows, with
the three header forms real pages use, colspan and rowspan expanded so
the result is rectangular, nested tables kept separate, layout tables
skipped and spans capped; the --check rule-based checks (no text, under
fifty words, duplicate text); per-table CSV and JSON plus tables.jsonl
carrying source_url; and the scrape folder feeding back in as a run's
input so collection and generation need not be one decision.

Cost control now records the v2.6.0 change: a budget bounds what a
working run spends but not what a broken one wastes, so credential and
connection failures are exempt from retry, propagate out of the
per-chunk handler, and stop dispatch on the first one.

Two corrections in Limitations. The claim that tabular data is
"flattened or lost" stopped being true in v2.4.5, but only halfway: the
extractor serves the no-LLM path, while the dataset pipeline still
flattens a table through the Markdown conversion. Routing those rows
into chunking is now named as the reachable half of that gap. The
lineage paragraph gained the v2.6.0 recurrence, where a try written for
one failure class had quietly become the handler for every class.

Smaller edits: Processing says what conversion does to a table and
points at both paths; the MCP section describes scrape_page as a library
call beside explore_site rather than leaving it a bare table row;
Discovery cross-references the new subsection; the abstract gains a
sentence on the no-LLM path; System overview states that the note
describes version 2.7.3.

Verified structurally only: refs, cites, braces and environments all
balance, but there is no LaTeX toolchain on this machine and no CI job
for the .tex, so it has not been through pdflatex.

The DOI in the preamble still points at the v1.1.0 record. A new Zenodo
version mints a new DOI, so that is set at publish time; the concept DOI
in CITATION.cff stays as it is.
The note had four tables and no figures, and three of its arguments are
easier to see than to read: the stage structure, why overlapping the
stages matters, and why a random split leaks. All five figures are TikZ
or pgfplots in the document itself, so there are no binary assets and
they diff in git.

- Figure 1 (System overview): the six stages, what passes between them,
  and where the guarantees attach. Bands mark the three stages that
  overlap as concurrent pools and the two that cannot. Complements
  Table 1 by showing structure rather than repeating its columns.
- Figure 2 (Concurrency): batch vs streaming as a timeline, with the
  LLM idle span marked on the batch row. Drawn schematically and the
  caption says so, because the benchmark ran in streaming mode only and
  the saving is a design rationale, not a measured result.
- Figure 3 (Leak-aware splitting): page to chunks to samples, then the
  same samples under a random split and under a group-aware split. The
  random panel marks one page landing on both sides of the boundary.
  This is the paper's central claim and the one that most wanted a
  picture.
- Figure 4 (Empirical evaluation): per-site stage timings as stacked
  bars, from evals/results/pipeline_metrics_20260922T202407Z.json.
  Shows USCIS bound by its Crawl-delay and the Python tutorial bound by
  the LLM, which the surrounding prose argues.
- Figure 5 (Empirical evaluation): the quality funnel, 831 generated to
  811 approved, with each gate's rejections broken out. Carries the
  ordering claim that the free filters run before the judge that costs
  money.

Every figure is referenced from the prose. Colours were checked for
colour-vision-deficiency separation against a light surface, and every
coloured mark also carries a direct label or a letter, so the figures
survive greyscale printing.

Verified: compiles clean under tectonic (XeTeX), 24 pages, no errors,
and each figure was rendered and inspected. Not yet compiled under
pdfLaTeX, which is what Overleaf and arXiv run; pgfplots compat is
pinned to 1.16 rather than a newer value so it builds on older TeX Live
as well.
@ianktoo ianktoo changed the title docs(paper): bring TECHNICAL.tex up to version 2.7.3 docs(paper): bring TECHNICAL.tex up to version 2.7.3, and add five figures Sep 25, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant