Repository navigation
Conversation
The Zenodo record (v1.1.0, 23 Sep) predates the scrape path and the v2.6.0 error handling work, so the note no longer described what the code does. Diffed the paper against HEAD and closed the gaps. Added a subsection, "The pipeline without the LLM" (sec:nollm), under System overview. `dataforge scrape` had only been mentioned in passing, in the agent guide list and as an MCP table row. It now covers: the same politeness gate as a run (robots.txt, per-domain rate limit, Crawl-delay); no key, session or database; page title, Markdown and text; every HTML table extracted as a header row plus data rows, with the three header forms real pages use, colspan and rowspan expanded so the result is rectangular, nested tables kept separate, layout tables skipped and spans capped; the --check rule-based checks (no text, under fifty words, duplicate text); per-table CSV and JSON plus tables.jsonl carrying source_url; and the scrape folder feeding back in as a run's input so collection and generation need not be one decision. Cost control now records the v2.6.0 change: a budget bounds what a working run spends but not what a broken one wastes, so credential and connection failures are exempt from retry, propagate out of the per-chunk handler, and stop dispatch on the first one. Two corrections in Limitations. The claim that tabular data is "flattened or lost" stopped being true in v2.4.5, but only halfway: the extractor serves the no-LLM path, while the dataset pipeline still flattens a table through the Markdown conversion. Routing those rows into chunking is now named as the reachable half of that gap. The lineage paragraph gained the v2.6.0 recurrence, where a try written for one failure class had quietly become the handler for every class. Smaller edits: Processing says what conversion does to a table and points at both paths; the MCP section describes scrape_page as a library call beside explore_site rather than leaving it a bare table row; Discovery cross-references the new subsection; the abstract gains a sentence on the no-LLM path; System overview states that the note describes version 2.7.3. Verified structurally only: refs, cites, braces and environments all balance, but there is no LaTeX toolchain on this machine and no CI job for the .tex, so it has not been through pdflatex. The DOI in the preamble still points at the v1.1.0 record. A new Zenodo version mints a new DOI, so that is set at publish time; the concept DOI in CITATION.cff stays as it is.
The note had four tables and no figures, and three of its arguments are easier to see than to read: the stage structure, why overlapping the stages matters, and why a random split leaks. All five figures are TikZ or pgfplots in the document itself, so there are no binary assets and they diff in git. - Figure 1 (System overview): the six stages, what passes between them, and where the guarantees attach. Bands mark the three stages that overlap as concurrent pools and the two that cannot. Complements Table 1 by showing structure rather than repeating its columns. - Figure 2 (Concurrency): batch vs streaming as a timeline, with the LLM idle span marked on the batch row. Drawn schematically and the caption says so, because the benchmark ran in streaming mode only and the saving is a design rationale, not a measured result. - Figure 3 (Leak-aware splitting): page to chunks to samples, then the same samples under a random split and under a group-aware split. The random panel marks one page landing on both sides of the boundary. This is the paper's central claim and the one that most wanted a picture. - Figure 4 (Empirical evaluation): per-site stage timings as stacked bars, from evals/results/pipeline_metrics_20260922T202407Z.json. Shows USCIS bound by its Crawl-delay and the Python tutorial bound by the LLM, which the surrounding prose argues. - Figure 5 (Empirical evaluation): the quality funnel, 831 generated to 811 approved, with each gate's rejections broken out. Carries the ordering claim that the free filters run before the judge that costs money. Every figure is referenced from the prose. Colours were checked for colour-vision-deficiency separation against a light surface, and every coloured mark also carries a direct label or a letter, so the figures survive greyscale printing. Verified: compiles clean under tectonic (XeTeX), 24 pages, no errors, and each figure was rendered and inspected. Not yet compiled under pdfLaTeX, which is what Overleaf and arXiv run; pgfplots compat is pinned to 1.16 rather than a newer value so it builds on older TeX Live as well.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Two commits. The first brings the technical note's text up to date with
the code; the second gives the note its first figures.
The Zenodo record (v1.1.0, 23 Sep) predates the scrape path and the
v2.6.0 error handling work, so the note no longer described what the
code does. This diffs the paper against HEAD (v2.7.3) and closes the
gaps.
1. Text: up to version 2.7.3
Added: "The pipeline without the LLM" (sec:nollm)
New subsection under System overview.
dataforge scrapehad only beenmentioned in passing, in the agent guide list and as an MCP table row,
so the capability was effectively undocumented in the paper. It now
covers:
Crawl-delay), with no key, session or database
as a header row plus data rows: the three header forms real pages
use, colspan and rowspan expanded so the result is rectangular,
nested tables kept separate, layout tables skipped, spans capped
--checkrule-based checks (no text, under fifty words,duplicate text) and why they are the failures the quality stage would
otherwise spend LLM calls to notice
generation need not be a single decision
Cost control: the v2.6.0 change
A budget bounds what a working run spends but not what a broken one
wastes. Credential and connection failures are now exempt from retry,
propagate out of the per-chunk handler, and stop dispatch on the first
one. Forty chunks with a bad key now ends in about three seconds with
one message, where it previously made three calls per chunk.
Limitations: two corrections
true in v2.4.5, but only halfway. The extractor serves the no-LLM
path, while the dataset pipeline still flattens a table through the
Markdown conversion. Routing those rows into chunking is now named
as the reachable half of that gap.
trywritten for one failure class had quietly become the handler for
every class.
Smaller edits
scrape_pageas a library call besideexplore_site, rather than leaving it a bare table row2. Figures
The note had four tables and no figures, and three of its arguments are
easier to see than to read. All five are TikZ or pgfplots in the
document itself, so there are no binary assets and they diff in git.
them, and where the guarantees attach. Bands mark the three stages
that overlap as concurrent pools and the two that cannot. Complements
Table 1 by showing structure rather than repeating its columns.
the LLM idle span marked on the batch row. Drawn schematically and
the caption says so, because the benchmark ran in streaming mode only
and the saving is a design rationale, not a measured result.
the same samples under a random split and under a group-aware split,
with one page marked landing on both sides of the boundary. This is
the paper's central claim and the one that most wanted a picture.
stacked bars, from
evals/results/pipeline_metrics_20260922T202407Z.json. Shows USCISbound by its Crawl-delay and the Python tutorial bound by the LLM,
which the surrounding prose argues.
generated to 811 approved, with each gate's rejections broken out.
Carries the ordering claim that the free filters run before the judge
that costs money.
Every figure is referenced from the prose. Colours were checked for
colour-vision-deficiency separation against a light surface, and every
coloured mark also carries a direct label or a letter, so the figures
survive greyscale printing.
Verification
Compiles clean under tectonic (XeTeX), 24 pages, no errors. Refs,
cites, braces and environments all balance. Every figure was rendered
and inspected, which caught two label collisions the compiler was
perfectly happy with.
Not yet compiled under pdfLaTeX, which is what Overleaf and arXiv run.
To keep that build uneventful: pgfplots compat is pinned to 1.16 rather
than a newer value so it builds on older TeX Live, nothing uses
\includegraphics, and a local style that shadowed TikZ's built-inthinwas renamed. Worth an Overleaf pass before the next Zenodoupload.
Not in this PR
The DOI in the preamble still points at the v1.1.0 record. A new Zenodo
version mints a new DOI, so that is set at publish time. The concept
DOI in CITATION.cff stays as it is.
The v2.7.1 to v2.7.3 changes are CLI path-resolution fixes and a retag,
judged below the paper's altitude and covered by the version statement
instead.