Skip to content

2.2: run scrapes from the service worker and store them in IndexedDB - #8

Open
EnesYilmazcode wants to merge 25 commits into
rebuild/p2-v2.1from
rebuild/p3-run-engine
Open

EnesYilmazcode wants to merge 25 commits into
rebuild/p2-v2.1from
rebuild/p3-run-engine

Conversation

@EnesYilmazcode

@EnesYilmazcode EnesYilmazcode commented Sep 24, 2026 •

Copy link
Copy Markdown
Owner

Stack, merge in order: #5 > #6 > #7 > #8 (this one) > #9. Based on #7, so the diff here is only this step.
The proscan-web stack: EnesYilmazcode/proscan-web#1, EnesYilmazcode/proscan-web#2.

2.2. The content script used to own the run, so run state lived in three places and a stopped worker or a second tab could corrupt it. Now the service worker owns it (F-10, F-11, F-13, F-14, F-18).

  • scripts/background/engine.js drives the run. The live record is in chrome.storage.session, and each page is written to IndexedDB in one transaction (scripts/background/db.js).
  • The content script only parses and reports. The worker opens the next page with tabs.update and parses only that page; a search typed in the tab while it loads ends the run. While a page is pending, the tab sends a heartbeat, which wakes a stopped worker. That is why no alarms permission is needed.
  • If session storage is emptied without an update or a restart (the extension disabled and enabled, or its process crashed), the run IndexedDB still has as live is ended as interrupted when the popup asks, and Stop ends it. The popup is no longer stuck on a run nothing can stop.
  • Starting a run keeps the 10 newest runs and removes older ones, except runs still waiting to sync. When storage is full at start it drops the rest and tries once more, so the storage full message ("starting a new scan removes older saved scans") is true.
  • Every message type is in scripts/lib/messages.js and goes through one router.
  • Migration 3 to 4 moves 2.0 and 2.1 data into IndexedDB. It commits there before removing anything (F-26, F-101, F-102). A retry that runs after a 2.2 run no longer points latestRunId back at 2.1 or overwrites outbox rows the engine wrote.

Tested: npm test 655 jest and 35 tool tests. Playwright with one worker, before this restack: 21 passed, 1 skipped (the same F-20 sync case). That includes killing the worker over CDP between pages and swapping the extension mid-run from both 2.0 and 2.1. After the restack I ran the run tab and 503 scenarios here (all pass) and the full suite on #7 and #9. The page cap test had a race on slow runners, fixed in the last commit before this restack.

Not sure about: Chrome ignores a CDP quota override for extension origins, so storage_full is only covered in jest. I also have not checked whether a real Amazon replaceState can fire a loading event and end a run by mistake. The harness cannot show that.

Does not change permissions. The content script list loses run.js, flags.js and delta.js and gains messages.js:

[permission-lock] OK manifest.json
[permission-lock] OK dist\manifest.json
[version-gate] OK: 2.2.0 > live 2.0

The content script no longer writes storage or navigates. It says
PAGE_READY on load, parses when the worker asks, sends PAGE_RESULT and
keeps a heartbeat while the next page is pending.

The run scenarios of scraper-run.test.js and every test of
scraper-dedupe.test.js now run in engine.test.js, over this script and
the engine, with the same names and checks. scraper-run.test.js keeps
the content script's own contract.

The characterization test drives the engine instead of seeding storage.
Products, pages, next links and helpers are unchanged on every corpus
page. The snapshot changes only in the messages sent and in non-search
pages, which Start now refuses instead of being forced into a run.
…der the 3 to 4 migration

With no session record (the extension was disabled and enabled, or its
process crashed), GET_STATE now ends the durable run as interrupted and
Stop ends it as stopped, so the popup is no longer stuck on a run that
cannot be stopped. Starting a run removes runs past the newest 10,
except ones still in the outbox, and a full disk at start drops the rest
and tries once more, so the storage full message is true. The 3 to 4
migration no longer overwrites latestRunId or outbox rows the engine
wrote first. The engine also checks that the page a tab reports is the
one it opened.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant