Skip to content

Add a saved-page corpus and a Chromium test harness - #6

Open
EnesYilmazcode wants to merge 15 commits into
rebuild/p0-guardrailsfrom
rebuild/p1-tests
Open

EnesYilmazcode wants to merge 15 commits into
rebuild/p0-guardrailsfrom
rebuild/p1-tests

Conversation

@EnesYilmazcode

@EnesYilmazcode EnesYilmazcode commented Sep 24, 2026 •

Copy link
Copy Markdown
Owner

Stack, merge in order: #5 > #6 (this one) > #7 > #8 > #9. Based on #5, so the diff here is only this step.
The proscan-web stack: EnesYilmazcode/proscan-web#1, EnesYilmazcode/proscan-web#2.

Before changing the scraper I wanted tests that run against real Amazon markup instead of hand-written snippets, which is how F-27 and F-16 went unnoticed.

  • tests/pages: sanitized saved pages (search, product, AOD fragment, captcha) with an expected.json for each. Scripts, tokens and session values are stripped.
  • Search and offer parsing moved into scripts/lib/parsers.js. A characterization snapshot taken before the move still matches after it.
  • A Playwright harness loads the built extension in Chromium against a local fake Amazon. It covers pagination, a stopped service worker, Stop, second tabs, storage quota and upgrading in place from the real 7c2ba1c (v2.0) build.

Bugs the corpus found are marked as expected failures with their F-id in the name (F-11, F-12, F-13, F-15, F-16, F-17, F-20, F-26, F-27, F-30, F-37, F-100, F-101). The header result count ("of over 10,000 results" read as 0) has no audit finding, so it is tagged NEW-PARSE-1. Later PRs in the stack flip them to passing. A test.failing passes on any error, so every corpus page also has plain checks that must pass today: the parsers run without throwing and return the right shape. A crash in the offer parser now fails the corpus instead of hiding behind a known failure. Each Playwright scenario checks its setup before the buggy assertion, so a broken harness still fails loudly. (F-81, F-83)

Tested: npm test is 377 jest (19 of them test.failing) and 35 tool tests. npx playwright test is 13/13, 10 of them expected failures. The suite runs on Windows locally and on Linux (ubuntu-latest) in CI.

Does not change permissions. The manifest diff only adds parsers.js to the content script list:

[permission-lock] OK manifest.json
[permission-lock] OK dist\manifest.json
[version-gate] OK: 2.1.0 > live 2.0

Strips scripts, frames and inline handlers, and scrubs session and
tracking tokens, so captured pages can live in the repo as fixtures.
Real captures from the audit (a yoga mat search page, the Akamai
interstitial, an aodAjaxMain fragment and a product page), hand-written
pages for a captcha, title-recipe cards and a missing pagination strip,
and copies of the old fixtures. Each page has an expected.json written
from the page, not from the current code.
A characterization test, bugs included, so extracting the parsers can
prove it changed nothing. Real pages carry CSS jsdom cannot parse, so the
content-script loader now keeps that error out of the test output.
The content scripts, the popup and the tests now share one pure module
(global Parsers, CommonJS in Node). scraper.js and offer-fetcher.js keep
their function names as thin delegates, and Analyzer's coercions live
there too. No behavior change: the corpus snapshot is unchanged.
Adds a parseDoc helper that builds a jsdom Document with the same
innerText fallback the content-script loader uses.
Each check a page's expected.json makes (page kind, ASIN dedupe, sponsored
count, ad links, next page, total, fill rate, spot products, seller
prices) runs per page. The ones today's code fails run as test.failing
with their finding id, so the fix that makes one pass has to take it
off the list.
At most 5 amazon.com fetches, 3 seconds apart. Each page is sanitized
and compared with the previous capture before it is written, and only
after a y. It refuses to run in CI.
global-setup builds the prod extension, lib/extension.mjs loads it into
Chromium with a fresh profile and drives the real popup, and
lib/amazon.mjs answers every request from generated pages or the corpus,
so no request reaches Amazon. One Chromium at a time.
A clean run, the saved yoga mat page against the parser, a captcha,
duplicates, Stop, a second search tab, a product tab, a slow page, a
full storage quota, two runs without an export and a stopped service
worker. Scenarios that fail on today's code call bug() with the finding
id right before the failing checks, so they run as expected failures.
Target.closeTarget returns before the worker is gone, which made the
service worker scenario flaky.
The harness checks out commit 7c2ba1c (the store build) with a temporary
index, loads it unpacked, scrapes, then swaps in the current build and
reloads, like an auto-update. Both scenarios are known failures: nothing
migrates v2.0 data (F-100), and a run in flight is resumed with no run
id (F-101).
A separate job installs the full Chromium build for the pinned
Playwright version and runs it in the new headless mode, which loads
extensions. Full history is fetched for the upgrade scenarios.
A test.failing passes on any error, so a crash in the offer parser left
the aod-pinned-only page green. Each page now also checks that the
parsers run without throwing and return the right shape. The header
count check was tagged F-25, which is the dashboard schema finding; it
is now NEW-PARSE-1, since no audit finding covers it.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant