Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
25 commits
Select commit Hold shift + click to select a range
cb5b19d
Add a message catalog and a typed router for the service worker
EnesYilmazcode Sep 24, 2026
24f2b00
Add an IndexedDB store for runs, products, pages, lastValues and the …
EnesYilmazcode Sep 24, 2026
90d5036
Give the run record states from idle to done and a reason for how it …
EnesYilmazcode Sep 24, 2026
c6b38ab
Run scrapes from a service worker engine that owns the run and its data
EnesYilmazcode Sep 24, 2026
077e61c
Move run data out of storage.local into IndexedDB as schema 4
EnesYilmazcode Sep 24, 2026
ec83596
Answer every worker message through the router and wire in the engine
EnesYilmazcode Sep 24, 2026
c3b18fc
Make the scraper parse only and report pages to the worker
EnesYilmazcode Sep 24, 2026
5caeaf7
Send spread results to the worker instead of writing storage from the…
EnesYilmazcode Sep 24, 2026
cb63631
Render the popup from the worker's state and start and stop through it
EnesYilmazcode Sep 24, 2026
a2c26a3
Ship messages.js to pages, drop unused content globals, bump to 2.2.0
EnesYilmazcode Sep 24, 2026
0cc1f7e
Tell extension pages from content scripts by URL, since a popup can o…
EnesYilmazcode Sep 24, 2026
01b58b3
Open the database again when another context closes it, and end the r…
EnesYilmazcode Sep 24, 2026
14a2528
Show how a run ended even when its database cannot be opened
EnesYilmazcode Sep 24, 2026
51329a7
Read run state from session storage and IndexedDB in the Chromium har…
EnesYilmazcode Sep 24, 2026
2029a6f
Upgrade harness: expect 2.2.0 and schema 4, and cut a 2.1 run off mid…
EnesYilmazcode Sep 24, 2026
0dfb3a9
Keep the specs and tests of checked-out older builds out of the test …
EnesYilmazcode Sep 24, 2026
be49263
Chromium harness: a run tab taken elsewhere ends the run and stays wh…
EnesYilmazcode Sep 24, 2026
7fde97c
Document the service worker run engine, IndexedDB storage and schema 4
EnesYilmazcode Sep 24, 2026
7502721
End a run cut off by an update as updated, not interrupted
EnesYilmazcode Sep 24, 2026
f9b9b73
Ask a page again when it has not reported 20 seconds after the worker…
EnesYilmazcode Sep 24, 2026
433a794
End a run whose tab left Amazon while its next page was loading
EnesYilmazcode Sep 24, 2026
95e0204
Report a page again when the worker re-asks after every try failed
EnesYilmazcode Sep 24, 2026
b96d1fa
Let the install defaults land before the page cap test sets its own
EnesYilmazcode Sep 24, 2026
e2f5ad4
End runs left live only in IndexedDB, keep the newest 10 runs, and or…
EnesYilmazcode Sep 24, 2026
7ce2963
Read the run-tab and 503 scenarios through the 2.2 run record
EnesYilmazcode Sep 24, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
89 changes: 61 additions & 28 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,29 +4,33 @@

Chrome extension (Manifest V3) that scrapes Amazon seller product listings, provides analytics for resellers/arbitrage, and includes a floating AI chatbot on Amazon pages powered by Gemini API. No server required — everything runs client-side.

## Architecture (v2.0)
## Architecture (v2.2)

```text
AmazonSellerScraper/
├── manifest.json # Extension config (v2.0)
├── manifest.json # Extension config (v2.2)
├── popup/ # UI Layer
│ ├── popup.html # Popup interface (dashboard + settings)
│ ├── popup.css # Popup styling
│ ├── popup.js # UI logic
│ └── ai-key.js # Gemini key field (AI chat settings)
├── scripts/
│ ├── content/
│ │ ├── scraper.js # DOM scraping on Amazon pages
│ │ ├── scraper.js # Parses a search page, reports it to the SW (parse only)
│ │ ├── chatbot.js # Floating AI chatbot widget (Shadow DOM)
│ │ └── offer-fetcher.js # Seller price fetching for spread analysis
│ ├── lib/
│ │ ├── parsers.js # Pure search/offer parsing (global Parsers)
│ │ ├── run.js # Scrape run record, bound to one tab (global Run)
│ │ ├── flags.js # Build flags (global Flags); CLOUD_SYNC is off in 2.1
│ │ ├── messages.js # Every message type and who may send it (global Msg)
│ │ ├── run.js # Run state machine, bound to one tab (global Run)
│ │ ├── flags.js # Build flags (global Flags); CLOUD_SYNC is off until 2.3
│ │ ├── migrate.js # Storage schema migrations, run by the SW
│ │ └── chat.js # Gemini request builder, run scoping, error text
│ ├── background/
│ │ └── service-worker.js # Message routing + Gemini API calls
│ │ ├── service-worker.js # Wires router, engine, chat, auth and migrations
│ │ ├── router.js # The one typed message router
│ │ ├── engine.js # The run engine; the only writer of run data
│ │ └── db.js # IndexedDB: runs, products, placements, lastValues, outbox, spread
│ └── modules/
│ ├── storage.js # Chrome storage wrapper
│ ├── analyzer.js # Data analysis & insights
Expand Down Expand Up @@ -59,37 +63,52 @@ AmazonSellerScraper/

| Module | Purpose |
| -------------------- | --------------------------------------------------------- |
| `scraper.js` | DOM scraping with cascading fallback selectors |
| `scraper.js` | Parses a page with `Parsers`, sends `PAGE_RESULT` |
| `engine.js` | Run state machine, page saves, navigation, heartbeat |
| `chatbot.js` | Floating AI chatbot widget on Amazon pages (Shadow DOM) |
| `offer-fetcher.js` | Fetches seller offer pages for price spread analysis |
| `chatbot.css` | Widget styles loaded into Shadow DOM |
| `storage.js` | Async wrapper for chrome.storage.local |
| `analyzer.js` | Opportunity scoring, insights, statistics |
| `spread-analyzer.js` | Price spread statistics (CV, std dev, arbitrage scoring) |
| `exporter.js` | Multi-format export (Excel, CSV, JSON) |
| `service-worker.js` | Message routing, Gemini API calls |
| `service-worker.js` | Router wiring, Gemini API calls, migrations |

## Data Flow

### Scraping

1. User clicks "Start Scraping" in popup
2. `popup.js` pings the tab; with no answer it offers `chrome.tabs.reload` and starts after the reload
3. `popup.js` writes a run record (`scripts/lib/run.js`, key `run`) bound to the tab id, then sends `START_SCRAPING`
4. `scraper.js` classifies the page (results, last, empty, captcha, interstitial, signin, unknown) and extracts products
5. Results and run progress stored in `chrome.storage.local` in one write per page
6. Follows the page's Next link after 2 to 4 seconds, up to `settings.maxPages`; on each load the content script asks the SW for its tab id (`WHO_AM_I`) and acts only if its tab owns the run
7. The run ends with a reason; Stop goes to the run's tab and cancels the pending navigation
8. `analyzer.js` generates insights and opportunity scores
9. User exports via `exporter.js` (Excel/CSV/JSON)
The service worker is the only writer (`scripts/background/engine.js`).

1. User clicks "Start Scraping" in popup; `popup.js` sends `START_RUN {tabId}` to the SW
2. The SW pings the tab. No answer: the popup offers `chrome.tabs.reload` and starts after the reload.
A page that is not a search (captcha, sign-in, product page) is refused and nothing changes
3. The SW writes the run record (`scripts/lib/run.js`) to `chrome.storage.session` under `run`,
bound to the tab id, and asks the tab to parse (`PARSE_PAGE`)
4. `scraper.js` classifies the page (results, last, empty, captcha, interstitial, signin, unknown),
extracts products and sends `PAGE_RESULT`. It never writes storage and never navigates
5. The SW folds the page into the run and writes products, the page, lastValues and the run
to IndexedDB in one transaction, then prunes lastValues
6. After 2 to 4 seconds the SW opens the page's Next link with `tabs.update`, up to `settings.maxPages`.
The new page says `PAGE_READY`; only the run's tab is asked to parse
7. While a page is pending the tab sends `HEARTBEAT` every 2 seconds. That wakes a stopped SW,
which reads the run from session storage and opens the next page when it is due
8. The run ends with a reason; Stop goes to the SW, which cancels the pending page. Closing the
run's tab or taking it elsewhere ends the run as `interrupted`
9. `analyzer.js` generates insights and opportunity scores
10. User exports via `exporter.js` (Excel/CSV/JSON)

Run states: idle, starting, running, stopping, then stopped, blocked, failed or done.
`reason` says why it ended: complete, stopped, blocked, selectors_broken, storage_full,
storage_error, interrupted or updated.

### AI Chatbot (client-side, no server)

1. `chatbot.js` injects a floating widget (bottom-right) on Amazon pages with product listings
2. Widget uses Shadow DOM to isolate styles from Amazon's CSS
3. User types a question (e.g. "What's the best deal under $30?")
4. `chatbot.js` sends `CHAT_MESSAGE` to `service-worker.js` with the question and the last few turns
5. The service worker reads the user's key and the current run from `chrome.storage.local`; the content script never sees the key
5. The service worker reads the user's key from `chrome.storage.local` and the latest run from IndexedDB; the content script never sees the key
6. `scripts/lib/chat.js` builds the request: model id in `GEMINI_MODEL`, key in the `x-goog-api-key` header, titles in a fenced JSON block marked untrusted
7. Response displayed in chat bubble as text, never HTML

Expand All @@ -114,16 +133,23 @@ npm run test:coverage # With coverage report
**Test architecture:**
- Pure logic modules (`analyzer.js`, `spread-analyzer.js`) are tested via `require()` directly
- Content scripts (`scraper.js`, `offer-fetcher.js`) have no `module.exports` — loaded via `vm.runInContext` into a JSDOM context with Chrome API mocks and an `innerText` polyfill
- `tests/setup/engine-rig.js` runs the real router and engine over fake-indexeddb with the real `scraper.js` in a JSDOM page per tab; `tests/unit/engine.test.js` drives runs through it
- `npm run test:e2e` runs the same scenarios in Chromium (`tests/e2e/`), including the worker stopped via CDP between pages and 2.0 and 2.1 builds updated mid-run
- HTML fixtures in `tests/fixtures/` match the exact CSS selectors the code uses
- Chrome APIs (`storage`, `runtime`, `tabs`, `downloads`) are mocked in `tests/setup/chrome-mock.js`
- XLSX is mocked with jest.fn() stubs in exporter.test.js; export-xlsx.test.js uses the real libs/xlsx.full.min.js

## Chrome APIs Used

- `chrome.storage.local` — state persistence
- `chrome.runtime.sendMessage/onMessage` — inter-script communication
- `chrome.downloads` — file downloads
- `chrome.tabs` — active tab messaging
- `chrome.storage.local`: settings, the Gemini key and `schemaVersion` only
- `chrome.storage.session`: the live run record (content scripts cannot read it)
- IndexedDB: run data, in the extension origin; needs no permission
- `chrome.runtime.sendMessage/onMessage`: messages, all in `scripts/lib/messages.js`
- `chrome.downloads`: file downloads
- `chrome.tabs`: `sendMessage`, `update`, `onRemoved`, `onUpdated`; none need the `tabs` permission

No permission was added for the run engine: no `alarms`, `scripting`, `offscreen` or
`unlimitedStorage`. `npm run lock` fails on any addition.

## DOM Selectors (Amazon-specific, updated Sep 2026)

Expand Down Expand Up @@ -175,17 +201,24 @@ One record per ASIN per run, from `Parsers.parseSearchPage` and the scraper:

## Storage

- `chrome.storage.local` carries `schemaVersion` (3). Storage without it came from 2.0.
- `chrome.storage.local` carries `schemaVersion` (4) and settings. Storage without it came from 2.0.
`scripts/lib/migrate.js` brings it up to date on install, update and browser start;
each step is idempotent. 2 to 3 adds `priceCents`, nulls and `/dp/` URLs to 2.0 rows,
seeds `lastValues` from them (with `firstSeenAt`), ends a run the update cut off as
`updated`, and drops the untouched 2.0 default settings.
`updated`, and drops the untouched 2.0 default settings. 3 to 4 moves results, pages,
lastValues, spread data and the run into IndexedDB (a run still running is `updated`);
the IndexedDB writes commit before anything is removed.
- IndexedDB `proscan` (`scripts/background/db.js`): `runs`, `products` (key `[runId, n]`,
n is the order found), `placements` (one record per page), `lastValues` (by ASIN),
`outbox` (pages waiting for sync), `spread` and `meta` (`latestRunId`). The popup shows
the latest run's products.
`tests/fixtures/v2.0-storage.json` is what the live 2.0 build stores, captured with
`node tests/e2e/capture-v20-storage.mjs`.
- `lastValues` is capped by `Delta.prune`: 5,000 ASINs, none older than a year.
- `Storage` rejects when `chrome.runtime.lastError` is set (code `storage_full` on a
quota error). The popup warns from 90% of the quota.
- Cloud sync is off in 2.1 (`Flags.CLOUD_SYNC`): nothing goes into `syncQueue`, the SW
- `lastValues` is pruned after every page: 5,000 ASINs, none older than a year
(`Delta.MAX_ENTRIES`, `Delta.MAX_AGE_DAYS`).
- A page write that fails ends the run as `storage_full` (quota) or `storage_error`.
The popup warns from 90% of `navigator.storage.estimate()`.
- Cloud sync is off in 2.1 and 2.2 (`Flags.CLOUD_SYNC`): nothing goes into the outbox, the SW
refuses `PROSCAN_EXPORT`, and the popup hides sign-in and Export to ProScan.

## Export
Expand Down
35 changes: 22 additions & 13 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -58,11 +58,14 @@ ProScan is a Chrome extension that scrapes Amazon product listings across multip

```
User clicks "Start Scraping"
→ popup.js pings the tab (offers a reload if ProScan is not loaded there)
→ popup.js creates a run bound to that tab and sends START_SCRAPING
→ scraper.js classifies the page, then extracts products
→ Results and run progress stored in chrome.storage.local
→ Follows the page's Next link after a 2 to 4 second delay
→ popup.js sends START_RUN to the service worker
→ The worker pings the tab (the popup offers a reload if ProScan is not
loaded there), makes a run bound to that tab and asks it to parse
→ scraper.js classifies the page, extracts products and sends PAGE_RESULT;
it never writes storage and never navigates
→ The worker saves the page to IndexedDB in one transaction
→ After a 2 to 4 second delay the worker opens the page's Next link
with tabs.update, and the new page reports in
→ Repeats until the last page or the page cap (settings.maxPages, default 20)
→ The run ends with a reason: complete, stopped, blocked (captcha,
bot check, sign-in), selectors_broken, storage_full, interrupted or
Expand All @@ -76,7 +79,7 @@ User clicks "Start Scraping"
```
User types question in floating widget
→ chatbot.js sends CHAT_MESSAGE (question + last few turns) to service-worker.js
→ Service worker reads the user's key and the current run from storage
→ Service worker reads the user's key and the latest run from its storage
→ scripts/lib/chat.js builds the Gemini request (model id is GEMINI_MODEL there)
→ Response displayed in chat bubble as plain text
```
Expand Down Expand Up @@ -173,10 +176,12 @@ The Jest suite includes a golden corpus of saved Amazon pages (`tests/pages/`, s

## Chrome APIs Used

- `chrome.storage.local` -- Persistent state and product data
- `chrome.runtime.sendMessage` / `onMessage` -- Inter-script messaging
- `chrome.storage.local` -- Settings and the schema version only
- `chrome.storage.session` -- The live run record, written by the service worker
- IndexedDB (the extension's own origin, no permission) -- Runs, products, pages, lastValues and the sync outbox. Starting a run keeps the 10 newest runs and removes older ones, except runs still waiting to sync.
- `chrome.runtime.sendMessage` / `onMessage` -- Messages, all listed in `scripts/lib/messages.js`
- `chrome.downloads` -- File export downloads
- `chrome.tabs` -- Active tab communication
- `chrome.tabs` -- Messages to the run's tab, `tabs.update` for the next page, and `onRemoved` / `onUpdated` to notice the tab going away (none of these need the `tabs` permission)

## Project Structure

Expand All @@ -189,17 +194,21 @@ AmazonSellerScraper/
│ └── popup.js # UI state management and export handling
├── scripts/
│ ├── content/
│ │ ├── scraper.js # Scrape loop: storage, messages, pagination
│ │ ├── scraper.js # Parses a search page and reports it to the worker
│ │ ├── chatbot.js # Floating AI chatbot (Shadow DOM)
│ │ └── offer-fetcher.js # Seller offer page fetching for spread analysis
│ ├── lib/
│ │ ├── parsers.js # Pure search and offer page parsing
│ │ ├── run.js # The scrape run record and its end reasons
│ │ ├── flags.js # Build flags (cloud sync is off in 2.1)
│ │ ├── messages.js # Every message type and who may send it
│ │ ├── run.js # The run state machine and its end reasons
│ │ ├── flags.js # Build flags (cloud sync is off until 2.3)
│ │ ├── migrate.js # Storage schema migrations (schemaVersion)
│ │ └── chat.js # Gemini request builder and run scoping
│ ├── background/
│ │ └── service-worker.js # Message routing + Gemini API
│ │ ├── service-worker.js # Wires the router, the engine, the chat and migrations
│ │ ├── router.js # The one message router
│ │ ├── engine.js # Runs scrapes; the only writer of run data
│ │ └── db.js # IndexedDB stores
│ └── modules/
│ ├── storage.js # Chrome storage abstraction layer
│ ├── analyzer.js # Analytics engine + opportunity scoring
Expand Down
3 changes: 3 additions & 0 deletions jest.config.js
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,9 @@ module.exports = {
roots: ['<rootDir>/tests'],
setupFiles: ['<rootDir>/tests/setup/chrome-mock.js'],
testMatch: ['**/*.test.js'],
// Older builds checked out by the e2e upgrade harness have their own tests.
testPathIgnorePatterns: ['/node_modules/', '<rootDir>/tests/e2e/\\.build/'],
modulePathIgnorePatterns: ['<rootDir>/tests/e2e/\\.build/'],
// sync.js and service-worker.js are authored as ESM (esbuild bundles
// them). A tiny scoped transform rewrites their import/export to
// CommonJS so they can be unit-tested; every other file keeps the
Expand Down
6 changes: 2 additions & 4 deletions manifest.json
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
{
"manifest_version": 3,
"name": "ProScan - Amazon Product Scraper",
"version": "2.1.0",
"version": "2.2.0",
"description": "Scrape Amazon seller products with analytics and AI-powered product insights",
"permissions": [
"storage",
Expand Down Expand Up @@ -30,9 +30,7 @@
"js": [
"scripts/modules/price.js",
"scripts/lib/parsers.js",
"scripts/lib/run.js",
"scripts/lib/flags.js",
"scripts/modules/delta.js",
"scripts/lib/messages.js",
"scripts/content/scraper.js",
"scripts/content/chatbot.js",
"scripts/content/offer-fetcher.js"
Expand Down
11 changes: 11 additions & 0 deletions package-lock.json

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

1 change: 1 addition & 0 deletions package.json
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,7 @@
"@playwright/test": "1.63.0",
"adm-zip": "^0.5.17",
"esbuild": "^0.28.0",
"fake-indexeddb": "^6.2.5",
"jest": "^29.7.0",
"jest-environment-jsdom": "^29.7.0",
"jsdom": "^20.0.3"
Expand Down
2 changes: 2 additions & 0 deletions playwright.config.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,8 @@ import { defineConfig } from '@playwright/test';
export default defineConfig({
testDir: 'tests/e2e',
testMatch: '*.spec.mjs',
// Older builds checked out for the upgrade harness bring their own specs.
testIgnore: '**/.build/**',
globalSetup: './tests/e2e/global-setup.mjs',
workers: 1,
fullyParallel: false,
Expand Down
1 change: 1 addition & 0 deletions popup/popup.html
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,7 @@
<script src="../libs/xlsx.full.min.js"></script>
<script src="../scripts/modules/price.js"></script>
<script src="../scripts/lib/parsers.js"></script>
<script src="../scripts/lib/messages.js"></script>
<script src="../scripts/lib/run.js"></script>
<script src="../scripts/lib/flags.js"></script>
<script src="../scripts/modules/storage.js"></script>
Expand Down
Loading
Loading