Skip to content

Browser: hidden tabs pause and sleep, plugin engines for agents' background tabs, background screenshots; Ubuntu 24.04 bwrap fallback; Windows pipe host restart - #119

Merged
howdeploy merged 7 commits into
howdeploy:mainfrom
BIackFIame:pr/7-browser-engines
Oct 3, 2026

Conversation

@BIackFIame

Copy link
Copy Markdown
Contributor

Branch: pr/7-browser-engines → base main (b322234). 5 commits.

Why

  • A browser tab nobody looks at still cost as much as a visible one: a hidden tab was only throttled, so a few background pages kept spending CPU and hundreds of MB.
  • Agents that read many pages in the background paid for a full Chromium renderer per tab, even for plain text. There was no way for a plugin to offer a lighter engine.
  • An agent's screenshot of a background tab failed with VIEWPORT_UNAVAILABLE whenever the Browser card was off-screen, which is most of the time when an agent works in the background.
  • Ubuntu 24.04 and later block unprivileged user namespaces through AppArmor, so an installed bubblewrap fails to start and isolated launches failed instead of falling back.
  • A Windows pipe host that kept crashing right after start-up was restarted forever.

What changes for users after merging

Hidden browser tabs stop working in the background and wake when they are needed; long-hidden ones sleep with a preview. Agents can take screenshots of their background tabs while the Browser card is off-screen. On Ubuntu 24.04 subagents run in Manual with a clear reason instead of failing, and the docs say how to allow bubblewrap. With an optional engine plugin installed, agents read background pages much more cheaply.

Hidden tabs (macOS, hidden built app, 5 hidden tabs each with a 16 ms timer and a rAF loop, 2 runs; CPU in % of one core over 30 s, memory over the app alone):

5 hidden tabs CPU Memory Agent command reaches the page
Before: throttled only 9.7–13.4 % +340–496 MB immediately
After: frozen (paused, after 30 s) 0.4 % unchanged 13–19 ms
After: sleeping (after 10 min) 0.2 % +20–30 MB 80–87 ms, incl. reload

Lightpanda engine (optional plugin, not part of this PR; macOS, stub agent, live network, one run per page; Lightpanda 0.4.1 vs CanvasTTY's Chromium in a hidden window):

Lightpanda Chromium
Public pages read correctly 17/21 (3 bot walls, 1 heavy JS app) 21/21
CPU for the 21 pages 5.1 s 21.8 s
Median time to content 571 ms 781 ms
Cold start to first page ready 85–130 ms 580–780 ms
5 pages open and idle 91 MB 1,064 MB
Added per page inside CanvasTTY ~13 MB + 25 MB process ~150–160 MB

When the engine cannot do something (screenshots, bot walls, thin text, a missing CDP method, a crash, showing the tab), the tab moves to Chromium under the same tab id and the agent gets a notice, so these failures cost time, not results.

What's inside

  1. Windows pipe host restart (fix(orchestration)): the browser gateway restarts a failed host on the pipe name it already published, so helpers reconnect with their tokens (--pipe-name, accepted only in the host's own generated form). Both gateways keep their three-retry budget while a replacement dies within a minute of starting; a host that served longer earns a fresh budget. Orchestration keeps main's one-run leases.
  2. Ubuntu 24.04 bubblewrap fallback (fix(isolation)): a cached probe (once, again after a minute on failure) runs bubblewrap with a launch's namespaces. If bwrap is installed but cannot create a user namespace, it counts as no isolation layer: agents that need one run in Manual and the card shows what bwrap said and where the docs are. The docs (en/ru/zh-CN) show an AppArmor profile for bwrap alone, or the sysctl.
  3. Hidden tab freeze and sleep (perf(browser)): a tab hidden for 30 s with no agent on it is paused through CDP (Page.setWebLifecycleState) and resumed before any agent command. After 10 minutes, or when more than 6 hidden tabs are alive, the least recently used one sleeps: its WebContents closes, its address, history and a small picture stay, and it reloads on the next command or when shown (the agent result says so). Tabs playing media, downloading, showing a dialog or with a beforeunload handler are left alone. One toggle in Settings, on by default.
  4. Plugin browser engines and background screenshots (feat(browser), fix(browser)): a plugin can contribute a browser:engine; browser_new_tab takes engine (auto, chromium, or an engine id). Engine tabs get no cookies or profile and never show; the person's tabs always use Chromium. Screenshots of background and moved tabs work with the card off-screen: the tab's view is put in the window with only its top-left pixel inside the bottom-right corner, under a one-pixel cover, and taken out right after.

How to verify

npm ci
npm run build:helpers -- --host
npm run typecheck
node --test --test-concurrency=2 tests/*.test.mjs
npm run build
npm run audit:secrets

Local: full suite 1666 tests, 1663 pass, 3 skipped (two history tests need Node ≥ 22.13 for node:sqlite; they pass with --experimental-sqlite on the local Node 23.3). CI on the fork: verify, macos-cli-resolution and windows-pipe-host pass.

Hidden app run (built app, Lightpanda plugin, stub agent, Browser card panned off-screen): 5/5 screenshots of moved engine tabs and Chromium background tabs succeed, in two runs; engine tabs saw no cookie the Chromium session held.

Manual: open a few pages, hide the Browser card, and watch CPU drop; ask an agent for a screenshot of its background tab while the card is off-screen.

Risks

  • A frozen tab cannot answer a native file chooser; CanvasTTY cannot detect one, so a page waiting on it stays frozen until shown.
  • beforeunload is checked in the top frame only before a tab sleeps.
  • The Windows host restart is verified in CI only, not on a Windows machine.
  • The tab shown in the card still needs the card in view for a screenshot, as before.
  • Lightpanda (optional plugin): bot walls (its User-Agent cannot pretend to be Chrome), no real layout (clicks go through the DOM), no faithful screenshots. It is AGPL-3.0, installed by the user from the official release and never shipped with CanvasTTY. Plugin: https://github.com/BIackFIame/canvastty-plugin-lightpanda

…and a crashing one is not restarted forever

The browser gateway restarts a failed Windows pipe host on the pipe name
the first host published, so agent helpers that already hold the address
and their reconnect token come back. The native host takes an optional
--pipe-name argument and accepts only names in the form it generates
itself; the transport passes it and fails if the ready frame names a
different pipe.

Both gateways now keep their three-retry budget while a replacement host
dies within a minute of starting, so a host that crashes right after
start-up is given up on instead of being restarted every few seconds; a
host that served longer than that earns a fresh budget. The orchestration
gateway keeps its one-run leases: a restarted host serves only newly
launched orchestrators.
Background throttling leaves a hidden tab's requestAnimationFrame loop
and one timer wake-up per second running. A tab that stays hidden, with
no agent driving it, for 30 s is now frozen through its existing
debugger attachment (CDP Page.setWebLifecycleState). After 10 minutes,
or when more than 6 hidden tabs are alive, the least recently used one
goes to sleep: its WebContents closes. The tab keeps its address, title,
favicon, back/forward history (with each entry's scroll and form state)
and a small picture for the card.

Showing a tab wakes it, and so does any command for it. A paused tab
resumes before any automation CDP command reaches the page. A sleeping
tab reloads on its history entry, and the agent result says "Browser tab
was reloaded" (details.tabReloaded on errors). Closing or reloading a
sleeping tab does not wake it first. A tab is never paused while it
plays media, downloads, has a dialog open, or is being captured or
inspected. It is never put to sleep while an agent is on it or while
the page has a beforeunload handler. An unanswered check counts as
"has one". A blocked tab is checked again later and never forced.

Settings: one toggle, "Pause hidden browser tabs" (on by default). The
card marks tabs as "paused" or "sleeping" and shows a sleeping tab's
last picture. Agents never receive that picture.

Measured on the hidden built app with a fake HOME, 2 runs: 5 hidden
tabs, each with a 16 ms timer and a rAF loop. CPU over 30 s: throttling
only 13.4 / 9.7 % of a core, paused 0.4 / 0.4 %, sleeping 0.2 / 0.2 %
(the app alone uses 0.1 %). RSS: throttling only and paused +496 / +340
MB over the app alone; sleeping +30 / +20 MB. A command reaches a paused
tab in 13-19 ms and a sleeping one in 80-87 ms (reloaded, with the
notice).

Electron's view.webContents getter returns undefined once a WebContents
has closed. Each tab's view now pins its WebContents object, so a closed
tab still answers isDestroyed(). This also covers destroyed tabs, which
already had the problem.
A tab is now driven through a small TabDriver: the Electron WebContents
(ElectronTabDriver, unchanged behaviour) or a local CDP WebSocket
(CdpTabDriver, one connection per tab, Target.createTarget plus a flat
session, loopback ws:// endpoints only).

A plugin service can declare `browserEngine` (permission
`browser:engine`). The host calls canvastty.browserEngine.openTab
{ engineId, tabId } for a tab and gets { webSocketUrl };
canvastty.browserEngine.closeTab is a notification. The core keeps the
policy (BrowserEngineTabs):
- only an agent's new tab may use an engine; the person's tabs and
  engine "chromium" always use Chromium;
- an engine tab is never shown, gets no cookies, profile or
  credentials, and counts toward the tab limit.

browser_new_tab takes `engine` ("auto" by default, "chromium", or an
engine id). An engine tab opens in the background and the active tab
stays; commands without a tab id go to the tab the agent opened last.
Engines without layout get observation without geometry and clicks
and hovers through element.click() / DOM events.

The tab moves to Chromium under the same id (next revision, refs
stale, a notice in the result, details.movedToChromium on errors) for:
- screenshots, drags and downloads;
- bot walls (challenge title or text, a /cdn-cgi/challenge-platform/
  request, a 403/429/503 document);
- text too thin for the page's size;
- an unsupported CDP method (-32601);
- an engine disconnect or crash;
- being shown.

A site that failed this way goes straight to Chromium for the rest of
the session, and an engine that cannot open a tab is skipped by auto
for a minute. A moved tab is drawn under the active tab for a
screenshot, so it can be captured without being shown.

Docs: manifest schema, plugin API types and the plugin, browser and
changelog pages in three languages, plus the agent browser skill.
Tests use a fake CDP engine behind a real loopback WebSocket.
… as no isolation layer

Ubuntu 24.04 and later set kernel.apparmor_restrict_unprivileged_userns=1,
so an installed bwrap fails to start unless AppArmor allows it. Before a
Linux launch CanvasTTY now runs bubblewrap once with the namespaces a
launch uses. A working result is kept; a failure is checked again after a
minute, so allowing it takes effect without a restart. When it fails, the
launch is treated exactly like a computer without an isolation layer:
subagents and plugin-started agents run in Manual and the card says what
bubblewrap reported and where the docs explain how to allow it.

The docs (en, ru, zh-CN) show both ways: an AppArmor profile for bwrap
alone, or the sysctl for every program. Tests use a fake probe for the
decision and a fake bwrap for the probe itself.
…ard is off-screen

A screenshot of an agent's background tab, or of a tab that moved to
Chromium from a contributed engine, failed with VIEWPORT_UNAVAILABLE
whenever the Browser card was off-screen: Chromium paints a view only while
some of it lies inside a shown window, and the old lend put the tab under
the active tab, which was hidden too.

Now the tab's view, at its page size, is put in the window for the capture
with only its top-left pixel inside the window's bottom-right corner, under
a one-pixel cover in the window's background color, and taken out right
after; the capture asks Chromium to paint it without making the page
visible. The clip view, the canvas and the person's tabs are not touched.
The tab shown in the card still needs the card in view, as before.

Checked in a hidden app run (macOS, built app, Lightpanda plugin, stub
agent, card panned off-screen): 5/5 screenshots of moved engine tabs and
Chromium background tabs succeed, in two runs; engine tabs still see no
cookies the person's Chromium session holds.
…k processes exit

The hook processes run in the plugin folder and can still be exiting when
the test ends; Windows then refuses the folder removal with EBUSY.

Copy link
Copy Markdown
Owner

Reviewed head bde3877f7eb89793f047e6b4cfaad06bf475f0de against current main (b3222349235257680bb03ddb6ff972be7ae54ff7). All three PR checks are green. I found two failure-path issues that should be fixed before merging.

P2 — Failed lifecycle transitions can retry continuously without backoff.

In BrowserTabLifecycle.ts:253–271, a rejection from host.freeze() leaves the entry active and does not advance freezeNotBefore. The finally block calls schedule(), which computes its delay from the same expired deadline and schedules another attempt with zero delay. The outer catch only logs the failure.

This is reachable through the real host: BrowserService.freezeTab() awaits BrowserAutomationService.setLifecycleState(), whose debugger attachment / Page.setWebLifecycleState command can reject. A persistent debugger failure therefore turns the idle-tab optimization into a rapid retry and warning loop. An exception in a due discard has the same missing-backoff problem.

Please advance the appropriate retry deadline when a transition rejects, while preserving the actual lifecycle state. Ensure all already-expired deadlines are handled so advancing only the freeze deadline cannot leave the discard deadline scheduling immediate attempts.

A useful regression should make freeze reject repeatedly after its deadline, check that the next attempt is delayed by a bounded retry interval, and then allow a later attempt to succeed. Cover a rejected discard as well. The current lifecycle tests cover blockers and successful transitions, but not these rejected-transition paths.

P2 — A connection that completes after the open timeout leaks its socket.

In CdpTabDriver.ts:132–159, the timeout may win while connect() is still pending. The catch calls driver.close() before there is a socket, marking the driver destroyed. If the connection resolves later, its socket is assigned to driver.socket, then the destroyed check throws. The late rejection is swallowed by opening.catch(); nobody closes the newly acquired socket. The engine tab has already failed to open, so normal tab disposal cannot clean up that driver.

Please explicitly close a socket acquired after cancellation/timeout, or make the connection attempt cancellable with ownership of cleanup. Add a delayed-connection regression: let the open timeout reject first, then resolve the socket factory and assert that the returned socket is closed and no target is created.

The broader changes have useful safeguards: native-code/service trust for contributed engines, loopback endpoint validation, no browser-profile forwarding, stale references after fallback, and bounded Windows host restart attempts. These do not remove the two resource/retry issues above.

This feedback is from source review and inspection of the existing tests and GitHub CI; I did not run local tests or an Electron/UI session.

Copy link
Copy Markdown
Owner

The two P2 findings in my previous review still apply to head bde3877f7eb89793f047e6b4cfaad06bf475f0de. Please fix both before we merge this PR:

  1. Add bounded retry/backoff for rejected freeze/discard transitions, accounting for all expired deadlines so the scheduler cannot fall into immediate repeated attempts. Preserve the actual lifecycle state and add regressions for persistent failure followed by recovery.
  2. Close a CDP socket that resolves after the open timeout/cancellation. Add a regression where connection completion happens after timeout, and assert that the late socket is closed without creating a target.

Please push these fixes and rerun CI. The earlier comment links the exact code paths and explains both failure sequences. We are proceeding with #120 and #121 separately; #119 remains pending these corrections.

…eploy#119)

Back off failed freeze/discard transitions across both deadlines, including
hidden-limit discards, while retaining the completed lifecycle state.
Close sockets acquired after an engine open has already timed out.
Add regressions for repeated failures, recovery, and late connection ownership.

Note: an existing browser-engine click revision assertion failed once in the
full suite; the focused browser rerun passed all 32 tests.
@BIackFIame

Copy link
Copy Markdown
Contributor Author

Failed lifecycle transitions can retry continuously without backoff.

A connection that completes after the open timeout leaks its socket.

Addressed both findings and the follow-up request in 684cd3a.

  • Rejected freeze/discard transitions now defer both retry deadlines before rescheduling, preserving the last completed lifecycle state. The same backoff applies to discards queued by the hidden-tab limit.
  • A CDP socket acquired after timeout or cancellation is closed before the driver takes ownership or creates a target.
  • Added six regressions covering repeated rejection followed by recovery, deadlines expiring during an in-flight discard, hidden-limit retries, and delayed socket acquisition. The lifecycle and late-socket regressions fail against the original implementation.

Validation: all 32 browser tests and typecheck pass. The full run reported 1666 passed, 3 skipped and 3 failures: two suites required the experimental SQLite flag on this local Node 23.3 runtime, and an existing click/revision assertion failed once. The SQLite suites then passed all 31 tests with the flag; the full browser group passed on rerun, and the focused click test passed five additional runs plus a clean original-head snapshot. CI starts again on the updated PR head.

@howdeploy
howdeploy merged commit 4472c21 into howdeploy:main Oct 3, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants