V2 media observation - #31
Closed
Manni-MinM wants to merge 154 commits into
Closed
Manni-MinM wants to merge 154 commits into
Manni-MinM wants to merge 154 commits into
Conversation
Add desktop application automation (macOS) via Accessibility API + host bridge
Add verified two-person Google Meet workflow
…ap-integration Jenil prajapati/backend swap integration
…esearch docs - core/: compiled-workflow schema as pydantic models — states carry trigger conditions, transitions carry one atomic action (closed 7-op set), locator chains replayed by whitelist reflection; artifacts load from JSON or YAML - executor/: control-sequence engine — recognize state, dispatch edge action, await target conditions, record per-edge outcome/latency; fails loudly - browser/: minimal Playwright session (dispatch + polling trigger engine, typed errors); import boundaries (core←browser←executor←agent) enforced by test - cli/: Typer app (run/generate/eval/doctor); `netgent run` executes artifacts end-to-end; doctor checks env/keys/browser/credentials - tests/: unit + env-gated integration split (BQT style); 20 passing incl. real-Chromium e2e fixture replay - docs/: OVERVIEW, browser-layer design conclusion, 32-repo survey, 16 research files (incl. team design record ported from snl5) - also: repo restructure moving v1 into v1/ (pre-staged) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ors/logger refactor
Browser layer:
- browser/dom/: StealthProfile (webdriver/plugins/webgl/headers hardening, no
patched binary, CAPTCHA out of scope) + DomSnapshot (shadow-DOM-piercing
interactive-element observer with candidate selectors). Stealth on by default.
Trajectory (the viewable agent run):
- EdgeRecord enriched with per-edge screenshots + which state conditions held;
executor writes a trajectory bundle (record.json + screenshots/) on --trajectory;
`netgent trajectory` renders a text timeline or a self-contained HTML page.
Eval harness:
- netgent evalharness + `netgent eval`: serve a dataset dir locally, substitute
{base}, run each *.workflow.yaml, success = reached accepting state. Committed
forms dataset (vanilla / shadow-DOM / multi-step) all pass; raw results in
evals/results/ (screenshots gitignored).
Structure/refactor (earlier in the wave):
- schema/ package now holds all pydantic models (moved from core/); core/ keeps
errors.py (typed taxonomy) + logger.py (secret-redacting). Workflow gains
`version`. go_back + hover actions added (9-op closed set). netgent schema
command (on-demand JSON Schema, nothing committed). Import boundaries
(schema<-browser<-executor<-agent) enforced by test.
Docs: OVERVIEW.md, stealth-browser.md, long-horizon-agents.md, langchain-evals.md.
35 tests (unit + env-gated real-Chromium integration), ruff clean.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Browser agent (src/netgent/agent/):
- BrowserAgent: LLM-driven loop over a stealth BrowserSession — snapshot the
interactive DOM, ask the LLM for one atomic action (element-indexed decision),
resolve to a durable locator, dispatch, record a trajectory step. Long-horizon
safety: step cap + repeat-action loop detection. CAPTCHA out of scope (prompt
instructs stop; nothing attempts a challenge).
- LLM seam: one decide() method; LangChainLLM (structured output, langchain
imported lazily = generate extra) + FakeLLM for tests. `make_llm("provider/model")`.
- `netgent agent TASK --url ... --model ... --trajectory ...`.
- Full loop tested end-to-end against a real form fixture with a scripted FakeLLM
(no API key): fills fields by observed index, submits, done; plus CAPTCHA-stop
and stuck-loop guard tests.
Control program (schema/control.py) — long-horizon workflow structure:
- EdgeStep/Repeat/Branch/Call union + Param + Milestone. Workflow gains
params/control/accept_states/milestones; control_sequence kept (deprecated).
Executor interprets the program (loops with until/count/max_iterations,
guard-dispatched branches); Call deferred (typed error). Param ${name}
substitution via resolve_params + `netgent run --param`.
38 tests, ruff clean.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Real-model testing (claude-haiku-4-5 against a form fixture) surfaced two DOM bugs:
- <select> accessible name is an option dump ("Plan choose Pro Free"), so
get_by_role(name=...) can't match. Prefer a stable selector for selects, and
collapse whitespace / drop nested-control text in accName so names are clean.
The agent otherwise completed the multi-field signup end to end.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- core/settings.py: Settings(BaseSettings) is the config source of truth for the .env.example surface (provider keys, generator/secondary model, browser, credentials, log level). GEMINI_API_KEY falls back to GOOGLE_API_KEY via an alias. Real env vars win over .env (pydantic-settings default precedence). - sync_provider_keys() publishes keys under the names LLM SDKs read (gemini → GOOGLE_API_KEY for langchain), so `netgent agent` picks up .env keys with no shell sourcing. - CLI root loads settings + syncs keys + sets log level; `netgent doctor` and the agent now see .env automatically. `netgent agent` defaults its model to NETGENT_GENERATOR_MODEL. 44 tests, ruff clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…-tests Running the agent against browser-use/stress-tests challenges surfaced real gaps; each fix is a genuine capability improvement: - Observation now includes salient VISIBLE TEXT (headings, status/alert messages), so the agent can read success/confirmation — not just interactive elements. This alone unblocked the checkbox/radio/form challenges (agent couldn't see "completed"). - Elements show input type (input[date]/[file]/[email]), checkbox/radio [checked] state, and [disabled] — fixing a date-fill loop (needs YYYY-MM-DD) and letting the agent skip disabled/already-set controls. - Correct ARIA roles for inputs by type (radio/checkbox/email/number/...), fixing locator resolution that mapped every <input> to 'textbox'. - New set_checked action (Playwright .set_checked: label-aware, idempotent) + agent check/uncheck decisions — reliable for custom/hidden checkboxes and radios. Verified against real challenges with claude-haiku-4-5: checkbox, radio-selection, dropdown-selections, and a full multi-field signup all complete; hidden-labels fills all 5 field types and is blocked only by a required file upload (upload_file is the known remaining gap). 47 tests, ruff clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A headed run surfaced the last select gap: the agent couldn't see a dropdown's
valid option values, so it guessed ("United States") and select_option failed.
The snapshot now includes each <select>'s option values and the observation shows
them (options=[USA, UK, Canada]), so the agent picks a valid value first try.
Verified headful end-to-end: fill name + email, open/select the country dropdown,
check terms via set_checked, submit, read the success sentinel — 6 steps. 47 tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Two features the stress-tests' challenge.html / forms-comparison.html need: - upload_file action (Playwright .set_input_files) + agent kind="upload". The agent offers a sample file to any file input autonomously (override via BrowserAgent(upload_file=...)). Verified: file-upload challenge passes in 1 step. - The DOM snapshot now descends into same-origin iframes (cross-origin skipped), recording a frame-selector path per element; locators prepend a frame_locator chain so the executor pierces frames. Verified: forms-comparison.html exposes 186 elements across 23 iframes; the agent filled 4 fields inside a nested-iframe form end-to-end. Known: contenteditable fields (no .value) can cause a re-fill loop; challenge.html is an aggregate of ~20 independent mini-challenges, not a sequential flow. 49 tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The snapshot found shadow-DOM elements, but locator resolution picked get_by_role(role) from a placeholder-only field — Playwright's role-name matching ignores placeholder, so it matched nothing and fills timed out. Reorder _locator_for: prefer a simple #id css selector (precise, and Playwright's css engine pierces open shadow roots) over role; use role only WITH a genuine accessible name; skip nameless role locators entirely. Verified against browser-use stress-tests shadow-dom-form: email fill, country select, date fill, and file upload all now work inside the shadow DOM (previously the first fill timed out). 49 tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
frameSelector emitted iframe:nth-of-type(n), which is sibling-scoped — on pages where each iframe sits in its own container (forms-comparison.html: 23 iframes, each alone in a div) it matched ALL 23, so frame_locator was ambiguous and every in-frame action failed. Use a document-unique cssPath instead (id/name when present). This matches how Notte/WebArena handle frames (a frame_locator chain over unique CSS paths per iframe level). Verified: the frame selector now resolves to exactly 1 iframe, and the agent fills a full form inside an iframe (email, select, date, radio, file upload). 49 tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Native HTML5 validation blocks form submit with a tooltip that isn't in the DOM, so the agent would click Submit forever without knowing which field was empty. The snapshot now reports each field's required + validity state, shown as [required] / [invalid: still needs a valid value], and the prompt tells the agent to fix those before retrying Submit. Verified on forms-comparison.html: the first iframe form (email, select, date, radio, file upload) now fills and submits in one pass — the agent's reasoning cites the [invalid] markers to know what to fill. 50 tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Correction to an earlier limitation: Playwright's action layer already pierces cross-origin frames (it drives the browser over CDP, bypassing same-origin policy). The gap was only in observation — the snapshot walked iframe.contentDocument from the TOP frame's injected JS, which same-origin policy blocks. Fix: stop descending iframes in-JS; instead iterate page.frames and evaluate the DOM walk inside EACH frame's own context (Playwright runs it there via CDP, cross-origin included). The frame_locator path per frame is computed from Playwright's frame tree (frame_element in the parent frame), so it works regardless of origin. New test test_cross_origin_iframe: a child served on a different port (distinct origin) embedded in a parent — the input is both observed and filled through the frame_locator chain. 50 tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Long pages (forms-comparison.html: 23 forms/212 elements) exceeded the flat 80-element cap, so the agent never saw forms past the first few and scroll did nothing. Three coordinated changes fix long-page navigation: - Viewport-aware observation: element bbox.y is normalized to top-viewport coordinates (across frames), and the observation shows only the near-viewport slice with a POSITION line (top/middle/bottom) + "N above / N below" counts. Scrolling now shifts which slice is shown, so it makes real progress. - Scroll action adopts browser-use's model: ScrollAction(down: bool, pages: float) converted to pixels from the live viewport height, replacing raw delta_y. No pixel guessing; direction is explicit (the old raw-pixels design let the model try to "scroll to top" with a negative amount = a no-op loop). - Stuck detection is now observation-based: an action that changes nothing on screen counts as no-progress; a scroll that reveals a new batch does not. Verified on forms-comparison.html: the agent pages through the whole document, working ~13 forms deep (element 186/212) in one run vs stalling immediately before. 50 tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…obustness)
Addresses the two ways a long run stalled:
1. The agent never saw its own action errors — the history line dropped them, so it
repeated a failing action until the no-change detector killed the run. History now
includes "-> FAILED: <error>" so the agent can recover.
2. Silent wrong-action no-ops. to_action now guards element/action-type mismatches with
corrective messages: select on a non-<select> ("use fill, dates YYYY-MM-DD"), fill on a
dropdown ("use select"), check on a non-checkbox/radio, and a select value not in the
element's options (lists them). These surface via the feedback channel above.
3. Survey-scrolling. The agent would scroll the whole page "to understand the layout"
instead of acting. A firm prompt plus a scroll guard (reject scroll-down while
[required]/[invalid] fields remain in view, naming them) forces it to fill before paging.
Result on forms-comparison.html: no scroll-survey stall and no wrong-action loop — the run
goes from 8 steps (all scrolling) to 78 steps working forms across the page with 2 scrolls.
Remaining limitation is honest self-assessment: the agent still over-claims "all submitted"
after ~a handful of forms, which is why the compiled-NFA design uses explicit accept
conditions rather than agent self-report. 51 tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…a time "Do all the forms" reliably isn't one long free-form run (which over-claims done); it's deterministic orchestration + LLM per unit + verified outcomes — NetGent's thesis. - BrowserAgent.run gains frame_filter: scope the observation to one iframe so the agent works a single form at a time (DomSnapshot.scoped_to). TextBlock carries its frame_path so success can be checked per form. - agent/sweep.py: enumerate the form frames, run a fresh agent scoped to each, then VERIFY submission by a success marker in that form's own text (not self-report). Reports per-form submitted vs agent-claimed — catching the gap in both directions. - `netgent forms-sweep URL`. Live on forms-comparison.html: 14/24 forms verified submitted, each attempted and checked independently (vs a single run that submitted ~a handful then declared all done). The 10 misses are harder forms that hit the per-form step cap or use trickier widgets — improving them is per-form robustness, not orchestration. 52 tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…e reasons Diagnosing why the sweep hit 14/24: the snapshot treated ANY element with a role as interactive, so container roles (radiogroup, group, list, tablist) were listed and the agent wasted steps clicking a radiogroup wrapper until it timed out. isInteractive now uses an INTERACTIVE_ROLES whitelist (button/link/checkbox/radio/textbox/... — not containers). Also: FormResult now records stopped_reason + last_error, so a sweep says WHY each form failed. Evidence from the 10 misses: ~5 custom framework radios that don't toggle on a synthetic click (the hard widget class), 1 radiogroup-container click (fixed here), and 3 complex forms that hit the per-form step cap. 52 tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Drop the check/uncheck decision kinds (13 -> 11). The agent just clicks a checkbox/radio; to_action routes it to the robust set_checked path — clicking a checkbox toggles it (checked = not current), clicking a radio selects it. This is browser-use's model (one click verb, checkbox smarts underneath), while the schema keeps SetCheckedAction as the replay-safe primitive the executor dispatches. 52 tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…atch The checkbox/radio smarts now live in the click dispatch (session._click), keyed on the live element: a checkbox toggles (set_checked to the opposite of its current state), a radio selects (check) — both via Playwright's label-aware, verified path — and everything else is a plain click. So there's one ClickAction, no SetCheckedAction. Schema action set 11 -> 10; the agent already had one click verb. Integration test proves click toggles a checkbox both ways, selects a radio, and clicks a button. Note: clicking a checkbox is now a toggle (reads live state), so a compiled workflow that clicks a checkbox is only deterministic if the page's initial checkbox state is stable — an accepted trade for the simpler single-action model. 53 tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Push toward completing more forms without changing the action set: - session._click: after set_checked/check on a checkbox/radio, verify the state actually changed; if not (a custom control whose real <input> is hidden behind a styled label), click the associated label in the element's own context via JS — which fires the framework's own listener. This is the biggest failure bucket on the stress-test forms (Angular/custom radios). browser-use uses the same fallback. - sweep_forms gains retries (default 2): each form is attempted with a fresh agent and a larger budget each try, stopping as soon as a success marker is verified — absorbing LLM run-to-run variance. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…it button The enumeration counted any frame with a button as a form, so the top document (page chrome around the embedded form iframes) became "form 1" — the agent then clicked page headings forever. Require an input/select/textarea AND a button; excludes the top frame (24 -> 21 real forms on forms-comparison). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…rash) The viewport paging added for whole-page navigation regressed single-form filling: scoped to one tall form, fields page out of view while their label text remains, so the agent thinks an input is missing and scroll-thrashes for one it already filled. scoped_to now zeroes viewport_height, showing the whole (bounded) form at once. Diagnosed from a clean sweep: 12/21 verified (~57%), down from ~67% on real forms before paging — this reverts that regression for the scoped case. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ur's banner verified two broken fixtures) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ic oracles, feedback contracts; NetGent verifier design Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…on, replay determinism, feedback forms Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… wired as an orchestrator node; repeated-action guard agent/verifier: Verdict + Evidence (task, params, action log, final observation, texts seen, this run's dialogs, final URL, last 3 screenshots; never the explorer's reasoning — the judge failure modes are sycophancy toward it), judge_trajectory via a new LLM.judge() seam method. Orchestrator: explore → verify → generate; 'not achieved' re-explores once with the unmet points appended; the replay still gates. `netgent generate --judge/--no-judge`; NETGENT_JUDGE=1 on `eval stress sweep` scores the judge against page truth. Measured (21-form sweep, Haiku 4.5): precision 100%, recall 100% after two fixes the measurement forced — an outcome-not-fields rule, and dialogs scoped to the run. Explorer: same-action-repeated guard (nudge at 3, stop at 6) — measured 12 identical clicks on a YouTube overlay that observation-equality never caught. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…d-mirrored copy of the loop's structure) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…--graph` Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…iled graph; explorer refactor sketch Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…s and the settle watcher move out of Agent The one class the LangGraph refactor keeps (langgraph-agent-structure.md §5.6 step 1): it owns an asyncio.Task with a lifecycle, so it can live neither in graph state nor in a checkpoint. Agent delegates to it for now; note/drain_noticed/start_watch/stop_watch are verbatim. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…frozen run context instead of closed-over names Step 2 of langgraph-agent-structure.md §5.6: build_agent_graph still compiles per run, but it now builds an ExplorerContext (session, llm, memory, task, knobs — with Agent.__init__'s validation) and every node reads ctx.*; upload_path/capture_screenshot become module functions taking the context. No behaviour change; the Runtime plumbing comes next. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…(), module-level EXPLORER, explore() Step 3 of langgraph-agent-structure.md §5.6. observe/decide/act are module-level async functions taking (state, runtime: Runtime[ExplorerContext]); the graph is compiled once at import (StateGraph(AgentState, context_schema=ExplorerContext).compile(name="explorer")) and a run passes its live session/LLM/memory as context= — never checkpointed. explore(...) is the single run API (Agent.run's body, minus the per-run rebuild); Agent.run delegates to it. langgraph is imported at module level in graph.py only; netgent.agent still loads without it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…graph; models → explorer/models.py; every call site on explore()/ExplorerAgent Step 4 of langgraph-agent-structure.md §5.6, amended: the class is not deleted but reduced to a thin façade — ExplorerAgent(llm, *, max_steps, run_dir, allowed_kinds, max_actions_per_step, upload_file, memory) holds the knobs and ONE ExplorerMemory, and run() delegates to graph.explore(); no loop logic. StepRecord/AgentStep/AgentTrajectory move to models.py (MAX_REPEAT to graph.py, the fold constants to memory.py). explorer/__init__ and agent/__init__ re-export ExplorerAgent/ExplorerMemory/ExplorerContext + models, and resolve explore/ create_explorer_agent/EXPLORER lazily (PEP 562) so the package still loads without langgraph. Call sites: orchestrator → explore(); CLI, sweep and stress → ExplorerAgent(...).run(...). Tests follow; one agent test now exercises explore() with an explicit memory. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… invoke, context/façade validation; netgent.agent import boundary Step 5 of langgraph-agent-structure.md §5.6. EXPLORER.get_graph().draw_mermaid() must show observe/decide/act with exactly the loop and its exits — generated from the nodes' Command annotations, the honest replacement for the diagrams dropped in 70a3a3b/0a70be2. A budget-0 run invokes the module-level graph with no browser and no key (context, not closure). test_import_boundaries now also pins that importing netgent.agent loads no langchain/langgraph. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…tic graph nests the explorer explore() takes an optional graph= (default EXPLORER; a checkpointer-compiled one persists) and the orchestrator passes EXPLORER explicitly. LangGraph's subgraph discovery reads the node's source for names it loads (get_function_nonlocals → find_subgraph_pregel), so an import inside the node body — or a call through explore() alone — leaves get_subgraphs() empty; with the name present, get_subgraphs() lists 'explore' and get_graph(xray=True) draws observe → decide → act inside the pipeline (langgraph-agent-structure.md §3d probe C → A, §5.5 item 1). Pinned by test. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…verifier_agent(), module-level VERIFIER, verify() judge.py splits into models.py (Evidence, Verdict), prompt.py (JUDGE_SYSTEM, build_judge_content), context.py (VerifierContext: the LLM and screenshot dir as Runtime.context) and graph.py — two function nodes, gather (trajectory → Evidence, pure) and judge (the one LLM call), compiled once as VERIFIER with verify(traj, task, *, llm, params, run_dir, graph=None) as the single run API. The package resolves the langgraph-importing names lazily, as the explorer does. Orchestrator and sweep call verify(); the orchestrator names VERIFIER in its node so xray nests gather → judge beside observe → decide → act. Existing imports (Evidence, Verdict, build_judge_content, judge_trajectory) keep working. Tests: mermaid snapshot, a no-key run through FakeLLM, context validation, and the pipeline's two subgraphs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…dator removed — pipeline is explore → verify → generate planner/: models.py (Plan, PlanStep), prompt.py (PLANNER_SYSTEM, build_planner_content), context.py (PlannerContext), graph.py (draft node, create_planner_agent(), module-level PLANNER, plan()), agent.py (PlannerAgent). One LLM call through the seam's structured judge(); not yet wired into the orchestrator — docs/OVERVIEW.md leaves how a plan feeds the explorer fleet open. verifier/agent.py: VerifierAgent(llm, *, run_dir, max_screenshots).run(traj, task, params) → verify(). agent/validator/ and its test are deleted; the orchestrator loses the validate node, GenerateRequest .validate_replay and GenerateResult.report/.validated; `netgent generate` loses --validate/--no-validate (replay proof is `netgent run` on the artifact). Tests for both new packages; CLAUDE.md updated. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…e breakpoints, injectable chat model Provider resolution is LangChain's: `provider:model` (a `/` is rewritten to `:`), so _PROVIDER_ALIAS and the partition logic go; the Anthropic temperature quirk stays (2 lines). Default and all call sites/docs use `google_genai:gemini-…` / `anthropic:claude-…`; doctor and settings accept the new prefix. NOTE: the old `gemini/…` spelling no longer resolves — LangChain infers Vertex AI for a bare `gemini-*` name. The explicit Anthropic cache_control breakpoint is dropped; messages are a plain SystemMessage(static) + HumanMessage(dynamic). Usage counters still report whatever the provider cached implicitly. LangChainLLM(model: str | BaseChatModel) — a built chat model is used as-is, which is the seam LangChain's unit-testing guide describes: a GenericFakeChatModel now drives the real with_structured_output → parser → retry-ladder path in tests, no key, no network. RetryWithErrorOutputParser (langchain_classic) was evaluated and not adopted: it re-parses a text completion and needs a text parser; tool-calling structured output has no completion to re-parse. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…middleware, same explore()/ExplorerAgent API (A/B arm) agent/explorer_v2/: tools.py (one tool per atomic kind + done; each executes against runtime.context.session, records the AgentStep with its durable locator and the StepRecord, runs under a per-memory lock in call order since ToolNode gathers a turn's calls concurrently), middleware.py (before_model = v1 observe: budget, scoped snapshot, stuck detection, texts_seen; wrap_model_call = v1's prompt layout rebuilt every turn from the cross-run memory, the accumulated messages are never sent; after_model = decide's guards: truncate to max_actions_per_step, intercept `done`, repeated-action nudge/stop, re-observe on a text reply), state.py, prompt.py (v1's prompt with the fields restated as tool calls), graph.py (create_explorer_agent(model, allowed_kinds, max_actions) → create_agent(...), cached per tool set; explore() with v1's signature; usage via LangChain's UsageMetadataCallbackHandler), agent.py (ExplorerAgent façade). Reuses v1's ExplorerContext/ExplorerMemory/models; ExplorerContext.llm is now typed Any so pydantic can build the tools' arg schemas. NETGENT_EXPLORER=v2 switches the sweep/stress arm; SweepResult.usage carries v2's token totals. Known gaps vs v1: no settle watcher, no structured-output retry ladder (a no-tool-call turn costs the step), no module-level graph. Browser tests mirror test_agent.py with GenericFakeChatModel scripting tool calls. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…on decisions/steps, no PARAMETERS prompt section, no planner params
The explorer no longer declares which ${name} a value came from: AgentAction/AgentDecision lose
`param`, AgentStep loses `param`, the prompt loses its PARAMETERS section and field, explorer_v2's
tools lose the `param` argument. The planner loses its `params` input (state, prompt, plan(),
PlannerAgent.run). The compiler's structural binding pass goes with it; `-p name=sample` still
works through the case-insensitive literal sweep over value fields and state conditions (never
locators), and the orchestrator still tells the explorer the sample values to use verbatim.
Tests: test_param_binding rewritten for sweep-only semantics; prompt/planner/compiler tests adjusted.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ation re-attributed), induction literature, typed-key merge proposal Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…me-family tasks with proposed param values, normalized in code
One LLM call (VARIATION_PLANNER, the planner's graph shape) returns [{task_text, values}];
normalize_variation_plan (pure) forces variation 1 to the base task verbatim, makes run-1's
value names canonical, applies --variation pins to variation 2, and pads/truncates to N.
The names are hypotheses only — the merge confirms them as Params when values vary in the
trajectories.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…der <name>.trajectories run-<k>/ holds trajectory.json + screenshots (explore writes them), variation.json and verdict.json (achieved flag, attempts); a retried run's first attempt is stashed, not lost; generalized.json at the root is where the merge's induced memory lands. Failed runs are stored too, marked — they are memory. Pure file I/O, zero LLM. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…e generalized NFA, pure code
Align achieved runs on (action type, durable target key) (Needleman-Wunsch over typed sigs;
AWM's abstract-trajectory signature with durable locators). Conditions by version-space
intersection: url_matches/anchors/dialogs survive only when they held in EVERY run, with
support counts in the evidence trail. Divergence gets the four dispositions of
trajectory-memory.md §C.1.3:
1. value-varies + matches a planner value -> Param (fill text/goto url substitution,
dwell -> Repeat(count="${watch_time}") of 1s slices behind a noop edge, and role-name
targets containing the value -> get_by_role(role, name="${param}") + nth(0));
2. present-in-k-and-dismissal-shaped -> Interrupt (cross-run presence primary, text
heuristics tie-break; is_interruption_step lifted to module level in compiler.py);
3. genuine downstream divergence with distinguishing first targets -> Branch, one arm per
continuation, converging on one state;
4. else reject: warning naming the column, spine (run 1) kept so the artifact replays.
Failed runs contribute nothing structural (failures poison trajectory-shaped memory).
accept_states = the final state's intersected conditions, never invented. generalized.json
(GeneralizedTrajectory) is the induced memory. Executor: Repeat.count numeric-string
coercion, since resolve_params substitutes ${p} into count's str type.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ped merge -> zero-LLM replay check New LangGraph (runs=1 keeps the single-run graph untouched): plan (variation planner, one LLM call; -p names join as proposals) -> explore_run xN (fresh ExplorerMemory per run via explore(); per-run verify with ONE private retry whose unmet-points suffix never crosses runs; every attempt persisted in the TrajectoryStore) -> merge (merge_trajectories, pure code; generalized.json written; artifact dumped) -> replay (agent/replay.py: metamorphic zero-LLM check — run-1's and run-2's value sets must walk the same state sequence, interrupt edges excluded, dwell self-loops collapsed; mismatch sets result.error). Independence policy: runs share only a HINTS line of previously seen dismissal anchors. CLI: --runs / --variation name=value. Integration tests: N=2 fixture e2e (params inferred, store layout, both-value replays), private retry suffix, no-achieved-run stop. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…matches a 10.0 dwell, stored as the bare number for Repeat.count Observed live: the variation planner writes durations with units; the count feeds 1s slices and must resolve numerically. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
1. Alignment scoring is locator-shape-aware: same-type substitution scores 2 only when the
target shape agrees (both role=link, both css, ...), else 0 — the shape-blind +1 aligned
run-1's css play-button click with the other runs' video-title links, dropping the real
video click and leaving an unguarded artifact that passed replay vacuously.
2. Disposition 4 minority steps are DROPPED (structural intersection), not kept: the runs
without them achieved the task, and keeping run-1-only steps (suggestion click, play
click) made replay time out on elements that never reappear. Branch synthesis is now
presence-based only (all-gap regions, the login-wall shape) — a substitution column in a
region is value variance, and arms guarded on per-run targets would freeze the values.
After the fixes the stored 3-run YouTube memory re-merges (offline, zero LLM) to:
goto -> fill ${video_query} -> click search -> click video (target-varies, run-1 kept) ->
5s dwell; replay set 1 walks s1..s5 to the real watch page; set 2 fails honestly at the
value-dependent click (titles do not contain the query - the known open gap).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…or shape Measured on the second real YouTube merge: run A clicks the video as a role link, run B as a >> filter(has_text) — shapes differ, so the shape-only rule paired run B's video click with run A's play-button click and the compiled word never opened a video. Each step now carries its URL effect (did the base change, where to); a substitution pairing scores 2 with shape+effect, 1 with either, 0 with neither. With the fix the stored 2-achieved-run memory re-merges offline to a workflow with BOTH params (video_query -> fill text, watch_time -> Repeat.count) and the zero-LLM metamorphic replay PASSES: both value sets walk s1..s5 through the real url_matches(/watch) gate, dwelling 10 vs 5 slices. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…t-aware stuck detection
The observation layer lied to the agent on video pages, three ways at once
(diagnosed on live YouTube runs, runs/youtube-warpigs trajectory as evidence):
- YouTube freezes its control-bar DOM (time display, Play/Pause labels) while
the controls are auto-hidden, and the walker's visible() only checked the
element's OWN computed style — ancestor opacity:0 doesn't propagate into it —
so frozen, invisible controls were reported as live UI. The explorer clicked
pause, was shown the stale pre-click label, and toggled itself into the
repeat-stop (2 runs); the frozen timer made playing ads compare byte-equal
across steps and fire the no-change stuck stop (2 runs).
- Nothing read the <video> element's own properties, the one signal that never
freezes.
Fixes, one per layer:
- snapshot.js: emit media state (currentTime/duration/paused/ended/muted) per
visible-or-audible <video>/<audio>, read-only property access; visible() now
uses native checkVisibility({opacityProperty, visibilityProperty}) so
ancestor-faded controls drop out of the observation instead of lying.
- models/observer/serializer: MediaState -> DomSnapshot.media -> a
"MEDIA: video PLAYING at 0:29 / 7:56" line under TITLE. Its ticking position
also keeps a playing page from ever comparing equal to its previous
observation.
- explorer/graph.py: no_progress increments only when the rendered observation
is unchanged AND no never-seen text appeared — texts_seen was recording ad
captions advancing at the very moment the stuck stop declared "no change on
screen".
Measured effect: the War Pigs intent went from 4 consecutive failed
generations (pause-toggle x2, false-stuck x2) to a clean judge-approved
compile whose zero-LLM replay passes end-to-end; exploration reasoning now
cites playback state instead of guessing it from button labels.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.