Skip to content

V2 media observation - #31

Closed
Manni-MinM wants to merge 154 commits into
SNL-UCSB:mainfrom
jaber-the-great:v2-media-observation
Closed

Manni-MinM wants to merge 154 commits into
SNL-UCSB:mainfrom
jaber-the-great:v2-media-observation

Conversation

@Manni-MinM

Copy link
Copy Markdown
Member

No description provided.

sthanavm and others added 30 commits July 20, 2026 09:45
Add desktop application automation (macOS) via Accessibility API + host bridge
Add verified two-person Google Meet workflow
…ap-integration

Jenil prajapati/backend swap integration
…esearch docs

- core/: compiled-workflow schema as pydantic models — states carry trigger
  conditions, transitions carry one atomic action (closed 7-op set), locator
  chains replayed by whitelist reflection; artifacts load from JSON or YAML
- executor/: control-sequence engine — recognize state, dispatch edge action,
  await target conditions, record per-edge outcome/latency; fails loudly
- browser/: minimal Playwright session (dispatch + polling trigger engine,
  typed errors); import boundaries (core←browser←executor←agent) enforced by test
- cli/: Typer app (run/generate/eval/doctor); `netgent run` executes artifacts
  end-to-end; doctor checks env/keys/browser/credentials
- tests/: unit + env-gated integration split (BQT style); 20 passing incl.
  real-Chromium e2e fixture replay
- docs/: OVERVIEW, browser-layer design conclusion, 32-repo survey, 16 research
  files (incl. team design record ported from snl5)
- also: repo restructure moving v1 into v1/ (pre-staged)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ors/logger refactor

Browser layer:
- browser/dom/: StealthProfile (webdriver/plugins/webgl/headers hardening, no
  patched binary, CAPTCHA out of scope) + DomSnapshot (shadow-DOM-piercing
  interactive-element observer with candidate selectors). Stealth on by default.

Trajectory (the viewable agent run):
- EdgeRecord enriched with per-edge screenshots + which state conditions held;
  executor writes a trajectory bundle (record.json + screenshots/) on --trajectory;
  `netgent trajectory` renders a text timeline or a self-contained HTML page.

Eval harness:
- netgent evalharness + `netgent eval`: serve a dataset dir locally, substitute
  {base}, run each *.workflow.yaml, success = reached accepting state. Committed
  forms dataset (vanilla / shadow-DOM / multi-step) all pass; raw results in
  evals/results/ (screenshots gitignored).

Structure/refactor (earlier in the wave):
- schema/ package now holds all pydantic models (moved from core/); core/ keeps
  errors.py (typed taxonomy) + logger.py (secret-redacting). Workflow gains
  `version`. go_back + hover actions added (9-op closed set). netgent schema
  command (on-demand JSON Schema, nothing committed). Import boundaries
  (schema<-browser<-executor<-agent) enforced by test.

Docs: OVERVIEW.md, stealth-browser.md, long-horizon-agents.md, langchain-evals.md.
35 tests (unit + env-gated real-Chromium integration), ruff clean.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Browser agent (src/netgent/agent/):
- BrowserAgent: LLM-driven loop over a stealth BrowserSession — snapshot the
  interactive DOM, ask the LLM for one atomic action (element-indexed decision),
  resolve to a durable locator, dispatch, record a trajectory step. Long-horizon
  safety: step cap + repeat-action loop detection. CAPTCHA out of scope (prompt
  instructs stop; nothing attempts a challenge).
- LLM seam: one decide() method; LangChainLLM (structured output, langchain
  imported lazily = generate extra) + FakeLLM for tests. `make_llm("provider/model")`.
- `netgent agent TASK --url ... --model ... --trajectory ...`.
- Full loop tested end-to-end against a real form fixture with a scripted FakeLLM
  (no API key): fills fields by observed index, submits, done; plus CAPTCHA-stop
  and stuck-loop guard tests.

Control program (schema/control.py) — long-horizon workflow structure:
- EdgeStep/Repeat/Branch/Call union + Param + Milestone. Workflow gains
  params/control/accept_states/milestones; control_sequence kept (deprecated).
  Executor interprets the program (loops with until/count/max_iterations,
  guard-dispatched branches); Call deferred (typed error). Param ${name}
  substitution via resolve_params + `netgent run --param`.

38 tests, ruff clean.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Real-model testing (claude-haiku-4-5 against a form fixture) surfaced two DOM bugs:
- <select> accessible name is an option dump ("Plan choose Pro Free"), so
  get_by_role(name=...) can't match. Prefer a stable selector for selects, and
  collapse whitespace / drop nested-control text in accName so names are clean.
The agent otherwise completed the multi-field signup end to end.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- core/settings.py: Settings(BaseSettings) is the config source of truth for the
  .env.example surface (provider keys, generator/secondary model, browser,
  credentials, log level). GEMINI_API_KEY falls back to GOOGLE_API_KEY via an
  alias. Real env vars win over .env (pydantic-settings default precedence).
- sync_provider_keys() publishes keys under the names LLM SDKs read (gemini →
  GOOGLE_API_KEY for langchain), so `netgent agent` picks up .env keys with no
  shell sourcing.
- CLI root loads settings + syncs keys + sets log level; `netgent doctor` and the
  agent now see .env automatically. `netgent agent` defaults its model to
  NETGENT_GENERATOR_MODEL.

44 tests, ruff clean.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…-tests

Running the agent against browser-use/stress-tests challenges surfaced real gaps;
each fix is a genuine capability improvement:
- Observation now includes salient VISIBLE TEXT (headings, status/alert messages),
  so the agent can read success/confirmation — not just interactive elements. This
  alone unblocked the checkbox/radio/form challenges (agent couldn't see "completed").
- Elements show input type (input[date]/[file]/[email]), checkbox/radio [checked]
  state, and [disabled] — fixing a date-fill loop (needs YYYY-MM-DD) and letting the
  agent skip disabled/already-set controls.
- Correct ARIA roles for inputs by type (radio/checkbox/email/number/...), fixing
  locator resolution that mapped every <input> to 'textbox'.
- New set_checked action (Playwright .set_checked: label-aware, idempotent) + agent
  check/uncheck decisions — reliable for custom/hidden checkboxes and radios.

Verified against real challenges with claude-haiku-4-5: checkbox, radio-selection,
dropdown-selections, and a full multi-field signup all complete; hidden-labels fills
all 5 field types and is blocked only by a required file upload (upload_file is the
known remaining gap). 47 tests, ruff clean.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A headed run surfaced the last select gap: the agent couldn't see a dropdown's
valid option values, so it guessed ("United States") and select_option failed.
The snapshot now includes each <select>'s option values and the observation shows
them (options=[USA, UK, Canada]), so the agent picks a valid value first try.

Verified headful end-to-end: fill name + email, open/select the country dropdown,
check terms via set_checked, submit, read the success sentinel — 6 steps. 47 tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Two features the stress-tests' challenge.html / forms-comparison.html need:

- upload_file action (Playwright .set_input_files) + agent kind="upload". The
  agent offers a sample file to any file input autonomously (override via
  BrowserAgent(upload_file=...)). Verified: file-upload challenge passes in 1 step.
- The DOM snapshot now descends into same-origin iframes (cross-origin skipped),
  recording a frame-selector path per element; locators prepend a frame_locator
  chain so the executor pierces frames. Verified: forms-comparison.html exposes
  186 elements across 23 iframes; the agent filled 4 fields inside a nested-iframe
  form end-to-end.

Known: contenteditable fields (no .value) can cause a re-fill loop; challenge.html
is an aggregate of ~20 independent mini-challenges, not a sequential flow. 49 tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The snapshot found shadow-DOM elements, but locator resolution picked
get_by_role(role) from a placeholder-only field — Playwright's role-name
matching ignores placeholder, so it matched nothing and fills timed out.

Reorder _locator_for: prefer a simple #id css selector (precise, and Playwright's
css engine pierces open shadow roots) over role; use role only WITH a genuine
accessible name; skip nameless role locators entirely.

Verified against browser-use stress-tests shadow-dom-form: email fill, country
select, date fill, and file upload all now work inside the shadow DOM (previously
the first fill timed out). 49 tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
frameSelector emitted iframe:nth-of-type(n), which is sibling-scoped — on pages
where each iframe sits in its own container (forms-comparison.html: 23 iframes,
each alone in a div) it matched ALL 23, so frame_locator was ambiguous and every
in-frame action failed. Use a document-unique cssPath instead (id/name when present).

This matches how Notte/WebArena handle frames (a frame_locator chain over unique
CSS paths per iframe level). Verified: the frame selector now resolves to exactly
1 iframe, and the agent fills a full form inside an iframe (email, select, date,
radio, file upload). 49 tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Native HTML5 validation blocks form submit with a tooltip that isn't in the DOM,
so the agent would click Submit forever without knowing which field was empty.
The snapshot now reports each field's required + validity state, shown as
[required] / [invalid: still needs a valid value], and the prompt tells the agent
to fix those before retrying Submit.

Verified on forms-comparison.html: the first iframe form (email, select, date,
radio, file upload) now fills and submits in one pass — the agent's reasoning cites
the [invalid] markers to know what to fill. 50 tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Correction to an earlier limitation: Playwright's action layer already pierces
cross-origin frames (it drives the browser over CDP, bypassing same-origin policy).
The gap was only in observation — the snapshot walked iframe.contentDocument from
the TOP frame's injected JS, which same-origin policy blocks.

Fix: stop descending iframes in-JS; instead iterate page.frames and evaluate the DOM
walk inside EACH frame's own context (Playwright runs it there via CDP, cross-origin
included). The frame_locator path per frame is computed from Playwright's frame tree
(frame_element in the parent frame), so it works regardless of origin.

New test test_cross_origin_iframe: a child served on a different port (distinct
origin) embedded in a parent — the input is both observed and filled through the
frame_locator chain. 50 tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Long pages (forms-comparison.html: 23 forms/212 elements) exceeded the flat
80-element cap, so the agent never saw forms past the first few and scroll did
nothing. Three coordinated changes fix long-page navigation:

- Viewport-aware observation: element bbox.y is normalized to top-viewport
  coordinates (across frames), and the observation shows only the near-viewport
  slice with a POSITION line (top/middle/bottom) + "N above / N below" counts.
  Scrolling now shifts which slice is shown, so it makes real progress.
- Scroll action adopts browser-use's model: ScrollAction(down: bool, pages: float)
  converted to pixels from the live viewport height, replacing raw delta_y. No
  pixel guessing; direction is explicit (the old raw-pixels design let the model
  try to "scroll to top" with a negative amount = a no-op loop).
- Stuck detection is now observation-based: an action that changes nothing on
  screen counts as no-progress; a scroll that reveals a new batch does not.

Verified on forms-comparison.html: the agent pages through the whole document,
working ~13 forms deep (element 186/212) in one run vs stalling immediately before.
50 tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…obustness)

Addresses the two ways a long run stalled:

1. The agent never saw its own action errors — the history line dropped them, so it
   repeated a failing action until the no-change detector killed the run. History now
   includes "-> FAILED: <error>" so the agent can recover.

2. Silent wrong-action no-ops. to_action now guards element/action-type mismatches with
   corrective messages: select on a non-<select> ("use fill, dates YYYY-MM-DD"), fill on a
   dropdown ("use select"), check on a non-checkbox/radio, and a select value not in the
   element's options (lists them). These surface via the feedback channel above.

3. Survey-scrolling. The agent would scroll the whole page "to understand the layout"
   instead of acting. A firm prompt plus a scroll guard (reject scroll-down while
   [required]/[invalid] fields remain in view, naming them) forces it to fill before paging.

Result on forms-comparison.html: no scroll-survey stall and no wrong-action loop — the run
goes from 8 steps (all scrolling) to 78 steps working forms across the page with 2 scrolls.
Remaining limitation is honest self-assessment: the agent still over-claims "all submitted"
after ~a handful of forms, which is why the compiled-NFA design uses explicit accept
conditions rather than agent self-report. 51 tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…a time

"Do all the forms" reliably isn't one long free-form run (which over-claims done);
it's deterministic orchestration + LLM per unit + verified outcomes — NetGent's thesis.

- BrowserAgent.run gains frame_filter: scope the observation to one iframe so the
  agent works a single form at a time (DomSnapshot.scoped_to). TextBlock carries its
  frame_path so success can be checked per form.
- agent/sweep.py: enumerate the form frames, run a fresh agent scoped to each, then
  VERIFY submission by a success marker in that form's own text (not self-report).
  Reports per-form submitted vs agent-claimed — catching the gap in both directions.
- `netgent forms-sweep URL`.

Live on forms-comparison.html: 14/24 forms verified submitted, each attempted and
checked independently (vs a single run that submitted ~a handful then declared all
done). The 10 misses are harder forms that hit the per-form step cap or use trickier
widgets — improving them is per-form robustness, not orchestration. 52 tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…e reasons

Diagnosing why the sweep hit 14/24: the snapshot treated ANY element with a role as
interactive, so container roles (radiogroup, group, list, tablist) were listed and the
agent wasted steps clicking a radiogroup wrapper until it timed out. isInteractive now
uses an INTERACTIVE_ROLES whitelist (button/link/checkbox/radio/textbox/... — not
containers).

Also: FormResult now records stopped_reason + last_error, so a sweep says WHY each form
failed. Evidence from the 10 misses: ~5 custom framework radios that don't toggle on a
synthetic click (the hard widget class), 1 radiogroup-container click (fixed here), and 3
complex forms that hit the per-form step cap. 52 tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Drop the check/uncheck decision kinds (13 -> 11). The agent just clicks a
checkbox/radio; to_action routes it to the robust set_checked path — clicking a
checkbox toggles it (checked = not current), clicking a radio selects it. This is
browser-use's model (one click verb, checkbox smarts underneath), while the schema
keeps SetCheckedAction as the replay-safe primitive the executor dispatches.

52 tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…atch

The checkbox/radio smarts now live in the click dispatch (session._click), keyed on
the live element: a checkbox toggles (set_checked to the opposite of its current
state), a radio selects (check) — both via Playwright's label-aware, verified path —
and everything else is a plain click. So there's one ClickAction, no SetCheckedAction.

Schema action set 11 -> 10; the agent already had one click verb. Integration test
proves click toggles a checkbox both ways, selects a radio, and clicks a button.

Note: clicking a checkbox is now a toggle (reads live state), so a compiled workflow
that clicks a checkbox is only deterministic if the page's initial checkbox state is
stable — an accepted trade for the simpler single-action model. 53 tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Push toward completing more forms without changing the action set:

- session._click: after set_checked/check on a checkbox/radio, verify the state
  actually changed; if not (a custom control whose real <input> is hidden behind a
  styled label), click the associated label in the element's own context via JS —
  which fires the framework's own listener. This is the biggest failure bucket on
  the stress-test forms (Angular/custom radios). browser-use uses the same fallback.
- sweep_forms gains retries (default 2): each form is attempted with a fresh agent
  and a larger budget each try, stopping as soon as a success marker is verified —
  absorbing LLM run-to-run variance.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…it button

The enumeration counted any frame with a button as a form, so the top document
(page chrome around the embedded form iframes) became "form 1" — the agent then
clicked page headings forever. Require an input/select/textarea AND a button;
excludes the top frame (24 -> 21 real forms on forms-comparison).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…rash)

The viewport paging added for whole-page navigation regressed single-form filling:
scoped to one tall form, fields page out of view while their label text remains, so
the agent thinks an input is missing and scroll-thrashes for one it already filled.
scoped_to now zeroes viewport_height, showing the whole (bounded) form at once.

Diagnosed from a clean sweep: 12/21 verified (~57%), down from ~67% on real forms
before paging — this reverts that regression for the scoped case.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
EugeneVuong and others added 29 commits August 28, 2026 14:13
…ur's banner verified two broken fixtures)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ic oracles, feedback contracts; NetGent verifier design

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…on, replay determinism, feedback forms

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… wired as an orchestrator node; repeated-action guard

agent/verifier: Verdict + Evidence (task, params, action log, final observation, texts seen,
this run's dialogs, final URL, last 3 screenshots; never the explorer's reasoning — the judge
failure modes are sycophancy toward it), judge_trajectory via a new LLM.judge() seam method.
Orchestrator: explore → verify → generate; 'not achieved' re-explores once with the unmet
points appended; the replay still gates. `netgent generate --judge/--no-judge`;
NETGENT_JUDGE=1 on `eval stress sweep` scores the judge against page truth.
Measured (21-form sweep, Haiku 4.5): precision 100%, recall 100% after two fixes the
measurement forced — an outcome-not-fields rule, and dialogs scoped to the run.
Explorer: same-action-repeated guard (nudge at 3, stop at 6) — measured 12 identical clicks
on a YouTube overlay that observation-equality never caught.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…d-mirrored copy of the loop's structure)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…--graph`

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…iled graph; explorer refactor sketch

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…s and the settle watcher move out of Agent

The one class the LangGraph refactor keeps (langgraph-agent-structure.md §5.6 step 1): it owns
an asyncio.Task with a lifecycle, so it can live neither in graph state nor in a checkpoint.
Agent delegates to it for now; note/drain_noticed/start_watch/stop_watch are verbatim.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…frozen run context instead of closed-over names

Step 2 of langgraph-agent-structure.md §5.6: build_agent_graph still compiles per run, but it
now builds an ExplorerContext (session, llm, memory, task, knobs — with Agent.__init__'s
validation) and every node reads ctx.*; upload_path/capture_screenshot become module
functions taking the context. No behaviour change; the Runtime plumbing comes next.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…(), module-level EXPLORER, explore()

Step 3 of langgraph-agent-structure.md §5.6. observe/decide/act are module-level async
functions taking (state, runtime: Runtime[ExplorerContext]); the graph is compiled once at
import (StateGraph(AgentState, context_schema=ExplorerContext).compile(name="explorer")) and
a run passes its live session/LLM/memory as context= — never checkpointed. explore(...) is
the single run API (Agent.run's body, minus the per-run rebuild); Agent.run delegates to it.
langgraph is imported at module level in graph.py only; netgent.agent still loads without it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…graph; models → explorer/models.py; every call site on explore()/ExplorerAgent

Step 4 of langgraph-agent-structure.md §5.6, amended: the class is not deleted but reduced to
a thin façade — ExplorerAgent(llm, *, max_steps, run_dir, allowed_kinds, max_actions_per_step,
upload_file, memory) holds the knobs and ONE ExplorerMemory, and run() delegates to
graph.explore(); no loop logic. StepRecord/AgentStep/AgentTrajectory move to models.py
(MAX_REPEAT to graph.py, the fold constants to memory.py). explorer/__init__ and agent/__init__
re-export ExplorerAgent/ExplorerMemory/ExplorerContext + models, and resolve explore/
create_explorer_agent/EXPLORER lazily (PEP 562) so the package still loads without langgraph.
Call sites: orchestrator → explore(); CLI, sweep and stress → ExplorerAgent(...).run(...).
Tests follow; one agent test now exercises explore() with an explicit memory.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… invoke, context/façade validation; netgent.agent import boundary

Step 5 of langgraph-agent-structure.md §5.6. EXPLORER.get_graph().draw_mermaid() must show
observe/decide/act with exactly the loop and its exits — generated from the nodes' Command
annotations, the honest replacement for the diagrams dropped in 70a3a3b/0a70be2. A budget-0
run invokes the module-level graph with no browser and no key (context, not closure).
test_import_boundaries now also pins that importing netgent.agent loads no langchain/langgraph.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…tic graph nests the explorer

explore() takes an optional graph= (default EXPLORER; a checkpointer-compiled one persists) and
the orchestrator passes EXPLORER explicitly. LangGraph's subgraph discovery reads the node's
source for names it loads (get_function_nonlocals → find_subgraph_pregel), so an import inside
the node body — or a call through explore() alone — leaves get_subgraphs() empty; with the name
present, get_subgraphs() lists 'explore' and get_graph(xray=True) draws observe → decide → act
inside the pipeline (langgraph-agent-structure.md §3d probe C → A, §5.5 item 1). Pinned by test.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…verifier_agent(), module-level VERIFIER, verify()

judge.py splits into models.py (Evidence, Verdict), prompt.py (JUDGE_SYSTEM, build_judge_content),
context.py (VerifierContext: the LLM and screenshot dir as Runtime.context) and graph.py — two
function nodes, gather (trajectory → Evidence, pure) and judge (the one LLM call), compiled once
as VERIFIER with verify(traj, task, *, llm, params, run_dir, graph=None) as the single run API.
The package resolves the langgraph-importing names lazily, as the explorer does. Orchestrator
and sweep call verify(); the orchestrator names VERIFIER in its node so xray nests gather → judge
beside observe → decide → act. Existing imports (Evidence, Verdict, build_judge_content,
judge_trajectory) keep working. Tests: mermaid snapshot, a no-key run through FakeLLM, context
validation, and the pipeline's two subgraphs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…dator removed — pipeline is explore → verify → generate

planner/: models.py (Plan, PlanStep), prompt.py (PLANNER_SYSTEM, build_planner_content), context.py
(PlannerContext), graph.py (draft node, create_planner_agent(), module-level PLANNER, plan()),
agent.py (PlannerAgent). One LLM call through the seam's structured judge(); not yet wired into the
orchestrator — docs/OVERVIEW.md leaves how a plan feeds the explorer fleet open.
verifier/agent.py: VerifierAgent(llm, *, run_dir, max_screenshots).run(traj, task, params) → verify().
agent/validator/ and its test are deleted; the orchestrator loses the validate node, GenerateRequest
.validate_replay and GenerateResult.report/.validated; `netgent generate` loses --validate/--no-validate
(replay proof is `netgent run` on the artifact). Tests for both new packages; CLAUDE.md updated.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…e breakpoints, injectable chat model

Provider resolution is LangChain's: `provider:model` (a `/` is rewritten to `:`), so _PROVIDER_ALIAS
and the partition logic go; the Anthropic temperature quirk stays (2 lines). Default and all
call sites/docs use `google_genai:gemini-…` / `anthropic:claude-…`; doctor and settings accept
the new prefix. NOTE: the old `gemini/…` spelling no longer resolves — LangChain infers Vertex AI
for a bare `gemini-*` name.
The explicit Anthropic cache_control breakpoint is dropped; messages are a plain
SystemMessage(static) + HumanMessage(dynamic). Usage counters still report whatever the provider
cached implicitly.
LangChainLLM(model: str | BaseChatModel) — a built chat model is used as-is, which is the seam
LangChain's unit-testing guide describes: a GenericFakeChatModel now drives the real
with_structured_output → parser → retry-ladder path in tests, no key, no network.
RetryWithErrorOutputParser (langchain_classic) was evaluated and not adopted: it re-parses a text
completion and needs a text parser; tool-calling structured output has no completion to re-parse.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…middleware, same explore()/ExplorerAgent API (A/B arm)

agent/explorer_v2/: tools.py (one tool per atomic kind + done; each executes against
runtime.context.session, records the AgentStep with its durable locator and the StepRecord, runs
under a per-memory lock in call order since ToolNode gathers a turn's calls concurrently),
middleware.py (before_model = v1 observe: budget, scoped snapshot, stuck detection, texts_seen;
wrap_model_call = v1's prompt layout rebuilt every turn from the cross-run memory, the accumulated
messages are never sent; after_model = decide's guards: truncate to max_actions_per_step,
intercept `done`, repeated-action nudge/stop, re-observe on a text reply), state.py, prompt.py
(v1's prompt with the fields restated as tool calls), graph.py (create_explorer_agent(model,
allowed_kinds, max_actions) → create_agent(...), cached per tool set; explore() with v1's
signature; usage via LangChain's UsageMetadataCallbackHandler), agent.py (ExplorerAgent façade).
Reuses v1's ExplorerContext/ExplorerMemory/models; ExplorerContext.llm is now typed Any so
pydantic can build the tools' arg schemas. NETGENT_EXPLORER=v2 switches the sweep/stress arm;
SweepResult.usage carries v2's token totals. Known gaps vs v1: no settle watcher, no
structured-output retry ladder (a no-tool-call turn costs the step), no module-level graph.
Browser tests mirror test_agent.py with GenericFakeChatModel scripting tool calls.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…on decisions/steps, no PARAMETERS prompt section, no planner params

The explorer no longer declares which ${name} a value came from: AgentAction/AgentDecision lose
`param`, AgentStep loses `param`, the prompt loses its PARAMETERS section and field, explorer_v2's
tools lose the `param` argument. The planner loses its `params` input (state, prompt, plan(),
PlannerAgent.run). The compiler's structural binding pass goes with it; `-p name=sample` still
works through the case-insensitive literal sweep over value fields and state conditions (never
locators), and the orchestrator still tells the explorer the sample values to use verbatim.
Tests: test_param_binding rewritten for sweep-only semantics; prompt/planner/compiler tests adjusted.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ation re-attributed), induction literature, typed-key merge proposal

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…me-family tasks with proposed param values, normalized in code

One LLM call (VARIATION_PLANNER, the planner's graph shape) returns [{task_text, values}];
normalize_variation_plan (pure) forces variation 1 to the base task verbatim, makes run-1's
value names canonical, applies --variation pins to variation 2, and pads/truncates to N.
The names are hypotheses only — the merge confirms them as Params when values vary in the
trajectories.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…der <name>.trajectories

run-<k>/ holds trajectory.json + screenshots (explore writes them), variation.json and
verdict.json (achieved flag, attempts); a retried run's first attempt is stashed, not lost;
generalized.json at the root is where the merge's induced memory lands. Failed runs are
stored too, marked — they are memory. Pure file I/O, zero LLM.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…e generalized NFA, pure code

Align achieved runs on (action type, durable target key) (Needleman-Wunsch over typed sigs;
AWM's abstract-trajectory signature with durable locators). Conditions by version-space
intersection: url_matches/anchors/dialogs survive only when they held in EVERY run, with
support counts in the evidence trail. Divergence gets the four dispositions of
trajectory-memory.md §C.1.3:
 1. value-varies + matches a planner value -> Param (fill text/goto url substitution,
    dwell -> Repeat(count="${watch_time}") of 1s slices behind a noop edge, and role-name
    targets containing the value -> get_by_role(role, name="${param}") + nth(0));
 2. present-in-k-and-dismissal-shaped -> Interrupt (cross-run presence primary, text
    heuristics tie-break; is_interruption_step lifted to module level in compiler.py);
 3. genuine downstream divergence with distinguishing first targets -> Branch, one arm per
    continuation, converging on one state;
 4. else reject: warning naming the column, spine (run 1) kept so the artifact replays.
Failed runs contribute nothing structural (failures poison trajectory-shaped memory).
accept_states = the final state's intersected conditions, never invented. generalized.json
(GeneralizedTrajectory) is the induced memory. Executor: Repeat.count numeric-string
coercion, since resolve_params substitutes ${p} into count's str type.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ped merge -> zero-LLM replay check

New LangGraph (runs=1 keeps the single-run graph untouched): plan (variation planner, one
LLM call; -p names join as proposals) -> explore_run xN (fresh ExplorerMemory per run via
explore(); per-run verify with ONE private retry whose unmet-points suffix never crosses
runs; every attempt persisted in the TrajectoryStore) -> merge (merge_trajectories, pure
code; generalized.json written; artifact dumped) -> replay (agent/replay.py: metamorphic
zero-LLM check — run-1's and run-2's value sets must walk the same state sequence,
interrupt edges excluded, dwell self-loops collapsed; mismatch sets result.error).
Independence policy: runs share only a HINTS line of previously seen dismissal anchors.
CLI: --runs / --variation name=value. Integration tests: N=2 fixture e2e (params inferred,
store layout, both-value replays), private retry suffix, no-achieved-run stop.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…matches a 10.0 dwell, stored as the bare number for Repeat.count

Observed live: the variation planner writes durations with units; the count feeds 1s slices
and must resolve numerically.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
1. Alignment scoring is locator-shape-aware: same-type substitution scores 2 only when the
   target shape agrees (both role=link, both css, ...), else 0 — the shape-blind +1 aligned
   run-1's css play-button click with the other runs' video-title links, dropping the real
   video click and leaving an unguarded artifact that passed replay vacuously.
2. Disposition 4 minority steps are DROPPED (structural intersection), not kept: the runs
   without them achieved the task, and keeping run-1-only steps (suggestion click, play
   click) made replay time out on elements that never reappear. Branch synthesis is now
   presence-based only (all-gap regions, the login-wall shape) — a substitution column in a
   region is value variance, and arms guarded on per-run targets would freeze the values.

After the fixes the stored 3-run YouTube memory re-merges (offline, zero LLM) to:
goto -> fill ${video_query} -> click search -> click video (target-varies, run-1 kept) ->
5s dwell; replay set 1 walks s1..s5 to the real watch page; set 2 fails honestly at the
value-dependent click (titles do not contain the query - the known open gap).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…or shape

Measured on the second real YouTube merge: run A clicks the video as a role link, run B as
a >> filter(has_text) — shapes differ, so the shape-only rule paired run B's video click
with run A's play-button click and the compiled word never opened a video. Each step now
carries its URL effect (did the base change, where to); a substitution pairing scores 2
with shape+effect, 1 with either, 0 with neither. With the fix the stored 2-achieved-run
memory re-merges offline to a workflow with BOTH params (video_query -> fill text,
watch_time -> Repeat.count) and the zero-LLM metamorphic replay PASSES: both value sets
walk s1..s5 through the real url_matches(/watch) gate, dwelling 10 vs 5 slices.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…t-aware stuck detection

The observation layer lied to the agent on video pages, three ways at once
(diagnosed on live YouTube runs, runs/youtube-warpigs trajectory as evidence):

- YouTube freezes its control-bar DOM (time display, Play/Pause labels) while
  the controls are auto-hidden, and the walker's visible() only checked the
  element's OWN computed style — ancestor opacity:0 doesn't propagate into it —
  so frozen, invisible controls were reported as live UI. The explorer clicked
  pause, was shown the stale pre-click label, and toggled itself into the
  repeat-stop (2 runs); the frozen timer made playing ads compare byte-equal
  across steps and fire the no-change stuck stop (2 runs).

- Nothing read the <video> element's own properties, the one signal that never
  freezes.

Fixes, one per layer:

- snapshot.js: emit media state (currentTime/duration/paused/ended/muted) per
  visible-or-audible <video>/<audio>, read-only property access; visible() now
  uses native checkVisibility({opacityProperty, visibilityProperty}) so
  ancestor-faded controls drop out of the observation instead of lying.
- models/observer/serializer: MediaState -> DomSnapshot.media -> a
  "MEDIA: video PLAYING at 0:29 / 7:56" line under TITLE. Its ticking position
  also keeps a playing page from ever comparing equal to its previous
  observation.
- explorer/graph.py: no_progress increments only when the rendered observation
  is unchanged AND no never-seen text appeared — texts_seen was recording ad
  captions advancing at the very moment the stuck stop declared "no change on
  screen".

Measured effect: the War Pigs intent went from 4 consecutive failed
generations (pause-toggle x2, false-stuck x2) to a clean judge-approved
compile whose zero-LLM replay passes end-to-end; exploration reasoning now
cites playback state instead of guessing it from button labels.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@Manni-MinM Manni-MinM closed this Sep 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants