From 2c3647aeac84c6474978f5e55f3907afbdbad825 Mon Sep 17 00:00:00 2001 From: rockymadden Date: Fri, 4 Sep 2026 09:38:58 -0600 Subject: [PATCH 1/7] feat(uipath-maestro-flow): state headlessness in the flow eval config, not in 68 task prompts MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit In the 2026-09-04 nightly, 5 of 8 `skill-flow-*` tasks built and validated a flow, reported success, and never executed it. The checker then ran `flow debug` and found a null End-node output mapping, a faulted script, and an empty result. `flow validate` had passed on all of them. Every one of those prompts contained "Do NOT ask for approval, confirmation, or feedback". That phrasing forbids asking. It does not say nobody is there to ask, so an agent can honor it and still stop at a consent gate waiting for a reply that never arrives. Measured across the 128-task flow suite: 0 tasks stated the run was headless, 51 of the 119 non-simulated tasks said nothing about autonomy at all, and the 68 that did were spread across 8 wording variants. The fact belongs in the harness config, the domain behaviour in the skill. Eval side: - tests/experiments/flow.yaml — nightly's runtime with a system prompt that states the run is headless and what to do about it. Derived from nightly.yaml so the runtime cannot drift. - run-coder-eval.yml — `experiment` input, defaulting to nightly.yaml. A config's defaults apply to whatever its run selects, so scoping is task_globs plus -e, not a fork of the workflow. - tests/Makefile — `make flow` runs the suite with that config and excludes interactive/, whose 9 tasks have a simulated user and must never be told nobody is present. The exclusion is the config's contract, so it lives with the target. - 69 task files — the hand-copied autonomy lines removed. Task-specific text that shared those sentences ("Do NOT substitute a mock", "Do NOT run or debug the flow") is preserved. - test-task-template.yaml — stop telling authors to add the line. Skill side, both independent of headlessness and true with a user watching: - `flow debug` consent comes from the mandate. A request to build something that does X is a request for it to work, and building plus validating does not discharge that. Debug also overwrites the Studio Web solution behind the local .uipx SolutionId, so never debug a solution this run did not scaffold. - "Publish to Studio Web" is no longer marked `(default)` in either What's next dropdown. Rule #5's non-interactive fallback takes the marked option, which would have auto-published to a tenant. Modifies Critical Rules 2 and 5, per CONTRIBUTING. Co-Authored-By: Claude Opus 5 (1M context) --- .github/workflows/run-coder-eval.yml | 13 ++- skills/uipath-maestro-flow/SKILL.md | 4 +- .../references/author/brownfield.md | 2 +- .../references/author/greenfield.md | 2 +- .../references/operate/run.md | 2 +- tests/Makefile | 9 ++- tests/experiments/flow.yaml | 79 +++++++++++++++++++ .../contractregistry_crud_filters.yaml | 3 - .../e2e_contract_intake_pipeline.yaml | 3 - .../integration_create_get.yaml | 3 - .../smoke_create_all_types.yaml | 3 - .../datafabric_connector/smoke_error.yaml | 3 - .../smoke_file_activities.yaml | 3 - .../datafabric_connector/smoke_query.yaml | 3 - .../datafabric_connector/smoke_update.yaml | 3 - .../smoke_update_existing_flow.yaml | 3 - .../trigger_lifecycle.yaml | 3 - .../connector_features/drive_to_slack.yaml | 2 +- .../generic_dynamic_node.yaml | 4 +- .../jdbc_databricks_query.yaml | 3 - .../non_catalog_http_fallback.yaml | 2 - .../paginated_reference_lookup.yaml | 2 - .../slack_http_fallback.yaml | 2 - .../testmanager_attachments.yaml | 3 - .../testmanager_crud_grounded.yaml | 3 - .../testmanager_execution_results.yaml | 3 - .../testmanager_generic_records.yaml | 3 - .../testmanager_requirement_lifecycle.yaml | 3 - .../testmanager_testcase_lifecycle.yaml | 3 - .../testmanager_testset_lifecycle.yaml | 3 - .../trigger_with_filter.yaml | 1 - .../webhook_waitfor_parallel.yaml | 2 - .../batch_transform/batch_transform.yaml | 2 - .../summarize/summarize.yaml | 2 - .../escalation_jira_ticket.yaml | 3 +- .../escalation_orchestrator_paths.yaml | 3 +- .../escalation_slack_alert.yaml | 6 +- .../edit/add_node/add_node.yaml | 1 - .../edit/add_output/add_output.yaml | 1 - .../group_to_subflow/group_to_subflow.yaml | 1 - .../edit/move_node/move_node.yaml | 1 - .../edit/remove_node/remove_node.yaml | 1 - .../edit/update_node/update_node.yaml | 1 - .../inline_agent_eval/inline_agent_eval.yaml | 4 +- .../hitl/smoke_01_hitl_node_placed.yaml | 1 - .../e2e_01_invoice_extraction_greenfield.yaml | 1 - .../e2e_03_project_creation_handoff.yaml | 4 +- .../bellevue_weather/bellevue_weather.yaml | 1 - .../multi_node/calculator/calculator.yaml | 1 - .../customer_escalation.yaml | 3 - .../multi_node/dice_roller/dice_roller.yaml | 1 - .../multi_node/feet_inches/feet_inches.yaml | 2 - .../loop_multiply/loop_multiply.yaml | 1 - .../multi_city_weather.yaml | 1 - .../multi_node/reading_list/reading_list.yaml | 1 - .../slack_channel_description.yaml | 1 - .../slack_weather_pipeline.yaml | 1 - .../wiki_pageviews/wiki_pageviews.yaml | 2 - .../api_workflow/api_workflow.yaml | 1 - .../single_node/coded_agent/coded_agent.yaml | 1 - .../single_node/decision/decision.yaml | 1 - .../single_node/delay/delay.yaml | 2 - .../file_attachment/file_attachment.yaml | 4 +- .../lowcode_agent/lowcode_agent.yaml | 1 - .../openmeteo_weather/openmeteo_weather.yaml | 4 +- .../outlook_trigger_inbox.yaml | 2 - .../outlook_waitfor_email.yaml | 3 - .../single_node/rpa/rpa.yaml | 1 - .../single_node/subflow/subflow.yaml | 1 - .../single_node/switch/switch.yaml | 1 - .../single_node/terminate/terminate.yaml | 1 - .../transform_filter/transform_filter.yaml | 4 +- .../transform_group_by.yaml | 2 - .../transform_map/transform_map.yaml | 2 - .../smoke/merge_parallel_sync.yaml | 5 +- .../smoke/scheduled_trigger.yaml | 5 +- tests/templates/test-task-template.yaml | 6 +- 77 files changed, 120 insertions(+), 154 deletions(-) create mode 100644 tests/experiments/flow.yaml diff --git a/.github/workflows/run-coder-eval.yml b/.github/workflows/run-coder-eval.yml index df40b5bb73..28f37e2289 100644 --- a/.github/workflows/run-coder-eval.yml +++ b/.github/workflows/run-coder-eval.yml @@ -23,6 +23,14 @@ on: description: 'REQUIRED. Space-separated globs under tests/. E.g. tasks/uipath-agents/**/*.yaml' type: string required: true + # Per-skill configs differ only in defaults the runtime shares, so the + # file is a parameter rather than a fork of this workflow. Scope it with + # task_globs: a config's defaults apply to whatever that run selects. + experiment: + description: 'Experiment YAML under tests/. E.g. experiments/flow.yaml for the flow zero-shot config.' + type: string + required: false + default: 'experiments/nightly.yaml' # Host-level concurrency for the Linux job's `coder-eval -j`. Default 4: # ubuntu-latest is 4 vCPU, and j=20 oversubscribes ~5:1 — agents miss the # 1200s turn timeout and ERROR (false negatives, including a bindings @@ -409,6 +417,7 @@ jobs: E2E_LONG_PROCESS_KEY: ${{ secrets.E2E_LONG_PROCESS_KEY }} TASK_GLOBS: ${{ needs.partition.outputs.linux_globs }} TASK_COUNT: ${{ needs.partition.outputs.linux_count }} + EXPERIMENT_YAML: ${{ inputs.experiment }} TASK_PARALLELISM: ${{ inputs.parallelism }} AGENT: ${{ inputs.agent }} AGENT_MODEL: ${{ inputs.agent_model }} @@ -456,7 +465,7 @@ jobs: env_lines=() for name in SKILLS_REPO_PATH API_BACKEND AWS_BEARER_TOKEN_BEDROCK AWS_REGION \ BEDROCK_MODEL ANTHROPIC_API_KEY E2E_PROCESS_KEY E2E_LONG_PROCESS_KEY \ - TASK_GLOBS TASK_COUNT TASK_PARALLELISM AGENT AGENT_MODEL \ + TASK_GLOBS TASK_COUNT TASK_PARALLELISM AGENT AGENT_MODEL EXPERIMENT_YAML \ CLAUDE_CODE_MODEL CODEX_MODEL ANTIGRAVITY_MODEL DELEGATE_MODEL \ CODEX_API_KEY CODEX_BASE_URL GEMINI_API_KEY; do env_lines+=("$name=${!name}") @@ -465,7 +474,7 @@ jobs: # Args appended to `coder-eval run`, ONE PER LINE, each verbatim — no # word splitting, no pathname expansion. That is what the delegate `-D` # override needs: to bash its `[...]` list is a character class. - args=(-e experiments/nightly.yaml) + args=(-e "${EXPERIMENT_YAML}") # Agent selection. codex authenticates via CODEX_API_KEY/CODEX_BASE_URL, # antigravity via GEMINI_API_KEY (SDK baked into the agent image), diff --git a/skills/uipath-maestro-flow/SKILL.md b/skills/uipath-maestro-flow/SKILL.md index 182d0e7959..630d6eda28 100644 --- a/skills/uipath-maestro-flow/SKILL.md +++ b/skills/uipath-maestro-flow/SKILL.md @@ -61,7 +61,7 @@ Guide for creating, editing, validating, debugging, publishing, diagnosing, and > **Tool vocabulary.** `Edit` means in-place replacement, `Write` a full-file write, `Read`/`Glob`/`Grep` file access, `Bash` shell, and a progress list the harness task list. Map them to equivalent tools elsewhere; preserve reviewable diffs and use shell file edits only as a last resort. 1. **Use `--output json`; prefer `--output-filter` for extraction.** Filters are global and run against the `Data` envelope, so expressions start at `Data` without a `Data.` prefix. Registry search returns a flat PascalCase array (`NodeType`, `DisplayName`, `Description`, `AvailableOnTenant`), not `Data.Nodes` or lowercase fields. Example: `uip maestro flow registry search --output json --output-filter "[*].{NodeType:NodeType,DisplayName:DisplayName,Description:Description,AvailableOnTenant:AvailableOnTenant}"`. With `--local`, omit `AvailableOnTenant`. Use `python3 -c` or `jq` only after verifying shape and when JMESPath cannot express the transform. See [cli-conventions.md §3](references/shared/cli-conventions.md#3-prefer---output-filter-for-extraction). -2. **Do not run `flow debug` without explicit user consent.** It executes the flow for real (sends emails, posts messages, calls APIs). +2. **`flow debug` consent comes from the mandate.** It executes the flow for real (sends emails, posts messages, calls APIs), so run it only when the request is for a flow that *works*: the user asked for something that does X, or said make it work, get it running, iterate until it passes. Building and validating does not discharge that, and a flow that was never executed is not finished. Ask when the request stops short of a working artifact (review this, add a node, validate only); with nobody to ask, report debug as the step not run. Debug also overwrites the Studio Web solution matching the local `.uipx` `SolutionId`, so never debug a solution this run did not scaffold. 3. **Search before creating or declaring resources absent.** For named agents, API workflows, RPA processes, and similar resources: (a) pull and search the tenant registry with `uip maestro flow registry pull --force && uip maestro flow registry search "" --output json`; pull first because the cache expires after 30 minutes, login is required, and only published resources are returned; (b) search locally with `uip maestro flow registry list --local --output json` or `search "" --local` (no login; returns sibling projects in the same `.uipx` solution); an empty keyword search does not prove absence, so confirm with `list --local`; (c) scaffold, mock, or create only when both searches find no match and the user explicitly requests embedding/creation or no published resource satisfies the need. "Coded" and "low-code" describe implementation style, not inline status. Use `uipath.agent.autonomous` only when explicitly asked to embed/inline/create an agent. Use `core.logic.mock` only when the resource is neither in the solution nor published. See [rpa](references/author/plugins/rpa/impl.md) and [agent](references/author/plugins/agent/impl.md). @@ -73,7 +73,7 @@ Guide for creating, editing, validating, debugging, publishing, diagnosing, and **Two tells that you skipped the search and took the brand-name shortcut — both are build defects, not valid manual-mode HTTP:** (a) you authored a manual-mode `core.action.http.v2` node whose `url` targets a well-known SaaS API domain that has a connector (`slack.com/api/*`, `api.github.com`, `*.salesforce.com`, `graph.microsoft.com`, …); (b) you declared an `in` variable to hold that service's API token or secret (e.g. a `slackToken` holding an `xoxb-…` bot token, an `apiKey`, a bearer token). A connector-backed flow never carries the raw credential — the IS connection does. If you find yourself writing either, **stop**: run `uip maestro flow registry search ""` and `uip is connections list "" --all-folders`, then use the connector activity (or connector-mode HTTP: `authentication:"connector"` + `targetConnector` + a bound `connectionId`/`folderKey`). Manual mode is legitimate only for a service the search proves has no connector. 4. **Never invoke other skills automatically** — when a flow needs an RPA process, agent, or app, identify the gap and provide handoff instructions. Let the user decide when to switch skills. **One exception — IXP extraction with documents in hand:** when the flow needs document extraction, the user supplied sample documents, and `registry search "uipath.ixp"` shows no extractor covering them, invoke the `uipath-ixp` skill to build and deploy the model, then resume the flow ([plugins/ixp/impl.md — If the Model Does Not Exist Yet](references/author/plugins/ixp/impl.md#if-the-model-does-not-exist-yet)). Resolve the target Orchestrator folder for the deployment before invoking — from the user's request when it names one, otherwise per rule #5 (its non-interactive fallback applies) — and pass it in the handoff; the sibling stops rather than guess a folder. There is deliberately no separate consent gate on the tenant writes this creates: the project and folder deployment fulfil the extraction request itself, and the one consequential choice — where the deployment lands (deployments have no delete API) — is exactly the folder decision rule #5 just routed. Do NOT drive `uip ixp` project or deployment commands from this skill instead of invoking it — the sibling's guides carry guardrails this skill does not. If `uipath-ixp` is unavailable in the session, fall back to `core.logic.mock` plus an Open Questions entry, exactly as when no documents were supplied. -5. **Always present finite decisions as a dropdown with a final "Something else" escape hatch.** Whenever the skill needs a decision (which solution, publish vs debug vs deploy, which connector, trigger type, or resource to bind, etc.), ask with the enumerated choices plus **"Something else"** last for free-form input; never ask open-ended in chat when a finite set of sensible defaults exists. If the user picks "Something else", parse their answer and continue. No structured-question facility on the harness → ask in chat as a numbered list with "Something else" last. Non-interactively (CI/headless, no user available) → take the marked recommended option, proceed, and record the decision prominently in the final report; if none is recommended, stop and report the open decision instead of guessing. Consent gates (`flow debug`, destructive operations) are never auto-answered — in non-interactive mode, stop and report the blocked step. These fallbacks define "ask the user" / "confirm with the user" wherever this skill's references require it. +5. **Always present finite decisions as a dropdown with a final "Something else" escape hatch.** Whenever the skill needs a decision (which solution, publish vs debug vs deploy, which connector, trigger type, or resource to bind, etc.), ask with the enumerated choices plus **"Something else"** last for free-form input; never ask open-ended in chat when a finite set of sensible defaults exists. If the user picks "Something else", parse their answer and continue. No structured-question facility on the harness → ask in chat as a numbered list with "Something else" last. These fallbacks define "ask the user" / "confirm with the user" wherever this skill's references require it. 6. **Discover the target solution before scaffolding.** A Flow project must use double nesting: `//.flow`. Before any new `uip solution init` or `uip maestro flow init`, run `find . -maxdepth 2 -type f -name '*.uipx' -print`. If a solution exists, stop and ask which to use: one option per solution, "Create a new solution", then "Something else". Do not silently adopt, initialize, delete, or repair an existing solution, even if a new one was requested. If creating one, ask for its name rather than defaulting to the Flow name. diff --git a/skills/uipath-maestro-flow/references/author/brownfield.md b/skills/uipath-maestro-flow/references/author/brownfield.md index 707203f2a1..8523dee50f 100644 --- a/skills/uipath-maestro-flow/references/author/brownfield.md +++ b/skills/uipath-maestro-flow/references/author/brownfield.md @@ -81,7 +81,7 @@ Authoring ends here. For any selected option, read [operate/CAPABILITY.md](../op | Option | What it does | |---|---| -| **Publish to Studio Web** (default) | Push the solution to Studio Web so the user can visualize, edit, and publish from the browser. | +| **Publish to Studio Web** | Push the solution to Studio Web so the user can visualize, edit, and publish from the browser. | | **Debug the solution** | Execute the flow end-to-end against real systems. Confirm consent first because debug has real side effects (see the consent-before-debug rule in [SKILL.md](../../SKILL.md)). | | **Deploy to Orchestrator** | Pack and publish directly to Orchestrator (bypasses Studio Web). Only when explicitly chosen; see [/uipath:uipath-platform](/uipath:uipath-platform). | | **Something else** | Last option. Accept free-form string input and act on it. | diff --git a/skills/uipath-maestro-flow/references/author/greenfield.md b/skills/uipath-maestro-flow/references/author/greenfield.md index 9f2d049dc1..769b1e89c8 100644 --- a/skills/uipath-maestro-flow/references/author/greenfield.md +++ b/skills/uipath-maestro-flow/references/author/greenfield.md @@ -373,7 +373,7 @@ Authoring terminates here. Each option below hands off to Operate — read [oper | Option | What it does | | --- | --- | -| **Publish to Studio Web** (default) | Push the solution to Studio Web so the user can visualize, edit, and publish from the browser. | +| **Publish to Studio Web** | Push the solution to Studio Web so the user can visualize, edit, and publish from the browser. | | **Debug the solution** | Execute the flow end-to-end against real systems. Confirm consent first — debug has real side effects (see the consent-before-debug rule in [SKILL.md](../../SKILL.md)). | | **Deploy to Orchestrator** | Pack and publish directly to Orchestrator (bypasses Studio Web). Only when explicitly chosen — see [/uipath:uipath-platform](/uipath:uipath-platform). | | **Something else** | Last option. Accept free-form string input and act on it (e.g., "just leave it", "pack but don't publish", "upload to a different tenant"). | diff --git a/skills/uipath-maestro-flow/references/operate/run.md b/skills/uipath-maestro-flow/references/operate/run.md index 0c808e8d61..042450301a 100644 --- a/skills/uipath-maestro-flow/references/operate/run.md +++ b/skills/uipath-maestro-flow/references/operate/run.md @@ -13,7 +13,7 @@ Execute a flow on demand and monitor progress. Three modes: **debug** (controlle ## Debug — controlled end-to-end run -> **Confirm consent first.** `flow debug` executes the flow for real — sends emails, posts messages, calls APIs. See the consent-before-debug rule in [SKILL.md](../../SKILL.md). Do not run without explicit user authorization. +> **Consent comes from the mandate.** `flow debug` executes the flow for real — sends emails, posts messages, calls APIs. Run it when the request is for a flow that works; ask when the request stops at build or validate. Never debug a solution this run did not scaffold: debug overwrites the Studio Web solution matching the local `.uipx` `SolutionId`. See the `flow debug` consent rule in [SKILL.md](../../SKILL.md). ```bash UIP_LOG_LEVEL=info uip maestro flow debug --output json diff --git a/tests/Makefile b/tests/Makefile index 72d9f1a0b4..307de79b12 100644 --- a/tests/Makefile +++ b/tests/Makefile @@ -1,9 +1,13 @@ -.PHONY: help install all smoke smoke_rpa e2e tags +.PHONY: help install all smoke smoke_rpa e2e tags flow SKILLS_REPO_PATH ?= $(shell cd .. && pwd) VENV := .venv CODER_EVAL := SKILLS_REPO_PATH=$(SKILLS_REPO_PATH) $(VENV)/bin/coder-eval TASKS := $(shell find tasks -name '*.yaml' -type f) +# Flow's zero-shot config states the run is headless, which is false for the +# simulated tasks under interactive/. The exclusion is the config's contract, +# so it lives with the target rather than in a comment someone has to find. +FLOW_TASKS := $(shell find tasks/uipath-maestro-flow -name '*.yaml' -type f -not -path '*/interactive/*') TASK_PARALLELISM ?= 1 help: ## Show available commands @@ -38,6 +42,9 @@ smoke_rpa: ## Run all Windows RPA smoke tests (tempdir) e2e: ## Run all end-to-end tests $(CODER_EVAL) run $(TASKS) -e experiments/default.yaml --tags e2e -j $(TASK_PARALLELISM) -v +flow: ## Run the flow suite zero-shot (headless system prompt; excludes interactive/) + $(CODER_EVAL) run $(FLOW_TASKS) -e experiments/flow.yaml -j $(TASK_PARALLELISM) -v + tags: ## Run tests matching one or more tags: make tags TAGS="connector-feature" [EXPERIMENT=experiments/default.yaml] @if [ -z "$(TAGS)" ]; then echo "Usage: make tags TAGS=\"tag1 tag2\" [EXPERIMENT=experiments/.yaml]"; exit 2; fi @TAGS="$(TAGS)" python3 -c 'import os,re,sys,glob; \ diff --git a/tests/experiments/flow.yaml b/tests/experiments/flow.yaml new file mode 100644 index 0000000000..e95c75fa8e --- /dev/null +++ b/tests/experiments/flow.yaml @@ -0,0 +1,79 @@ +experiment_id: skill-flow-zero-shot +description: > + Flow zero-shot config — identical runtime to nightly.yaml, with a system + prompt that states the run is headless. Pair it with a task glob that + excludes tests/tasks/uipath-maestro-flow/interactive/**, whose tasks have a + simulated user and must not be told nobody is present. + +defaults: + run_limits: + max_turns: 200 + task_timeout: 1200 + # A single turn can chain several cold-runner `uip rpa` calls (each 30-90s); + # per-turn budget tracks the slowest external call, not the suite's intent. + turn_timeout: 900 + + sandbox: + driver: docker + docker: + image: skills-image:latest + env_passthrough_extra: + - SKILLS_REPO_PATH + - BEDROCK_MODEL + - TASK_PARALLELISM + - UIPATH_CLI_ENABLE_ENV_AUTH + - UIPATH_CLI_AUTH_TOKEN + - UIPATH_CLI_ORGANIZATION_NAME + - UIPATH_CLI_ORGANIZATION_ID + - UIPATH_CLI_TENANT_NAME + - UIPATH_CLI_TENANT_ID + - E2E_PROCESS_KEY + - E2E_LONG_PROCESS_KEY + - CODEX_API_KEY + - CODEX_BASE_URL + extra_mounts: + - ~/.uipath:/.uipath:rw + + agent: + # No agent.type / agent.model here: the runner passes both as CLI flags, which outrank this file. + permission_mode: acceptEdits + allowed_tools: ["Skill", "Bash", "Read", "Write", "Edit", "Glob", "Grep"] + # Task YAMLs used to carry hand-copied "Do NOT ask for approval" lines for + # this. That phrasing forbids asking without saying nobody is there to ask, + # so an agent could honor it and still stop at a consent gate waiting for a + # reply that never arrives — 5 of 8 flow tasks did exactly that on + # 2026-09-04. Stated once here instead, and removed from the tasks. + system_prompt: | + You are a coding agent. Do not access files in sibling runs/* directories. Everywhere else is permitted. + + This run is headless. No user is present, and nobody will answer a question or grant an approval. + Do not wait for input and do not stop to ask. Take the best available option, supply the most + defensible value where one is missing, and keep going. Stop only when proceeding would be unsafe + or irreversible, or when you could not resolve a fact the work depends on — never substitute a + guess for it. Record every decision, assumption, and blocked step in your final response. + Instructions in the task take precedence over this paragraph. + plugins: + - type: "local" + path: "$SKILLS_REPO_PATH" + ignore_patterns: [] + + checker_context: + api_route: + route: litellm + model: gpt-5.6-luna + params: + api_version: "2024-05-01" + num_retries: 5 + env_params: + api_base: CODEX_BASE_URL + api_key: CODEX_API_KEY + + # Sandbox cleanup between tasks. Tasks that upload Studio Web solutions + # wire cleanup_solutions.py into their own post_run — that must run before + # this step (still needs npm-installed CLIs). + post_run: + - command: "find . -maxdepth 5 -type d \\( -name node_modules -o -name .npm-prefix -o -name .venv \\) -prune -exec rm -rf {} +" + timeout: 30 + +variants: + - variant_id: default diff --git a/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/contractregistry_crud_filters.yaml b/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/contractregistry_crud_filters.yaml index e26e458ccc..49a446dfc6 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/contractregistry_crud_filters.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/contractregistry_crud_filters.yaml @@ -35,9 +35,6 @@ initial_prompt: | 6. Map the created, queried, updated, and retrieved values to the flow output where the activity supports output mapping. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between - planning and implementation. Build the complete flow end-to-end in a - single pass. The Flow is not complete until `uip maestro flow validate` passes. success_criteria: diff --git a/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/e2e_contract_intake_pipeline.yaml b/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/e2e_contract_intake_pipeline.yaml index 420f028c7e..dc544a60bb 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/e2e_contract_intake_pipeline.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/e2e_contract_intake_pipeline.yaml @@ -41,9 +41,6 @@ initial_prompt: | Entity Record ONCE inside the loop, binding recordId to the loop item's Id. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between - planning and implementation. Build the complete flow end-to-end in a - single pass. The Flow is not complete until `uip maestro flow validate` passes. success_criteria: diff --git a/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/integration_create_get.yaml b/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/integration_create_get.yaml index 2b91cd6977..2b68f20942 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/integration_create_get.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/integration_create_get.yaml @@ -37,9 +37,6 @@ initial_prompt: | at expansionLevel=1, expansionLevel=2, and expansionLevel=3. Delete when done. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between - planning and implementation. Build the complete flow end-to-end in a - single pass. The Flow is not complete until `uip maestro flow validate` passes. success_criteria: diff --git a/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/smoke_create_all_types.yaml b/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/smoke_create_all_types.yaml index a93cfbad58..bce50ac79b 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/smoke_create_all_types.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/smoke_create_all_types.yaml @@ -30,9 +30,6 @@ initial_prompt: | * lastUpdated: "2024-03-10T09:00:00" (DATETIME) * externalId: "aaaaaaaa-bbbb-cccc-dddd-eeeeeeeeeeee" (UUID) - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between - planning and implementation. Build the complete flow end-to-end in a - single pass. The Flow is not complete until `uip maestro flow validate` passes. success_criteria: diff --git a/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/smoke_error.yaml b/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/smoke_error.yaml index db20860861..ec269379bf 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/smoke_error.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/smoke_error.yaml @@ -22,9 +22,6 @@ initial_prompt: | FlowCodeEvalEntity for title='ErrorTestRecord'. 3. In the same parallel branch, Query FlowCodeEvalEntity with no filter. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between - planning and implementation. Build the complete flow end-to-end in a - single pass. The Flow is not complete until `uip maestro flow validate` passes. success_criteria: diff --git a/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/smoke_file_activities.yaml b/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/smoke_file_activities.yaml index 281fa4a058..5b232376da 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/smoke_file_activities.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/smoke_file_activities.yaml @@ -36,9 +36,6 @@ initial_prompt: | 4. DeleteFileFromRecordFieldV2 — bind recordId to the same create activity output Id used in step 3. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between - planning and implementation. Build the complete flow end-to-end in a - single pass. The Flow is not complete until `uip maestro flow validate` passes. success_criteria: diff --git a/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/smoke_query.yaml b/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/smoke_query.yaml index cdc0523780..bcc6d8e525 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/smoke_query.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/smoke_query.yaml @@ -39,9 +39,6 @@ initial_prompt: | tree, sort, and limit, but use offset/start 2 for the next page. Query 3 must filter active equals true, sort by score descending, and use a limit of 4. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between - planning and implementation. Build the complete flow end-to-end in a - single pass. The Flow is not complete until `uip maestro flow validate` passes. success_criteria: diff --git a/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/smoke_update.yaml b/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/smoke_update.yaml index 30189f76ab..0dfae2ad9f 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/smoke_update.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/smoke_update.yaml @@ -24,9 +24,6 @@ initial_prompt: | 3. Retrieves the record by Id to confirm. 4. Deletes the record. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between - planning and implementation. Build the complete flow end-to-end in a - single pass. The Flow is not complete until `uip maestro flow validate` passes. success_criteria: diff --git a/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/smoke_update_existing_flow.yaml b/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/smoke_update_existing_flow.yaml index 5239218ff2..5d64379e88 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/smoke_update_existing_flow.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/smoke_update_existing_flow.yaml @@ -31,9 +31,6 @@ initial_prompt: | 3. Then revert both changes so the activity is back to no filter, no ordering, and its original settings. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between - planning and implementation. Build the complete flow end-to-end in a - single pass. The flow is not complete until the final state matches the initial state on those three fields and `uip maestro flow validate` passes. diff --git a/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/trigger_lifecycle.yaml b/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/trigger_lifecycle.yaml index 525edc9e8b..15d3e66620 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/trigger_lifecycle.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/trigger_lifecycle.yaml @@ -35,9 +35,6 @@ initial_prompt: | For both flows, keep the configured connection and correct entity on every trigger and Data Fabric activity. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between - planning and implementation. Build the complete flows end-to-end in a - single pass. Validate every final flow file. success_criteria: diff --git a/tests/tasks/uipath-maestro-flow/connector_features/drive_to_slack.yaml b/tests/tasks/uipath-maestro-flow/connector_features/drive_to_slack.yaml index 3bf8394980..b6dc1a1523 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/drive_to_slack.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/drive_to_slack.yaml @@ -20,7 +20,7 @@ run_limits: task_timeout: 1500 max_turns: 120 turn_timeout: 1500 -initial_prompt: "Create a new Flow project called \"DriveToSlackTest\" with a manual trigger.\nIt should download a file from Google Drive and then post that file into a \nSlack channel named \"coding-agent-testing\" using Slack.\nWire the file output of the download operation to the channel input of the slack send file. \nValidate the final flow file.\nFor the Google Drive Download File activity, use the first available file.\nFor Send File, send as `user`.\nFor google drive, use the connection name 'is.sandboxes.test@gmail.com', and for slack\nuse the connection name 'is-sandboxes'.\n\nDo NOT ask for approval, confirmation, or feedback. Do NOT pause between planning and implementation. Build the complete flow end-to-end in a single pass.\n" +initial_prompt: "Create a new Flow project called \"DriveToSlackTest\" with a manual trigger.\nIt should download a file from Google Drive and then post that file into a \nSlack channel named \"coding-agent-testing\" using Slack.\nWire the file output of the download operation to the channel input of the slack send file. \nValidate the final flow file.\nFor the Google Drive Download File activity, use the first available file.\nFor Send File, send as `user`.\nFor google drive, use the connection name 'is.sandboxes.test@gmail.com', and for slack\nuse the connection name 'is-sandboxes'.\n" success_criteria: - type: command_executed description: uip maestro flow validate was called diff --git a/tests/tasks/uipath-maestro-flow/connector_features/generic_dynamic_node/generic_dynamic_node.yaml b/tests/tasks/uipath-maestro-flow/connector_features/generic_dynamic_node/generic_dynamic_node.yaml index a25f18494f..63cee0e53e 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/generic_dynamic_node/generic_dynamic_node.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/generic_dynamic_node/generic_dynamic_node.yaml @@ -38,9 +38,7 @@ initial_prompt: | Use the connection present in Shared/uipath-maestro-flow. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between - planning and implementation. Build, validate, and run end-to-end in a single - pass. Do NOT substitute a mock for the + Do NOT substitute a mock for the connector. success_criteria: diff --git a/tests/tasks/uipath-maestro-flow/connector_features/jdbc_databricks_query/jdbc_databricks_query.yaml b/tests/tasks/uipath-maestro-flow/connector_features/jdbc_databricks_query/jdbc_databricks_query.yaml index 82db43ecfb..0ce792e8bb 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/jdbc_databricks_query/jdbc_databricks_query.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/jdbc_databricks_query/jdbc_databricks_query.yaml @@ -52,9 +52,6 @@ initial_prompt: | Validate the final flow file and debug it. - Do NOT ask for approval, confirmation, or feedback. - Do NOT pause between planning and implementation. - success_criteria: - type: command_executed description: "uip maestro flow validate was called" diff --git a/tests/tasks/uipath-maestro-flow/connector_features/non-catalog-http-fallback/non_catalog_http_fallback.yaml b/tests/tasks/uipath-maestro-flow/connector_features/non-catalog-http-fallback/non_catalog_http_fallback.yaml index 44e7d74bb8..d3a678ca79 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/non-catalog-http-fallback/non_catalog_http_fallback.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/non-catalog-http-fallback/non_catalog_http_fallback.yaml @@ -28,8 +28,6 @@ initial_prompt: | Validate the final flow file. Use `--output json` on every `uip` command whose output you parse. - Do NOT ask for approval, confirmation, or feedback — build and validate the - flow end-to-end in a single pass. success_criteria: - type: command_executed diff --git a/tests/tasks/uipath-maestro-flow/connector_features/paginated_reference_lookup.yaml b/tests/tasks/uipath-maestro-flow/connector_features/paginated_reference_lookup.yaml index bcdcbdb0d5..6e8f71551d 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/paginated_reference_lookup.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/paginated_reference_lookup.yaml @@ -34,8 +34,6 @@ initial_prompt: | For Slack, use the connection name `is-sandboxes`. If more than one connection matches that name, use any enabled one. Validate the flow when you're done. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between - planning and implementation. Build the complete flow end-to-end in a single pass. success_criteria: - type: command_executed diff --git a/tests/tasks/uipath-maestro-flow/connector_features/slack-http-fallback/slack_http_fallback.yaml b/tests/tasks/uipath-maestro-flow/connector_features/slack-http-fallback/slack_http_fallback.yaml index dcedeb9bba..f029f84495 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/slack-http-fallback/slack_http_fallback.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/slack-http-fallback/slack_http_fallback.yaml @@ -25,8 +25,6 @@ initial_prompt: | Validate the final flow file. Once it validates, debug it. Use `--output json` on every `uip` command whose output you parse. - Do NOT ask for approval, confirmation, or feedback — build, validate, and - debug the flow end-to-end in a single pass. success_criteria: - type: command_executed diff --git a/tests/tasks/uipath-maestro-flow/connector_features/testmanager_attachments/testmanager_attachments.yaml b/tests/tasks/uipath-maestro-flow/connector_features/testmanager_attachments/testmanager_attachments.yaml index 7c8a562ce1..37ad851663 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/testmanager_attachments/testmanager_attachments.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/testmanager_attachments/testmanager_attachments.yaml @@ -29,9 +29,6 @@ initial_prompt: | The task is not complete until validation passed for the flow. - Do NOT ask for approval, confirmation, or feedback. - Do NOT pause between planning and implementation. - success_criteria: - type: command_executed description: "Agent refreshed the node manifest (`uip maestro flow registry pull`) before searching" diff --git a/tests/tasks/uipath-maestro-flow/connector_features/testmanager_crud_grounded/testmanager_crud_grounded.yaml b/tests/tasks/uipath-maestro-flow/connector_features/testmanager_crud_grounded/testmanager_crud_grounded.yaml index 28df13cfa5..96baf70526 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/testmanager_crud_grounded/testmanager_crud_grounded.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/testmanager_crud_grounded/testmanager_crud_grounded.yaml @@ -34,9 +34,6 @@ initial_prompt: | The task is not complete until validation passed for the flow. - Do NOT ask for approval, confirmation, or feedback. - Do NOT pause between planning and implementation. - success_criteria: - type: command_executed description: "Agent refreshed the node manifest (`uip maestro flow registry pull`)" diff --git a/tests/tasks/uipath-maestro-flow/connector_features/testmanager_execution_results/testmanager_execution_results.yaml b/tests/tasks/uipath-maestro-flow/connector_features/testmanager_execution_results/testmanager_execution_results.yaml index 92aecc8cd6..4e34afb35f 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/testmanager_execution_results/testmanager_execution_results.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/testmanager_execution_results/testmanager_execution_results.yaml @@ -34,9 +34,6 @@ initial_prompt: | The task is not complete until validation passed for the flow. - Do NOT ask for approval, confirmation, or feedback. - Do NOT pause between planning and implementation. - success_criteria: - type: command_executed description: "Agent refreshed the node manifest (`uip maestro flow registry pull`) before searching" diff --git a/tests/tasks/uipath-maestro-flow/connector_features/testmanager_generic_records/testmanager_generic_records.yaml b/tests/tasks/uipath-maestro-flow/connector_features/testmanager_generic_records/testmanager_generic_records.yaml index dc8caf370e..9904502ac5 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/testmanager_generic_records/testmanager_generic_records.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/testmanager_generic_records/testmanager_generic_records.yaml @@ -30,9 +30,6 @@ initial_prompt: | The task is not complete until validation passed for the flow. - Do NOT ask for approval, confirmation, or feedback. - Do NOT pause between planning and implementation. - success_criteria: - type: command_executed description: "Agent refreshed the node manifest (`uip maestro flow registry pull`) before searching" diff --git a/tests/tasks/uipath-maestro-flow/connector_features/testmanager_requirement_lifecycle/testmanager_requirement_lifecycle.yaml b/tests/tasks/uipath-maestro-flow/connector_features/testmanager_requirement_lifecycle/testmanager_requirement_lifecycle.yaml index b02e5261b1..d85af72684 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/testmanager_requirement_lifecycle/testmanager_requirement_lifecycle.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/testmanager_requirement_lifecycle/testmanager_requirement_lifecycle.yaml @@ -31,9 +31,6 @@ initial_prompt: | The task is not complete until validation passed for the flow. - Do NOT ask for approval, confirmation, or feedback. - Do NOT pause between planning and implementation. - success_criteria: - type: command_executed description: "Agent refreshed the node manifest (`uip maestro flow registry pull`) before searching" diff --git a/tests/tasks/uipath-maestro-flow/connector_features/testmanager_testcase_lifecycle/testmanager_testcase_lifecycle.yaml b/tests/tasks/uipath-maestro-flow/connector_features/testmanager_testcase_lifecycle/testmanager_testcase_lifecycle.yaml index 83716c6a24..d2f04c4e42 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/testmanager_testcase_lifecycle/testmanager_testcase_lifecycle.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/testmanager_testcase_lifecycle/testmanager_testcase_lifecycle.yaml @@ -30,9 +30,6 @@ initial_prompt: | The task is not complete until validation passed for the flow. - Do NOT ask for approval, confirmation, or feedback. - Do NOT pause between planning and implementation. - success_criteria: - type: command_executed description: "Agent refreshed the node manifest (`uip maestro flow registry pull`) before searching" diff --git a/tests/tasks/uipath-maestro-flow/connector_features/testmanager_testset_lifecycle/testmanager_testset_lifecycle.yaml b/tests/tasks/uipath-maestro-flow/connector_features/testmanager_testset_lifecycle/testmanager_testset_lifecycle.yaml index cd366f1db8..7f236037ea 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/testmanager_testset_lifecycle/testmanager_testset_lifecycle.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/testmanager_testset_lifecycle/testmanager_testset_lifecycle.yaml @@ -32,9 +32,6 @@ initial_prompt: | The task is not complete until validation passed for the flow. - Do NOT ask for approval, confirmation, or feedback. - Do NOT pause between planning and implementation. - success_criteria: - type: command_executed description: "Agent refreshed the node manifest (`uip maestro flow registry pull`) before searching" diff --git a/tests/tasks/uipath-maestro-flow/connector_trigger/trigger_with_filter.yaml b/tests/tasks/uipath-maestro-flow/connector_trigger/trigger_with_filter.yaml index 3f23be96b4..765d5cacf2 100644 --- a/tests/tasks/uipath-maestro-flow/connector_trigger/trigger_with_filter.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_trigger/trigger_with_filter.yaml @@ -16,7 +16,6 @@ initial_prompt: | `uip maestro flow node configure --detail ''` to configure that trigger. Use placeholders for any connection/folder/event ids you cannot resolve offline. - success_criteria: # `check_trigger_filter.py` locates trigger_detail.json (root or nested in the # flow project dir) and asserts it is the --detail object: valid JSON, no diff --git a/tests/tasks/uipath-maestro-flow/connector_trigger/webhook_waitfor_parallel.yaml b/tests/tasks/uipath-maestro-flow/connector_trigger/webhook_waitfor_parallel.yaml index 2189db7937..05bb446a77 100644 --- a/tests/tasks/uipath-maestro-flow/connector_trigger/webhook_waitfor_parallel.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_trigger/webhook_waitfor_parallel.yaml @@ -45,8 +45,6 @@ initial_prompt: | After building, validate the flow. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between planning and - implementation. Build the complete flow end-to-end in a single pass. Use `--output json` on `uip` commands. success_criteria: diff --git a/tests/tasks/uipath-maestro-flow/context-grounding/batch_transform/batch_transform.yaml b/tests/tasks/uipath-maestro-flow/context-grounding/batch_transform/batch_transform.yaml index a81538c530..a87c7063f7 100644 --- a/tests/tasks/uipath-maestro-flow/context-grounding/batch_transform/batch_transform.yaml +++ b/tests/tasks/uipath-maestro-flow/context-grounding/batch_transform/batch_transform.yaml @@ -42,8 +42,6 @@ initial_prompt: | `outputs.result.source: "=js:$vars..output"`. Do NOT run flow debug — just validate the flow. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between - planning and implementation. Build the complete flow end-to-end in a single pass. Before starting, follow the documented workflow steps exactly — in particular, the batch-transform plugin's `outputColumns` shape (array of `{ name, description }`). diff --git a/tests/tasks/uipath-maestro-flow/context-grounding/summarize/summarize.yaml b/tests/tasks/uipath-maestro-flow/context-grounding/summarize/summarize.yaml index 74c13bbf2a..ed2cecf25f 100644 --- a/tests/tasks/uipath-maestro-flow/context-grounding/summarize/summarize.yaml +++ b/tests/tasks/uipath-maestro-flow/context-grounding/summarize/summarize.yaml @@ -44,8 +44,6 @@ initial_prompt: | resolve to `undefined` at runtime. Do NOT run flow debug — just validate the flow. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between - planning and implementation. Build the complete flow end-to-end in a single pass. Before starting, follow the documented workflow steps exactly — in particular, note that the wire-level node type is `uipath.pattern.deep-rag` even though the canvas display name is "Summarize". diff --git a/tests/tasks/uipath-maestro-flow/e2e/escalation_jira_ticket/escalation_jira_ticket.yaml b/tests/tasks/uipath-maestro-flow/e2e/escalation_jira_ticket/escalation_jira_ticket.yaml index c79643e94f..857421e170 100644 --- a/tests/tasks/uipath-maestro-flow/e2e/escalation_jira_ticket/escalation_jira_ticket.yaml +++ b/tests/tasks/uipath-maestro-flow/e2e/escalation_jira_ticket/escalation_jira_ticket.yaml @@ -54,8 +54,7 @@ initial_prompt: | correlationId), and jiraIssueKey (the created issue's key). Validate the flow. Do NOT run or debug the flow — the grader executes it with the - seeded inputs. Do NOT ask for approval and do NOT pause; build end-to-end in a - single pass. Before starting, load the uipath-maestro-flow skill and follow its + seeded inputs.Before starting, load the uipath-maestro-flow skill and follow its workflow. success_criteria: diff --git a/tests/tasks/uipath-maestro-flow/e2e/escalation_orchestrator_paths/escalation_orchestrator_paths.yaml b/tests/tasks/uipath-maestro-flow/e2e/escalation_orchestrator_paths/escalation_orchestrator_paths.yaml index bb191b3e1a..9e44b5a465 100644 --- a/tests/tasks/uipath-maestro-flow/e2e/escalation_orchestrator_paths/escalation_orchestrator_paths.yaml +++ b/tests/tasks/uipath-maestro-flow/e2e/escalation_orchestrator_paths/escalation_orchestrator_paths.yaml @@ -82,8 +82,7 @@ initial_prompt: | Salesforce/Jira/Drive are NOT part of this task — the match status is an input. Validate the flow. Do NOT run or debug the flow — the grader executes it with - seeded inputs. Do NOT ask for approval and do NOT pause; build the flow directly. - Before starting, load the uipath-maestro-flow skill and follow its workflow. + seeded inputs. Before starting, load the uipath-maestro-flow skill and follow its workflow. success_criteria: - type: command_executed diff --git a/tests/tasks/uipath-maestro-flow/e2e/escalation_slack_alert/escalation_slack_alert.yaml b/tests/tasks/uipath-maestro-flow/e2e/escalation_slack_alert/escalation_slack_alert.yaml index 0a8076c810..e455c7d0cf 100644 --- a/tests/tasks/uipath-maestro-flow/e2e/escalation_slack_alert/escalation_slack_alert.yaml +++ b/tests/tasks/uipath-maestro-flow/e2e/escalation_slack_alert/escalation_slack_alert.yaml @@ -56,10 +56,8 @@ initial_prompt: | (from the Slack Send Message node's output) Validate the flow. Do NOT run or debug the flow — the grader executes it with - seeded inputs. Do NOT ask for approval, confirmation, or feedback, and do NOT - pause between planning and implementation. Build the complete flow end-to-end - in a single pass. Before starting, load the uipath-maestro-flow skill and - follow its workflow. + seeded inputs. Before starting, load the uipath-maestro-flow skill and follow + its workflow. success_criteria: # ── The agent validated the flow (convention adherence) ───────────────── diff --git a/tests/tasks/uipath-maestro-flow/edit/add_node/add_node.yaml b/tests/tasks/uipath-maestro-flow/edit/add_node/add_node.yaml index 96ec38a796..7e1c927f57 100644 --- a/tests/tasks/uipath-maestro-flow/edit/add_node/add_node.yaml +++ b/tests/tasks/uipath-maestro-flow/edit/add_node/add_node.yaml @@ -18,7 +18,6 @@ initial_prompt: | Add a script node called "convertToCelsius" between getWeather and formatSummary that converts the temp from F to C. Return both values. Also update formatSummary to include the Celsius value in its summary string. Validate the flow. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between planning and implementation. Build the complete flow end-to-end in a single pass. success_criteria: # ── Flow file validity ───────────────────────────────────────────────── diff --git a/tests/tasks/uipath-maestro-flow/edit/add_output/add_output.yaml b/tests/tasks/uipath-maestro-flow/edit/add_output/add_output.yaml index 306bf90b46..3c35b97fd0 100644 --- a/tests/tasks/uipath-maestro-flow/edit/add_output/add_output.yaml +++ b/tests/tasks/uipath-maestro-flow/edit/add_output/add_output.yaml @@ -17,7 +17,6 @@ initial_prompt: | Add a "location" field with value "Bellevue, WA" to the summary output of both end nodes. Validate the flow. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between planning and implementation. Build the complete flow end-to-end in a single pass. success_criteria: - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/edit/group_to_subflow/group_to_subflow.yaml b/tests/tasks/uipath-maestro-flow/edit/group_to_subflow/group_to_subflow.yaml index 73aab55567..df70d565a0 100644 --- a/tests/tasks/uipath-maestro-flow/edit/group_to_subflow/group_to_subflow.yaml +++ b/tests/tasks/uipath-maestro-flow/edit/group_to_subflow/group_to_subflow.yaml @@ -19,7 +19,6 @@ initial_prompt: | Move getWeather and formatSummary into a subflow called "fetchAndFormat". The subflow should output the temperatureF value. The rest of the main flow (decision + end nodes) stays but reads from the subflow output. Validate the flow. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between planning and implementation. Build the complete flow end-to-end in a single pass. success_criteria: # ── Flow file validity ───────────────────────────────────────────────── diff --git a/tests/tasks/uipath-maestro-flow/edit/move_node/move_node.yaml b/tests/tasks/uipath-maestro-flow/edit/move_node/move_node.yaml index ddc5bb7045..b92f2e9466 100644 --- a/tests/tasks/uipath-maestro-flow/edit/move_node/move_node.yaml +++ b/tests/tasks/uipath-maestro-flow/edit/move_node/move_node.yaml @@ -17,7 +17,6 @@ initial_prompt: | Rearrange the flow so the decision happens right after the HTTP call. Both the true and false branches should merge back into formatSummary, which then goes to a single end node. formatSummary should pick the message ('nice day' or 'bring a jacket') based on which branch the decision took. Validate the flow. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between planning and implementation. Build the complete flow end-to-end in a single pass. success_criteria: - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/edit/remove_node/remove_node.yaml b/tests/tasks/uipath-maestro-flow/edit/remove_node/remove_node.yaml index 423a266a1a..8d610ce21f 100644 --- a/tests/tasks/uipath-maestro-flow/edit/remove_node/remove_node.yaml +++ b/tests/tasks/uipath-maestro-flow/edit/remove_node/remove_node.yaml @@ -18,7 +18,6 @@ initial_prompt: | Remove the formatSummary script node. The decision and end nodes should read temperature directly from the HTTP response instead. Validate the flow. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between planning and implementation. Build the complete flow end-to-end in a single pass. success_criteria: # ── Flow file validity ───────────────────────────────────────────────── diff --git a/tests/tasks/uipath-maestro-flow/edit/update_node/update_node.yaml b/tests/tasks/uipath-maestro-flow/edit/update_node/update_node.yaml index 7a150b6785..a46f6b5871 100644 --- a/tests/tasks/uipath-maestro-flow/edit/update_node/update_node.yaml +++ b/tests/tasks/uipath-maestro-flow/edit/update_node/update_node.yaml @@ -19,7 +19,6 @@ initial_prompt: | otherwise the message field should be 'go home'. Validate the flow. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between planning and implementation. Build the complete flow end-to-end in a single pass. success_criteria: # ── Flow file validity ───────────────────────────────────────────────── diff --git a/tests/tasks/uipath-maestro-flow/evaluate/inline_agent_eval/inline_agent_eval.yaml b/tests/tasks/uipath-maestro-flow/evaluate/inline_agent_eval/inline_agent_eval.yaml index 1d08b8c820..85d51cf57b 100644 --- a/tests/tasks/uipath-maestro-flow/evaluate/inline_agent_eval/inline_agent_eval.yaml +++ b/tests/tasks/uipath-maestro-flow/evaluate/inline_agent_eval/inline_agent_eval.yaml @@ -54,9 +54,7 @@ initial_prompt: | - Local CRUD only. Do NOT call `uip login`. Do NOT call `uip solution upload`. Do NOT call `uip maestro flow eval run *`. Do NOT run `uip maestro flow debug`. - - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between - planning and implementation. Build it end-to-end in a single pass. - + - success_criteria: - type: command_executed description: "Agent scaffolded the inline agent with uip agent init --inline-in-flow" diff --git a/tests/tasks/uipath-maestro-flow/hitl/smoke_01_hitl_node_placed.yaml b/tests/tasks/uipath-maestro-flow/hitl/smoke_01_hitl_node_placed.yaml index f2d648f834..2c04c7b849 100644 --- a/tests/tasks/uipath-maestro-flow/hitl/smoke_01_hitl_node_placed.yaml +++ b/tests/tasks/uipath-maestro-flow/hitl/smoke_01_hitl_node_placed.yaml @@ -40,7 +40,6 @@ success_criteria: weight: 3.0 pass_threshold: 1.0 - - type: run_command description: "uip flow validate passes" command: "python3 $SKILLS_REPO_PATH/tests/tasks/uipath-maestro-flow/_shared/validate_flow.py" diff --git a/tests/tasks/uipath-maestro-flow/ixp/e2e_01_invoice_extraction_greenfield.yaml b/tests/tasks/uipath-maestro-flow/ixp/e2e_01_invoice_extraction_greenfield.yaml index 1ad9e68f4e..f60f2b7a4f 100644 --- a/tests/tasks/uipath-maestro-flow/ixp/e2e_01_invoice_extraction_greenfield.yaml +++ b/tests/tasks/uipath-maestro-flow/ixp/e2e_01_invoice_extraction_greenfield.yaml @@ -30,7 +30,6 @@ initial_prompt: | Validate the final flow file — the task is not complete until `uip maestro flow validate` passes. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between planning and implementation. Build the complete flow end-to-end in a single pass. success_criteria: - type: run_command description: uip maestro flow validate passes on the generated flow file diff --git a/tests/tasks/uipath-maestro-flow/ixp/e2e_03_project_creation_handoff/e2e_03_project_creation_handoff.yaml b/tests/tasks/uipath-maestro-flow/ixp/e2e_03_project_creation_handoff/e2e_03_project_creation_handoff.yaml index 0fffd7d922..d4c537699f 100644 --- a/tests/tasks/uipath-maestro-flow/ixp/e2e_03_project_creation_handoff/e2e_03_project_creation_handoff.yaml +++ b/tests/tasks/uipath-maestro-flow/ixp/e2e_03_project_creation_handoff/e2e_03_project_creation_handoff.yaml @@ -97,9 +97,7 @@ initial_prompt: | - The `uip` CLI is already available and the runner is logged in. - Use `--output json` on all uip commands. - Do NOT run `uip maestro flow debug` or `uip maestro flow deploy`. - - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between - planning and implementation. Build it end-to-end in a single pass. - + - success_criteria: - type: skill_triggered description: "Agent invoked the uipath-maestro-flow Skill" diff --git a/tests/tasks/uipath-maestro-flow/multi_node/bellevue_weather/bellevue_weather.yaml b/tests/tasks/uipath-maestro-flow/multi_node/bellevue_weather/bellevue_weather.yaml index 6905a1a767..e97fc21017 100644 --- a/tests/tasks/uipath-maestro-flow/multi_node/bellevue_weather/bellevue_weather.yaml +++ b/tests/tasks/uipath-maestro-flow/multi_node/bellevue_weather/bellevue_weather.yaml @@ -17,7 +17,6 @@ initial_prompt: | Use a Managed HTTP Request node (core.action.http.v2) to call the open-meteo API directly — do not use an Integration Service connector. Validate the flow. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between planning and implementation. Build the complete flow end-to-end in a single pass. success_criteria: # ── Flow file validity ───────────────────────────────────────────────── diff --git a/tests/tasks/uipath-maestro-flow/multi_node/calculator/calculator.yaml b/tests/tasks/uipath-maestro-flow/multi_node/calculator/calculator.yaml index 5a6a109a34..077b79a4e8 100644 --- a/tests/tasks/uipath-maestro-flow/multi_node/calculator/calculator.yaml +++ b/tests/tasks/uipath-maestro-flow/multi_node/calculator/calculator.yaml @@ -14,7 +14,6 @@ initial_prompt: | returned as an output variable. Validate the flow. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between planning and implementation. Build the complete flow end-to-end in a single pass. success_criteria: # ── Flow file validity ───────────────────────────────────────────────── diff --git a/tests/tasks/uipath-maestro-flow/multi_node/customer_escalation/customer_escalation.yaml b/tests/tasks/uipath-maestro-flow/multi_node/customer_escalation/customer_escalation.yaml index c384d2849e..cb7480a632 100644 --- a/tests/tasks/uipath-maestro-flow/multi_node/customer_escalation/customer_escalation.yaml +++ b/tests/tasks/uipath-maestro-flow/multi_node/customer_escalation/customer_escalation.yaml @@ -32,9 +32,6 @@ initial_prompt: | reply to the sender via Outlook with the ticket details. Validate the flow. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between - planning and implementation. Build the complete flow end-to-end in a - single pass. success_criteria: # ── Flow file validity ──────────────────────────────────────────────── diff --git a/tests/tasks/uipath-maestro-flow/multi_node/dice_roller/dice_roller.yaml b/tests/tasks/uipath-maestro-flow/multi_node/dice_roller/dice_roller.yaml index dd75fa2522..aac5bf221d 100644 --- a/tests/tasks/uipath-maestro-flow/multi_node/dice_roller/dice_roller.yaml +++ b/tests/tasks/uipath-maestro-flow/multi_node/dice_roller/dice_roller.yaml @@ -15,7 +15,6 @@ initial_prompt: | die and outputs the result. Validate the flow. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between planning and implementation. Build the complete flow end-to-end in a single pass. success_criteria: # ── Flow file validity ───────────────────────────────────────────────── diff --git a/tests/tasks/uipath-maestro-flow/multi_node/feet_inches/feet_inches.yaml b/tests/tasks/uipath-maestro-flow/multi_node/feet_inches/feet_inches.yaml index d64eab27b1..ca028f38ef 100644 --- a/tests/tasks/uipath-maestro-flow/multi_node/feet_inches/feet_inches.yaml +++ b/tests/tasks/uipath-maestro-flow/multi_node/feet_inches/feet_inches.yaml @@ -20,8 +20,6 @@ initial_prompt: | ("f2i", "i2f", "y2f"); each branch applies the corresponding conversion (in a Script node) and converges on a single End node that returns the result. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between planning and implementation. Build the complete flow end-to-end in a single pass. - success_criteria: # ── Flow file validity ───────────────────────────────────────────────── - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/multi_node/loop_multiply/loop_multiply.yaml b/tests/tasks/uipath-maestro-flow/multi_node/loop_multiply/loop_multiply.yaml index ab64ffd4da..798d464e07 100644 --- a/tests/tasks/uipath-maestro-flow/multi_node/loop_multiply/loop_multiply.yaml +++ b/tests/tasks/uipath-maestro-flow/multi_node/loop_multiply/loop_multiply.yaml @@ -14,7 +14,6 @@ initial_prompt: | numbers [13, 15, 17] together using a Loop node and returns the product. Validate the flow. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between planning and implementation. Build the complete flow end-to-end in a single pass. success_criteria: # ── Flow file validity ───────────────────────────────────────────────── diff --git a/tests/tasks/uipath-maestro-flow/multi_node/multi_city_weather/multi_city_weather.yaml b/tests/tasks/uipath-maestro-flow/multi_node/multi_city_weather/multi_city_weather.yaml index 57e8b3e480..a98be8c1ce 100644 --- a/tests/tasks/uipath-maestro-flow/multi_node/multi_city_weather/multi_city_weather.yaml +++ b/tests/tasks/uipath-maestro-flow/multi_node/multi_city_weather/multi_city_weather.yaml @@ -14,7 +14,6 @@ initial_prompt: | Use a Managed HTTP Request node (core.action.http.v2) to call the open-meteo API directly — do not use an Integration Service connector. Validate the flow. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between planning and implementation. Build the complete flow end-to-end in a single pass. success_criteria: - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/multi_node/reading_list/reading_list.yaml b/tests/tasks/uipath-maestro-flow/multi_node/reading_list/reading_list.yaml index 305eccd2e3..40a4bc0efb 100644 --- a/tests/tasks/uipath-maestro-flow/multi_node/reading_list/reading_list.yaml +++ b/tests/tasks/uipath-maestro-flow/multi_node/reading_list/reading_list.yaml @@ -41,7 +41,6 @@ initial_prompt: | an inline `=js:[...]` array or `=js:$vars.` expression. Validate the flow. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between planning and implementation. Build the complete flow end-to-end in a single pass. success_criteria: # ── Flow file validity ───────────────────────────────────────────────── diff --git a/tests/tasks/uipath-maestro-flow/multi_node/slack_channel_description/slack_channel_description.yaml b/tests/tasks/uipath-maestro-flow/multi_node/slack_channel_description/slack_channel_description.yaml index 70df43e92c..c60193589b 100644 --- a/tests/tasks/uipath-maestro-flow/multi_node/slack_channel_description/slack_channel_description.yaml +++ b/tests/tasks/uipath-maestro-flow/multi_node/slack_channel_description/slack_channel_description.yaml @@ -15,7 +15,6 @@ initial_prompt: | the channel description of #office-bellevue and outputs it. Validate the flow. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between planning and implementation. Build the complete flow end-to-end in a single pass. success_criteria: # ── Flow file validity ───────────────────────────────────────────────── diff --git a/tests/tasks/uipath-maestro-flow/multi_node/slack_weather_pipeline/slack_weather_pipeline.yaml b/tests/tasks/uipath-maestro-flow/multi_node/slack_weather_pipeline/slack_weather_pipeline.yaml index 21f962aa1d..2f1b1df38d 100644 --- a/tests/tasks/uipath-maestro-flow/multi_node/slack_weather_pipeline/slack_weather_pipeline.yaml +++ b/tests/tasks/uipath-maestro-flow/multi_node/slack_weather_pipeline/slack_weather_pipeline.yaml @@ -16,7 +16,6 @@ initial_prompt: | All connections for this task live in Orchestrator folder `Shared/uipath-maestro-flow`. Validate the flow, then run `uip maestro flow debug` and iterate until it completes without incidents. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between planning and implementation. Build the complete flow end-to-end in a single pass. success_criteria: - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/multi_node/wiki_pageviews/wiki_pageviews.yaml b/tests/tasks/uipath-maestro-flow/multi_node/wiki_pageviews/wiki_pageviews.yaml index bfcf65dcd1..2865800b91 100644 --- a/tests/tasks/uipath-maestro-flow/multi_node/wiki_pageviews/wiki_pageviews.yaml +++ b/tests/tasks/uipath-maestro-flow/multi_node/wiki_pageviews/wiki_pageviews.yaml @@ -32,8 +32,6 @@ initial_prompt: | does not exist and the API returns an error), the flow must instead return the literal string `Article not found`. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between planning and implementation. Build the complete flow end-to-end in a single pass. - success_criteria: # ── Flow file validity ───────────────────────────────────────────────── - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/single_node/api_workflow/api_workflow.yaml b/tests/tasks/uipath-maestro-flow/single_node/api_workflow/api_workflow.yaml index eebfd349f8..232744e608 100644 --- a/tests/tasks/uipath-maestro-flow/single_node/api_workflow/api_workflow.yaml +++ b/tests/tasks/uipath-maestro-flow/single_node/api_workflow/api_workflow.yaml @@ -14,7 +14,6 @@ initial_prompt: | API workflow with the name 'tomasz' and returns his age as an output. Validate the flow. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between planning and implementation. Build the complete flow end-to-end in a single pass. success_criteria: # ── Flow file validity ───────────────────────────────────────────────── diff --git a/tests/tasks/uipath-maestro-flow/single_node/coded_agent/coded_agent.yaml b/tests/tasks/uipath-maestro-flow/single_node/coded_agent/coded_agent.yaml index 638e7ccf66..f053107444 100644 --- a/tests/tasks/uipath-maestro-flow/single_node/coded_agent/coded_agent.yaml +++ b/tests/tasks/uipath-maestro-flow/single_node/coded_agent/coded_agent.yaml @@ -26,7 +26,6 @@ initial_prompt: | number of people who approved as an integer flow output. Validate the flow. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between planning and implementation. Build the complete flow end-to-end in a single pass. success_criteria: - type: run_command description: uip maestro flow validate passes on the flow file diff --git a/tests/tasks/uipath-maestro-flow/single_node/decision/decision.yaml b/tests/tasks/uipath-maestro-flow/single_node/decision/decision.yaml index fb3a14b056..0a6ca379a0 100644 --- a/tests/tasks/uipath-maestro-flow/single_node/decision/decision.yaml +++ b/tests/tasks/uipath-maestro-flow/single_node/decision/decision.yaml @@ -17,7 +17,6 @@ initial_prompt: | the flow should output "warm". Otherwise it should output "cool". Validate the flow. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between planning and implementation. Build the complete flow end-to-end in a single pass. success_criteria: # -- Flow file validity -- diff --git a/tests/tasks/uipath-maestro-flow/single_node/delay/delay.yaml b/tests/tasks/uipath-maestro-flow/single_node/delay/delay.yaml index c8b35a7534..f33d200df1 100644 --- a/tests/tasks/uipath-maestro-flow/single_node/delay/delay.yaml +++ b/tests/tasks/uipath-maestro-flow/single_node/delay/delay.yaml @@ -28,8 +28,6 @@ initial_prompt: | the delay node id, and one whose `sourceNodeId` is the delay node id). Do NOT run flow debug — just validate the flow. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between - planning and implementation. Build the complete flow end-to-end in a single pass. Before starting, follow the documented workflow steps exactly — in particular, the delay plugin's `inputs` shape (`timerType` + `timerPreset`). diff --git a/tests/tasks/uipath-maestro-flow/single_node/file_attachment/file_attachment.yaml b/tests/tasks/uipath-maestro-flow/single_node/file_attachment/file_attachment.yaml index eb10dbc135..cc00850d8d 100644 --- a/tests/tasks/uipath-maestro-flow/single_node/file_attachment/file_attachment.yaml +++ b/tests/tasks/uipath-maestro-flow/single_node/file_attachment/file_attachment.yaml @@ -30,9 +30,7 @@ initial_prompt: | file to the input. The task is not complete until the debug run reports `finalStatus: "Completed"` and the output reflects the attached file's name. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between - planning and implementation. Build, validate, and run end-to-end in a single - pass. Before starting, follow the documented + Before starting, follow the documented workflow steps exactly — in particular the file-attachment binding pre-flight and the requirement to run `solution resources refresh` before debug. diff --git a/tests/tasks/uipath-maestro-flow/single_node/lowcode_agent/lowcode_agent.yaml b/tests/tasks/uipath-maestro-flow/single_node/lowcode_agent/lowcode_agent.yaml index 415736357b..2a263de6f6 100644 --- a/tests/tasks/uipath-maestro-flow/single_node/lowcode_agent/lowcode_agent.yaml +++ b/tests/tasks/uipath-maestro-flow/single_node/lowcode_agent/lowcode_agent.yaml @@ -21,7 +21,6 @@ initial_prompt: | in rather than creating a new one. Validate the flow. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between planning and implementation. Build the complete flow end-to-end in a single pass. success_criteria: # ── Flow file validity ───────────────────────────────────────────────── diff --git a/tests/tasks/uipath-maestro-flow/single_node/openmeteo_weather/openmeteo_weather.yaml b/tests/tasks/uipath-maestro-flow/single_node/openmeteo_weather/openmeteo_weather.yaml index a911f49ab7..bada110238 100644 --- a/tests/tasks/uipath-maestro-flow/single_node/openmeteo_weather/openmeteo_weather.yaml +++ b/tests/tasks/uipath-maestro-flow/single_node/openmeteo_weather/openmeteo_weather.yaml @@ -32,9 +32,7 @@ initial_prompt: | The task is not complete until `uip maestro flow validate` passes. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between - planning and implementation. Build and validate the complete flow end-to-end - in a single pass. Do NOT substitute a mock or a manual + Do NOT substitute a mock or a manual HTTP request for the connector. success_criteria: diff --git a/tests/tasks/uipath-maestro-flow/single_node/outlook_trigger_inbox/outlook_trigger_inbox.yaml b/tests/tasks/uipath-maestro-flow/single_node/outlook_trigger_inbox/outlook_trigger_inbox.yaml index 3f27f03f6c..c9c961551f 100644 --- a/tests/tasks/uipath-maestro-flow/single_node/outlook_trigger_inbox/outlook_trigger_inbox.yaml +++ b/tests/tasks/uipath-maestro-flow/single_node/outlook_trigger_inbox/outlook_trigger_inbox.yaml @@ -26,8 +26,6 @@ initial_prompt: | from that Script node's output. Validate the flow. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between - planning and implementation. Build the complete flow end-to-end in a single pass. Do not reuse any folder ID you may have seen before. diff --git a/tests/tasks/uipath-maestro-flow/single_node/outlook_waitfor_email/outlook_waitfor_email.yaml b/tests/tasks/uipath-maestro-flow/single_node/outlook_waitfor_email/outlook_waitfor_email.yaml index 52fc561599..dd43b87217 100644 --- a/tests/tasks/uipath-maestro-flow/single_node/outlook_waitfor_email/outlook_waitfor_email.yaml +++ b/tests/tasks/uipath-maestro-flow/single_node/outlook_waitfor_email/outlook_waitfor_email.yaml @@ -41,9 +41,6 @@ initial_prompt: | Validate the flow. Do not debug the flow. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between - planning and implementation. Build the complete flow end-to-end in a single - pass. success_criteria: # ── Structural: flow builds and validates ────────────────────────────── diff --git a/tests/tasks/uipath-maestro-flow/single_node/rpa/rpa.yaml b/tests/tasks/uipath-maestro-flow/single_node/rpa/rpa.yaml index 54904b6a08..52ddce3fe3 100644 --- a/tests/tasks/uipath-maestro-flow/single_node/rpa/rpa.yaml +++ b/tests/tasks/uipath-maestro-flow/single_node/rpa/rpa.yaml @@ -16,7 +16,6 @@ initial_prompt: | return it as an output. Validate the flow. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between planning and implementation. Build the complete flow end-to-end in a single pass. success_criteria: # ── Flow file validity ───────────────────────────────────────────────── diff --git a/tests/tasks/uipath-maestro-flow/single_node/subflow/subflow.yaml b/tests/tasks/uipath-maestro-flow/single_node/subflow/subflow.yaml index 7c359f77f6..50457c204a 100644 --- a/tests/tasks/uipath-maestro-flow/single_node/subflow/subflow.yaml +++ b/tests/tasks/uipath-maestro-flow/single_node/subflow/subflow.yaml @@ -18,7 +18,6 @@ initial_prompt: | Return the reversed string as an output. Validate the flow. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between planning and implementation. Build the complete flow end-to-end in a single pass. success_criteria: # -- Flow file validity -- diff --git a/tests/tasks/uipath-maestro-flow/single_node/switch/switch.yaml b/tests/tasks/uipath-maestro-flow/single_node/switch/switch.yaml index b35e47ae0b..25a23cd219 100644 --- a/tests/tasks/uipath-maestro-flow/single_node/switch/switch.yaml +++ b/tests/tasks/uipath-maestro-flow/single_node/switch/switch.yaml @@ -20,7 +20,6 @@ initial_prompt: | The flow should branch into separate cases for each quarter value. Validate the flow. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between planning and implementation. Build the complete flow end-to-end in a single pass. success_criteria: # -- Flow file validity -- diff --git a/tests/tasks/uipath-maestro-flow/single_node/terminate/terminate.yaml b/tests/tasks/uipath-maestro-flow/single_node/terminate/terminate.yaml index 33a2a66d03..edee904217 100644 --- a/tests/tasks/uipath-maestro-flow/single_node/terminate/terminate.yaml +++ b/tests/tasks/uipath-maestro-flow/single_node/terminate/terminate.yaml @@ -21,7 +21,6 @@ initial_prompt: | Both branches should start at the same time from the trigger node. Validate the flow. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between planning and implementation. Build the complete flow end-to-end in a single pass. success_criteria: # -- Flow file validity -- diff --git a/tests/tasks/uipath-maestro-flow/single_node/transform_filter/transform_filter.yaml b/tests/tasks/uipath-maestro-flow/single_node/transform_filter/transform_filter.yaml index d1cdf46d46..f6ff66e8c6 100644 --- a/tests/tasks/uipath-maestro-flow/single_node/transform_filter/transform_filter.yaml +++ b/tests/tasks/uipath-maestro-flow/single_node/transform_filter/transform_filter.yaml @@ -38,9 +38,7 @@ initial_prompt: | `core.action.transform.filter` — do not hardcode a guessed value. Do NOT run flow debug — just validate the flow. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between - planning and implementation. Build the complete flow end-to-end in a single - pass. Before starting, follow the documented workflow steps exactly — in particular the transform plugin's filter + Before starting, follow the documented workflow steps exactly — in particular the transform plugin's filter node shape (`operations[].type == "filter"` with `config.filters`). success_criteria: diff --git a/tests/tasks/uipath-maestro-flow/single_node/transform_group_by/transform_group_by.yaml b/tests/tasks/uipath-maestro-flow/single_node/transform_group_by/transform_group_by.yaml index 1e9feaa5dd..1350ab1932 100644 --- a/tests/tasks/uipath-maestro-flow/single_node/transform_group_by/transform_group_by.yaml +++ b/tests/tasks/uipath-maestro-flow/single_node/transform_group_by/transform_group_by.yaml @@ -39,8 +39,6 @@ initial_prompt: | `core.action.transform.group-by` — do NOT hardcode a guessed version. Validate the flow. Do NOT run flow debug — just validate. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between - planning and implementation. Build the complete flow end-to-end in a single pass. Before starting, follow the documented workflow steps exactly — in particular, the transform plugin's Group By JSON shape and the collection-input contract. diff --git a/tests/tasks/uipath-maestro-flow/single_node/transform_map/transform_map.yaml b/tests/tasks/uipath-maestro-flow/single_node/transform_map/transform_map.yaml index 5ee7f6e3e0..3acdb33bd8 100644 --- a/tests/tasks/uipath-maestro-flow/single_node/transform_map/transform_map.yaml +++ b/tests/tasks/uipath-maestro-flow/single_node/transform_map/transform_map.yaml @@ -37,8 +37,6 @@ initial_prompt: | skill's workflow teaches how to read the registry). Do NOT run flow debug — just validate the flow. - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between - planning and implementation. Build the complete flow end-to-end in a single pass. Before starting, follow the documented workflow steps exactly — in particular, the transform plugin's Map JSON shape and the `collection` path contract. diff --git a/tests/tasks/uipath-maestro-flow/smoke/merge_parallel_sync.yaml b/tests/tasks/uipath-maestro-flow/smoke/merge_parallel_sync.yaml index 6e8af776a5..3422337dfd 100644 --- a/tests/tasks/uipath-maestro-flow/smoke/merge_parallel_sync.yaml +++ b/tests/tasks/uipath-maestro-flow/smoke/merge_parallel_sync.yaml @@ -60,10 +60,7 @@ initial_prompt: | layout/variables, then validate with `uip maestro flow validate` — do NOT run `uip maestro flow debug` (the parallel join only synchronizes in a live engine). - - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between - planning and implementation. Build the complete flow end-to-end in a - single pass. - + - success_criteria: # ── Flow file validity ───────────────────────────────────────────────── - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/smoke/scheduled_trigger.yaml b/tests/tasks/uipath-maestro-flow/smoke/scheduled_trigger.yaml index 665a01ecb7..101d384027 100644 --- a/tests/tasks/uipath-maestro-flow/smoke/scheduled_trigger.yaml +++ b/tests/tasks/uipath-maestro-flow/smoke/scheduled_trigger.yaml @@ -48,10 +48,7 @@ initial_prompt: | - The `uip` CLI is already available; use `--output json` on uip commands. - Validate locally with `uip maestro flow validate` — do NOT run `uip maestro flow debug` (the schedule fires only in a live engine). - - Do NOT ask for approval, confirmation, or feedback. Do NOT pause between - planning and implementation. Build the complete flow end-to-end in a - single pass. - + - success_criteria: # ── Flow file validity ───────────────────────────────────────────────── - type: run_command diff --git a/tests/templates/test-task-template.yaml b/tests/templates/test-task-template.yaml index b28c427bf6..06e193b337 100644 --- a/tests/templates/test-task-template.yaml +++ b/tests/templates/test-task-template.yaml @@ -34,9 +34,11 @@ initial_prompt: | Goal-oriented prompt — describe what to build, not how. Let the skill teach the agent the workflow. + Do not tell the agent not to ask for approval or to avoid pausing. The + experiment config's system_prompt states that the run is headless; a task + that repeats it drifts from the wording every other task uses. + For e2e tests, append: - Do NOT ask for approval, confirmation, or feedback. - Do NOT pause between planning and implementation. Before starting, load the skill and follow its workflow. success_criteria: From 7b91b813af3f485ebb6a9698efef04a4228ff571 Mon Sep 17 00:00:00 2001 From: rockymadden Date: Fri, 4 Sep 2026 09:48:45 -0600 Subject: [PATCH 2/7] fix(tests): strengthen the flow headless prompt; "irreversible" was blocking debug Three problems with the first wording. "Stop only when proceeding would be unsafe or irreversible" collided with this PR's own rule #2, which tells the agent that `flow debug` overwrites the Studio Web solution behind the local .uipx SolutionId. An agent reading both had every reason to treat debug as off-limits, which is the exact 5-of-8 stall this PR exists to fix. "Unsafe" was also broad enough to make an agent balk at the e2e tasks whose entire point is creating a Jira issue or posting to Slack. The stripped task lines carried "Build the complete flow end-to-end in a single pass", which is a completeness instruction, not an autonomy one. "Keep going" did not replace it; "complete the whole task in one pass" does. "Could not resolve a fact -> stop" fired too early. slack-channel- description gave up when the first page of a lookup came back empty; slack-weather-pipeline proved the same fact was resolvable by paginating five pages. The bar is now exhausting the documented resolution path. Replaces the two vague stop conditions with two concrete ones: do not delete or overwrite what this run did not create, and do not invent a value for a lookup that genuinely failed. Everything else the task implies is authorized, including tenant writes and real messages. The task-precedence line stays, and carries skill-flow-eval-no-auto- upload, which deliberately asserts a refusal. Co-Authored-By: Claude Opus 5 (1M context) --- tests/experiments/flow.yaml | 19 ++++++++++++++----- 1 file changed, 14 insertions(+), 5 deletions(-) diff --git a/tests/experiments/flow.yaml b/tests/experiments/flow.yaml index e95c75fa8e..cf442ab16e 100644 --- a/tests/experiments/flow.yaml +++ b/tests/experiments/flow.yaml @@ -47,11 +47,20 @@ defaults: You are a coding agent. Do not access files in sibling runs/* directories. Everywhere else is permitted. This run is headless. No user is present, and nobody will answer a question or grant an approval. - Do not wait for input and do not stop to ask. Take the best available option, supply the most - defensible value where one is missing, and keep going. Stop only when proceeding would be unsafe - or irreversible, or when you could not resolve a fact the work depends on — never substitute a - guess for it. Record every decision, assumption, and blocked step in your final response. - Instructions in the task take precedence over this paragraph. + Do not ask, do not pause, and do not wait for input. Complete the whole task in one pass: take the + best available option, supply the most defensible value where one is missing, and carry on to the end. + + Assume the actions the task implies are authorized. Creating, running, debugging, and publishing + this run's own artifacts is the work, not something to seek permission for, even when it writes to + a tenant or sends a real message. + + Hold back on two things. Do not delete or overwrite anything this run did not create. And when a + lookup the work depends on comes back empty or fails, exhaust the documented way of resolving it + before giving up; only if that genuinely fails, stop on that one field rather than inventing a + value, and say so. + + Record every decision, assumption, and blocked step in your final response. Instructions in the + task take precedence over this paragraph. plugins: - type: "local" path: "$SKILLS_REPO_PATH" From 18d100a3f500dc10488fd8f4706aa054b203349d Mon Sep 17 00:00:00 2001 From: rockymadden Date: Fri, 4 Sep 2026 09:53:02 -0600 Subject: [PATCH 3/7] fix(tests): repair two spacing artifacts left by the prompt strip `escalation_jira_ticket` read "seeded inputs.Before starting" and `escalation_orchestrator_paths` had a double space in the same spot. Co-Authored-By: Claude Opus 5 (1M context) --- .../e2e/escalation_jira_ticket/escalation_jira_ticket.yaml | 2 +- .../escalation_orchestrator_paths.yaml | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/tests/tasks/uipath-maestro-flow/e2e/escalation_jira_ticket/escalation_jira_ticket.yaml b/tests/tasks/uipath-maestro-flow/e2e/escalation_jira_ticket/escalation_jira_ticket.yaml index 857421e170..b97b9f3779 100644 --- a/tests/tasks/uipath-maestro-flow/e2e/escalation_jira_ticket/escalation_jira_ticket.yaml +++ b/tests/tasks/uipath-maestro-flow/e2e/escalation_jira_ticket/escalation_jira_ticket.yaml @@ -54,7 +54,7 @@ initial_prompt: | correlationId), and jiraIssueKey (the created issue's key). Validate the flow. Do NOT run or debug the flow — the grader executes it with the - seeded inputs.Before starting, load the uipath-maestro-flow skill and follow its + seeded inputs. Before starting, load the uipath-maestro-flow skill and follow its workflow. success_criteria: diff --git a/tests/tasks/uipath-maestro-flow/e2e/escalation_orchestrator_paths/escalation_orchestrator_paths.yaml b/tests/tasks/uipath-maestro-flow/e2e/escalation_orchestrator_paths/escalation_orchestrator_paths.yaml index 9e44b5a465..899c5ddc81 100644 --- a/tests/tasks/uipath-maestro-flow/e2e/escalation_orchestrator_paths/escalation_orchestrator_paths.yaml +++ b/tests/tasks/uipath-maestro-flow/e2e/escalation_orchestrator_paths/escalation_orchestrator_paths.yaml @@ -82,7 +82,7 @@ initial_prompt: | Salesforce/Jira/Drive are NOT part of this task — the match status is an input. Validate the flow. Do NOT run or debug the flow — the grader executes it with - seeded inputs. Before starting, load the uipath-maestro-flow skill and follow its workflow. + seeded inputs. Before starting, load the uipath-maestro-flow skill and follow its workflow. success_criteria: - type: command_executed From 6689520cd86971c0b44a6226dacc5b79896b8d5d Mon Sep 17 00:00:00 2001 From: rockymadden Date: Fri, 4 Sep 2026 09:58:27 -0600 Subject: [PATCH 4/7] fix(tests): close the regression, the simulated-task hole, and the drift risk MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Review of the PR against how it would actually be run turned up four problems, one of which made the change a net regression. The experiment input defaults to nightly.yaml, and run-coder-eval.yml is dispatch-only with per-skill globs. Anyone dispatching a flow run without setting the new input would get task prompts with the autonomy lines removed and a config that never states the run is headless — strictly worse than main. The partition step now refuses that dispatch, and refuses the mirror mistake of running uipath-maestro-flow/interactive/ under flow.yaml, where the task has a simulated user and the prompt would tell it nobody is there. Both cases were table-tested across seven dispatch shapes. flow.yaml is a snapshot of nightly.yaml, not a derivation — the experiment schema has no `extends`, so 55 lines are duplicated and will drift. An earlier commit message claimed the opposite. A parity test now asserts every substantive line of nightly.yaml still appears in flow.yaml, verified by adding a fake secret to nightly's passthrough list and watching it fail. The system prompt relied on "instructions in the task take precedence" to protect skill-flow-eval-no-auto-upload, which deliberately asserts a refusal. A system prompt outranking a user prompt is the usual direction, so that was load-bearing on an assumption. Uploading or publishing to a shared destination unless the task asks for it is now a hold-back in its own right, and no longer depends on precedence holding. Also drops EXPERIMENT_YAML from the sandbox env passthrough; it is used host-side to build the CLI args and the agent has no use for it. Co-Authored-By: Claude Opus 5 (1M context) --- .github/workflows/run-coder-eval.yml | 7 +- tests/experiments/flow.yaml | 8 +-- .../_shared/test_flow_experiment_parity.py | 68 +++++++++++++++++++ 3 files changed, 76 insertions(+), 7 deletions(-) create mode 100644 tests/tasks/uipath-maestro-flow/_shared/test_flow_experiment_parity.py diff --git a/.github/workflows/run-coder-eval.yml b/.github/workflows/run-coder-eval.yml index 28f37e2289..5a5feea21a 100644 --- a/.github/workflows/run-coder-eval.yml +++ b/.github/workflows/run-coder-eval.yml @@ -171,8 +171,9 @@ jobs: - name: Resolve globs, split by tag, enforce large-run gate id: split env: - INPUT_GLOBS: ${{ inputs.task_globs }} - CONFIRMED: ${{ inputs.confirm_large_run }} + INPUT_GLOBS: ${{ inputs.task_globs }} + CONFIRMED: ${{ inputs.confirm_large_run }} + EXPERIMENT_YAML: ${{ inputs.experiment }} run: | set -euo pipefail if [ -z "$INPUT_GLOBS" ]; then @@ -465,7 +466,7 @@ jobs: env_lines=() for name in SKILLS_REPO_PATH API_BACKEND AWS_BEARER_TOKEN_BEDROCK AWS_REGION \ BEDROCK_MODEL ANTHROPIC_API_KEY E2E_PROCESS_KEY E2E_LONG_PROCESS_KEY \ - TASK_GLOBS TASK_COUNT TASK_PARALLELISM AGENT AGENT_MODEL EXPERIMENT_YAML \ + TASK_GLOBS TASK_COUNT TASK_PARALLELISM AGENT AGENT_MODEL \ CLAUDE_CODE_MODEL CODEX_MODEL ANTIGRAVITY_MODEL DELEGATE_MODEL \ CODEX_API_KEY CODEX_BASE_URL GEMINI_API_KEY; do env_lines+=("$name=${!name}") diff --git a/tests/experiments/flow.yaml b/tests/experiments/flow.yaml index cf442ab16e..10e9cf2d01 100644 --- a/tests/experiments/flow.yaml +++ b/tests/experiments/flow.yaml @@ -54,10 +54,10 @@ defaults: this run's own artifacts is the work, not something to seek permission for, even when it writes to a tenant or sends a real message. - Hold back on two things. Do not delete or overwrite anything this run did not create. And when a - lookup the work depends on comes back empty or fails, exhaust the documented way of resolving it - before giving up; only if that genuinely fails, stop on that one field rather than inventing a - value, and say so. + Hold back on three things. Do not delete or overwrite anything this run did not create. Do not + upload or publish to a shared destination unless the task asks for it. And when a lookup the work + depends on comes back empty or fails, exhaust the documented way of resolving it before giving up; + only if that genuinely fails, stop on that one field rather than inventing a value, and say so. Record every decision, assumption, and blocked step in your final response. Instructions in the task take precedence over this paragraph. diff --git a/tests/tasks/uipath-maestro-flow/_shared/test_flow_experiment_parity.py b/tests/tasks/uipath-maestro-flow/_shared/test_flow_experiment_parity.py new file mode 100644 index 0000000000..88917038fe --- /dev/null +++ b/tests/tasks/uipath-maestro-flow/_shared/test_flow_experiment_parity.py @@ -0,0 +1,68 @@ +"""`experiments/flow.yaml` must not drift from `experiments/nightly.yaml`. + +flow.yaml exists for one reason: a system prompt stating the run is headless, +which the flow task prompts no longer carry themselves. Everything else — the +docker image, `env_passthrough_extra`, mounts, `checker_context`, `run_limits`, +`post_run` — is copied, because the experiment schema has no `extends`. A secret +added to nightly's passthrough list and not to flow's breaks flow runs with a +missing environment variable and no obvious cause. + +This asserts every substantive line of nightly.yaml still appears in flow.yaml. +It does not assert the reverse: flow.yaml is allowed to add the headless +paragraph and its own id and description. + +Regex, not PyYAML: CI installs only pytest, and a module-level `import yaml` +would error at collection and take the suite with it (see test_criterion_budgets). +""" + +from __future__ import annotations + +import os + +_HERE = os.path.dirname(os.path.abspath(__file__)) +_EXPERIMENTS = os.path.normpath(os.path.join(_HERE, "..", "..", "..", "experiments")) +_NIGHTLY = os.path.join(_EXPERIMENTS, "nightly.yaml") +_FLOW = os.path.join(_EXPERIMENTS, "flow.yaml") + +# Lines that legitimately differ: the identity of the config itself. +_EXEMPT_PREFIXES = ("experiment_id:", "description:") + + +def _substantive(path: str) -> list[str]: + """Non-blank, non-comment lines, minus the config's own identity block.""" + out: list[str] = [] + in_description = False + for raw in open(path, encoding="utf-8"): + line = raw.rstrip("\n") + stripped = line.strip() + if stripped.startswith("description:"): + # Folded scalar: skip its indented continuation lines too. + in_description = stripped.endswith((">", "|")) + continue + if in_description: + if line and not line[0].isspace(): + in_description = False + else: + continue + if not stripped or stripped.startswith("#"): + continue + if stripped.startswith(_EXEMPT_PREFIXES): + continue + out.append(line) + return out + + +def test_flow_config_carries_every_nightly_setting(): + missing = [ln for ln in _substantive(_NIGHTLY) if ln not in _substantive(_FLOW)] + assert not missing, ( + "experiments/flow.yaml has drifted from nightly.yaml. Copy these lines over " + "(flow.yaml is a snapshot of nightly's runtime plus a headless system prompt):\n " + + "\n ".join(missing) + ) + + +def test_flow_config_states_the_run_is_headless(): + """The one thing flow.yaml exists to add. The task prompts no longer say it.""" + text = open(_FLOW, encoding="utf-8").read() + assert "This run is headless." in text + assert "No user is present" in text From 663896610f3fadbcd22246e4b158c82ba710e1ab Mon Sep 17 00:00:00 2001 From: rockymadden Date: Fri, 4 Sep 2026 10:02:50 -0600 Subject: [PATCH 5/7] fix(ci): actually add the dispatch guard, and fix its empty-array bug The previous commit claimed a guard that is not in the file. The patch script raised before its write, so nothing landed, and I table-tested a standalone copy of the block rather than the workflow. The commit message and the PR comment both asserted behaviour that did not exist. The guard is now in the split step, and the test harness extracts the block from the shipping YAML rather than retyping it, so this cannot recur. Seven dispatch shapes pass, including the empty selection. Its first draft also had a real defect: `ALL=("${LINUX[@]}" "${WINDOWS[@]}")` is an unbound-variable error under `set -u` when an array is empty, which is the normal case for one of the two platforms. Reproduced on bash 3.2. Now word-splits the same `${LINUX[*]:-}` form the surrounding code uses. Parity test no longer re-reads and re-parses flow.yaml once per line of nightly.yaml. Co-Authored-By: Claude Opus 5 (1M context) --- .github/workflows/run-coder-eval.yml | 21 +++++++++++++++++++ .../_shared/test_flow_experiment_parity.py | 3 ++- 2 files changed, 23 insertions(+), 1 deletion(-) diff --git a/.github/workflows/run-coder-eval.yml b/.github/workflows/run-coder-eval.yml index 5a5feea21a..b0b3c7bdef 100644 --- a/.github/workflows/run-coder-eval.yml +++ b/.github/workflows/run-coder-eval.yml @@ -209,6 +209,27 @@ jobs: echo "::error::Glob matches $TOTAL tasks (> 50). Tick confirm_large_run to authorize." exit 1 fi + # The flow suite's prompts carry no autonomy language; experiments/flow.yaml + # supplies it in the system prompt. Running flow tasks under any other config + # silently drops it, and running the simulated tasks under flow.yaml tells a + # task that HAS a user that nobody is present. Catch both here rather than + # discover them in the results. Word-split, not "${ARR[@]}": the arrays are + # empty on one platform and that is an unbound-variable error under `set -u`. + FLOW_N=0; SIM_N=0 + for f in ${LINUX[*]:-} ${WINDOWS[*]:-}; do + case "$f" in + *tasks/uipath-maestro-flow/interactive/*) SIM_N=$((SIM_N + 1)) ;; + *tasks/uipath-maestro-flow/*) FLOW_N=$((FLOW_N + 1)) ;; + esac + done + if [ "$FLOW_N" -gt 0 ] && [ "$EXPERIMENT_YAML" != "experiments/flow.yaml" ]; then + echo "::error::$FLOW_N flow task(s) selected under '$EXPERIMENT_YAML'. Flow task prompts carry no autonomy language; experiments/flow.yaml supplies it. Set the experiment input, and dispatch flow separately from other skills." + exit 1 + fi + if [ "$SIM_N" -gt 0 ] && [ "$EXPERIMENT_YAML" = "experiments/flow.yaml" ]; then + echo "::error::$SIM_N task(s) under uipath-maestro-flow/interactive/ have a simulated user and must not run under experiments/flow.yaml, which states no user is present. Exclude that directory from the glob." + exit 1 + fi # Space-separated; downstream jobs word-split on consumption. echo "linux_globs=${LINUX[*]:-}" >> "$GITHUB_OUTPUT" echo "windows_globs=${WINDOWS[*]:-}" >> "$GITHUB_OUTPUT" diff --git a/tests/tasks/uipath-maestro-flow/_shared/test_flow_experiment_parity.py b/tests/tasks/uipath-maestro-flow/_shared/test_flow_experiment_parity.py index 88917038fe..630ab2afec 100644 --- a/tests/tasks/uipath-maestro-flow/_shared/test_flow_experiment_parity.py +++ b/tests/tasks/uipath-maestro-flow/_shared/test_flow_experiment_parity.py @@ -53,7 +53,8 @@ def _substantive(path: str) -> list[str]: def test_flow_config_carries_every_nightly_setting(): - missing = [ln for ln in _substantive(_NIGHTLY) if ln not in _substantive(_FLOW)] + flow = set(_substantive(_FLOW)) + missing = [ln for ln in _substantive(_NIGHTLY) if ln not in flow] assert not missing, ( "experiments/flow.yaml has drifted from nightly.yaml. Copy these lines over " "(flow.yaml is a snapshot of nightly's runtime plus a headless system prompt):\n " From f26bbb383f5f78999c35f7b0466300113d0e92d5 Mon Sep 17 00:00:00 2001 From: rockymadden Date: Fri, 4 Sep 2026 10:06:55 -0600 Subject: [PATCH 6/7] =?UTF-8?q?fix:=20address=20Copilot=20review=20?= =?UTF-8?q?=E2=80=94=20dangling=20rule-5=20fallback,=20stale=20consent=20g?= =?UTF-8?q?ates,=20orphaned=20bullets?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Four valid findings. Rule #5's non-interactive fallback was deleted as collateral when the rule-13 experiment was reverted. The intended edit was only dropping `flow debug` from its consent-gate list. Rule #4 and ixp/impl.md:356 both route folder selection through "rule #5 (its non-interactive fallback applies)", so both pointed at behaviour that no longer existed. Restored, minus `flow debug`. Both authoring handoffs still gated Debug on consent ("Confirm consent first", "see the consent-before-debug rule"), which is the exact gate rule #2 stops applying. An agent following greenfield or brownfield to its What's next menu could still stall there. Both now say consent comes from the mandate and that selecting the option is the ask. The prompt strip left a bare list marker in four tasks, so every run sent an empty bullet: inline_agent_eval, e2e_03_project_creation_handoff, merge_parallel_sync, scheduled_trigger. Swept the whole suite; none left. The template's "do not add autonomy language" note was blanket, but only experiments/flow.yaml carries the headless system prompt. A future non-flow e2e task following it would have shipped with no autonomy instruction at all. Scoped per experiment. Not changed: flow.yaml being a copy of nightly.yaml rather than a derived config. There is no `extends` in the experiment schema, so the parity test added earlier is the remedy Copilot's own comment suggests. Co-Authored-By: Claude Opus 5 (1M context) --- skills/uipath-maestro-flow/SKILL.md | 2 +- .../uipath-maestro-flow/references/author/brownfield.md | 2 +- .../uipath-maestro-flow/references/author/greenfield.md | 2 +- .../evaluate/inline_agent_eval/inline_agent_eval.yaml | 1 - .../e2e_03_project_creation_handoff.yaml | 1 - .../uipath-maestro-flow/smoke/merge_parallel_sync.yaml | 1 - .../uipath-maestro-flow/smoke/scheduled_trigger.yaml | 1 - tests/templates/test-task-template.yaml | 8 +++++--- 8 files changed, 8 insertions(+), 10 deletions(-) diff --git a/skills/uipath-maestro-flow/SKILL.md b/skills/uipath-maestro-flow/SKILL.md index 630d6eda28..5c31f8de4a 100644 --- a/skills/uipath-maestro-flow/SKILL.md +++ b/skills/uipath-maestro-flow/SKILL.md @@ -73,7 +73,7 @@ Guide for creating, editing, validating, debugging, publishing, diagnosing, and **Two tells that you skipped the search and took the brand-name shortcut — both are build defects, not valid manual-mode HTTP:** (a) you authored a manual-mode `core.action.http.v2` node whose `url` targets a well-known SaaS API domain that has a connector (`slack.com/api/*`, `api.github.com`, `*.salesforce.com`, `graph.microsoft.com`, …); (b) you declared an `in` variable to hold that service's API token or secret (e.g. a `slackToken` holding an `xoxb-…` bot token, an `apiKey`, a bearer token). A connector-backed flow never carries the raw credential — the IS connection does. If you find yourself writing either, **stop**: run `uip maestro flow registry search ""` and `uip is connections list "" --all-folders`, then use the connector activity (or connector-mode HTTP: `authentication:"connector"` + `targetConnector` + a bound `connectionId`/`folderKey`). Manual mode is legitimate only for a service the search proves has no connector. 4. **Never invoke other skills automatically** — when a flow needs an RPA process, agent, or app, identify the gap and provide handoff instructions. Let the user decide when to switch skills. **One exception — IXP extraction with documents in hand:** when the flow needs document extraction, the user supplied sample documents, and `registry search "uipath.ixp"` shows no extractor covering them, invoke the `uipath-ixp` skill to build and deploy the model, then resume the flow ([plugins/ixp/impl.md — If the Model Does Not Exist Yet](references/author/plugins/ixp/impl.md#if-the-model-does-not-exist-yet)). Resolve the target Orchestrator folder for the deployment before invoking — from the user's request when it names one, otherwise per rule #5 (its non-interactive fallback applies) — and pass it in the handoff; the sibling stops rather than guess a folder. There is deliberately no separate consent gate on the tenant writes this creates: the project and folder deployment fulfil the extraction request itself, and the one consequential choice — where the deployment lands (deployments have no delete API) — is exactly the folder decision rule #5 just routed. Do NOT drive `uip ixp` project or deployment commands from this skill instead of invoking it — the sibling's guides carry guardrails this skill does not. If `uipath-ixp` is unavailable in the session, fall back to `core.logic.mock` plus an Open Questions entry, exactly as when no documents were supplied. -5. **Always present finite decisions as a dropdown with a final "Something else" escape hatch.** Whenever the skill needs a decision (which solution, publish vs debug vs deploy, which connector, trigger type, or resource to bind, etc.), ask with the enumerated choices plus **"Something else"** last for free-form input; never ask open-ended in chat when a finite set of sensible defaults exists. If the user picks "Something else", parse their answer and continue. No structured-question facility on the harness → ask in chat as a numbered list with "Something else" last. These fallbacks define "ask the user" / "confirm with the user" wherever this skill's references require it. +5. **Always present finite decisions as a dropdown with a final "Something else" escape hatch.** Whenever the skill needs a decision (which solution, publish vs debug vs deploy, which connector, trigger type, or resource to bind, etc.), ask with the enumerated choices plus **"Something else"** last for free-form input; never ask open-ended in chat when a finite set of sensible defaults exists. If the user picks "Something else", parse their answer and continue. No structured-question facility on the harness → ask in chat as a numbered list with "Something else" last. Non-interactively (CI/headless, no user available) → take the marked recommended option, proceed, and record the decision prominently in the final report; if none is recommended, stop and report the open decision instead of guessing. Consent gates (destructive operations, tenant writes) are never auto-answered — in non-interactive mode, stop and report the blocked step; `flow debug` is not one of them, and is governed by the mandate rule above. These fallbacks define "ask the user" / "confirm with the user" wherever this skill's references require it. 6. **Discover the target solution before scaffolding.** A Flow project must use double nesting: `//.flow`. Before any new `uip solution init` or `uip maestro flow init`, run `find . -maxdepth 2 -type f -name '*.uipx' -print`. If a solution exists, stop and ask which to use: one option per solution, "Create a new solution", then "Something else". Do not silently adopt, initialize, delete, or repair an existing solution, even if a new one was requested. If creating one, ask for its name rather than defaulting to the Flow name. diff --git a/skills/uipath-maestro-flow/references/author/brownfield.md b/skills/uipath-maestro-flow/references/author/brownfield.md index 8523dee50f..501122d6fc 100644 --- a/skills/uipath-maestro-flow/references/author/brownfield.md +++ b/skills/uipath-maestro-flow/references/author/brownfield.md @@ -82,7 +82,7 @@ Authoring ends here. For any selected option, read [operate/CAPABILITY.md](../op | Option | What it does | |---|---| | **Publish to Studio Web** | Push the solution to Studio Web so the user can visualize, edit, and publish from the browser. | -| **Debug the solution** | Execute the flow end-to-end against real systems. Confirm consent first because debug has real side effects (see the consent-before-debug rule in [SKILL.md](../../SKILL.md)). | +| **Debug the solution** | Execute the flow end-to-end against real systems. Consent comes from the mandate, not from this menu — see the `flow debug` rule in [SKILL.md](../../SKILL.md). Selecting it here is the user asking for a run. | | **Deploy to Orchestrator** | Pack and publish directly to Orchestrator (bypasses Studio Web). Only when explicitly chosen; see [/uipath:uipath-platform](/uipath:uipath-platform). | | **Something else** | Last option. Accept free-form string input and act on it. | diff --git a/skills/uipath-maestro-flow/references/author/greenfield.md b/skills/uipath-maestro-flow/references/author/greenfield.md index 769b1e89c8..162b44b3c0 100644 --- a/skills/uipath-maestro-flow/references/author/greenfield.md +++ b/skills/uipath-maestro-flow/references/author/greenfield.md @@ -374,7 +374,7 @@ Authoring terminates here. Each option below hands off to Operate — read [oper | Option | What it does | | --- | --- | | **Publish to Studio Web** | Push the solution to Studio Web so the user can visualize, edit, and publish from the browser. | -| **Debug the solution** | Execute the flow end-to-end against real systems. Confirm consent first — debug has real side effects (see the consent-before-debug rule in [SKILL.md](../../SKILL.md)). | +| **Debug the solution** | Execute the flow end-to-end against real systems. Consent comes from the mandate, not from this menu — see the `flow debug` rule in [SKILL.md](../../SKILL.md). Selecting it here is the user asking for a run. | | **Deploy to Orchestrator** | Pack and publish directly to Orchestrator (bypasses Studio Web). Only when explicitly chosen — see [/uipath:uipath-platform](/uipath:uipath-platform). | | **Something else** | Last option. Accept free-form string input and act on it (e.g., "just leave it", "pack but don't publish", "upload to a different tenant"). | diff --git a/tests/tasks/uipath-maestro-flow/evaluate/inline_agent_eval/inline_agent_eval.yaml b/tests/tasks/uipath-maestro-flow/evaluate/inline_agent_eval/inline_agent_eval.yaml index 85d51cf57b..951b788737 100644 --- a/tests/tasks/uipath-maestro-flow/evaluate/inline_agent_eval/inline_agent_eval.yaml +++ b/tests/tasks/uipath-maestro-flow/evaluate/inline_agent_eval/inline_agent_eval.yaml @@ -54,7 +54,6 @@ initial_prompt: | - Local CRUD only. Do NOT call `uip login`. Do NOT call `uip solution upload`. Do NOT call `uip maestro flow eval run *`. Do NOT run `uip maestro flow debug`. - - success_criteria: - type: command_executed description: "Agent scaffolded the inline agent with uip agent init --inline-in-flow" diff --git a/tests/tasks/uipath-maestro-flow/ixp/e2e_03_project_creation_handoff/e2e_03_project_creation_handoff.yaml b/tests/tasks/uipath-maestro-flow/ixp/e2e_03_project_creation_handoff/e2e_03_project_creation_handoff.yaml index d4c537699f..a1951cc845 100644 --- a/tests/tasks/uipath-maestro-flow/ixp/e2e_03_project_creation_handoff/e2e_03_project_creation_handoff.yaml +++ b/tests/tasks/uipath-maestro-flow/ixp/e2e_03_project_creation_handoff/e2e_03_project_creation_handoff.yaml @@ -97,7 +97,6 @@ initial_prompt: | - The `uip` CLI is already available and the runner is logged in. - Use `--output json` on all uip commands. - Do NOT run `uip maestro flow debug` or `uip maestro flow deploy`. - - success_criteria: - type: skill_triggered description: "Agent invoked the uipath-maestro-flow Skill" diff --git a/tests/tasks/uipath-maestro-flow/smoke/merge_parallel_sync.yaml b/tests/tasks/uipath-maestro-flow/smoke/merge_parallel_sync.yaml index 3422337dfd..be182fbb1c 100644 --- a/tests/tasks/uipath-maestro-flow/smoke/merge_parallel_sync.yaml +++ b/tests/tasks/uipath-maestro-flow/smoke/merge_parallel_sync.yaml @@ -60,7 +60,6 @@ initial_prompt: | layout/variables, then validate with `uip maestro flow validate` — do NOT run `uip maestro flow debug` (the parallel join only synchronizes in a live engine). - - success_criteria: # ── Flow file validity ───────────────────────────────────────────────── - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/smoke/scheduled_trigger.yaml b/tests/tasks/uipath-maestro-flow/smoke/scheduled_trigger.yaml index 101d384027..ad17a89179 100644 --- a/tests/tasks/uipath-maestro-flow/smoke/scheduled_trigger.yaml +++ b/tests/tasks/uipath-maestro-flow/smoke/scheduled_trigger.yaml @@ -48,7 +48,6 @@ initial_prompt: | - The `uip` CLI is already available; use `--output json` on uip commands. - Validate locally with `uip maestro flow validate` — do NOT run `uip maestro flow debug` (the schedule fires only in a live engine). - - success_criteria: # ── Flow file validity ───────────────────────────────────────────────── - type: run_command diff --git a/tests/templates/test-task-template.yaml b/tests/templates/test-task-template.yaml index 06e193b337..37dea1dd49 100644 --- a/tests/templates/test-task-template.yaml +++ b/tests/templates/test-task-template.yaml @@ -34,9 +34,11 @@ initial_prompt: | Goal-oriented prompt — describe what to build, not how. Let the skill teach the agent the workflow. - Do not tell the agent not to ask for approval or to avoid pausing. The - experiment config's system_prompt states that the run is headless; a task - that repeats it drifts from the wording every other task uses. + Autonomy language depends on the experiment the task runs under. Under + experiments/flow.yaml the system prompt already states the run is headless, + so do not repeat it — a task that does drifts from the wording every other + task uses. Under experiments/default.yaml or nightly.yaml it does not, so an + e2e task that must not stall on a consent gate still needs to say so itself. For e2e tests, append: Before starting, load the skill and follow its workflow. From 8f2b4b864a82c832ff0e92b6833547f9dcd9fbb5 Mon Sep 17 00:00:00 2001 From: rockymadden Date: Fri, 4 Sep 2026 10:24:24 -0600 Subject: [PATCH 7/7] docs(tests): document experiments/flow.yaml and the make flow target Both were missing from the README table of experiments and the list of make targets, which is where anyone looks to find out how to run a suite. Co-Authored-By: Claude Opus 5 (1M context) --- tests/README.md | 5 +++++ 1 file changed, 5 insertions(+) diff --git a/tests/README.md b/tests/README.md index abddabdff2..1763d63263 100644 --- a/tests/README.md +++ b/tests/README.md @@ -52,6 +52,10 @@ make smoke_rpa # Run e2e-tagged tests under the default config make e2e +# Run the flow suite zero-shot: experiments/flow.yaml states the run is headless, +# and the glob excludes interactive/, whose tasks have a simulated user +make flow + # Run tests matching a combination of tags (AND semantics — tasks must carry all listed tags) (defaults to experiments/default.yaml): make tags TAGS="integration connector-feature" # Optionally override the experiment config @@ -176,6 +180,7 @@ Run-time caps live under `defaults.run_limits` (see coder_eval `RunLimits`). | `default.yaml` | tempdir | Devs locally, ad-hoc runs | 200 | 1200s | 900s | | `nightly.yaml` | docker | Nightly cron (`daily.sh`) | 200 | 1200s | 900s | | `smoke.yaml` | docker | PR-gate smoke (Linux) | 40 | 900s | 900s | +| `flow.yaml` | docker | `make flow` / flow dispatches — nightly's runtime plus a system prompt stating the run is headless | 200 | 1200s | 900s | | `smoke-windows.yaml` | tempdir | PR-gate smoke (Windows RPA only) | 40 | 900s | 900s | | `activation.yaml` | tempdir | Skill activation classifier (benchmark) | 3 + early-stop | 360s | 120s | | `same-ground-headtohead.yaml` | docker | Campaign-only local comparison arm | 200 | 1200s | 900s |