Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
37 changes: 34 additions & 3 deletions .github/workflows/run-coder-eval.yml
Original file line number Diff line number Diff line change
Expand Up @@ -23,6 +23,14 @@ on:
description: 'REQUIRED. Space-separated globs under tests/. E.g. tasks/uipath-agents/**/*.yaml'
type: string
required: true
# Per-skill configs differ only in defaults the runtime shares, so the
# file is a parameter rather than a fork of this workflow. Scope it with
# task_globs: a config's defaults apply to whatever that run selects.
experiment:
description: 'Experiment YAML under tests/. E.g. experiments/flow.yaml for the flow zero-shot config.'
type: string
required: false
default: 'experiments/nightly.yaml'
# Host-level concurrency for the Linux job's `coder-eval -j`. Default 4:
# ubuntu-latest is 4 vCPU, and j=20 oversubscribes ~5:1 — agents miss the
# 1200s turn timeout and ERROR (false negatives, including a bindings
Expand Down Expand Up @@ -163,8 +171,9 @@ jobs:
- name: Resolve globs, split by tag, enforce large-run gate
id: split
env:
INPUT_GLOBS: ${{ inputs.task_globs }}
CONFIRMED: ${{ inputs.confirm_large_run }}
INPUT_GLOBS: ${{ inputs.task_globs }}
CONFIRMED: ${{ inputs.confirm_large_run }}
EXPERIMENT_YAML: ${{ inputs.experiment }}
run: |
set -euo pipefail
if [ -z "$INPUT_GLOBS" ]; then
Expand Down Expand Up @@ -200,6 +209,27 @@ jobs:
echo "::error::Glob matches $TOTAL tasks (> 50). Tick confirm_large_run to authorize."
exit 1
fi
# The flow suite's prompts carry no autonomy language; experiments/flow.yaml
# supplies it in the system prompt. Running flow tasks under any other config
# silently drops it, and running the simulated tasks under flow.yaml tells a
# task that HAS a user that nobody is present. Catch both here rather than
# discover them in the results. Word-split, not "${ARR[@]}": the arrays are
# empty on one platform and that is an unbound-variable error under `set -u`.
FLOW_N=0; SIM_N=0
for f in ${LINUX[*]:-} ${WINDOWS[*]:-}; do
case "$f" in
*tasks/uipath-maestro-flow/interactive/*) SIM_N=$((SIM_N + 1)) ;;
*tasks/uipath-maestro-flow/*) FLOW_N=$((FLOW_N + 1)) ;;
esac
done
if [ "$FLOW_N" -gt 0 ] && [ "$EXPERIMENT_YAML" != "experiments/flow.yaml" ]; then
echo "::error::$FLOW_N flow task(s) selected under '$EXPERIMENT_YAML'. Flow task prompts carry no autonomy language; experiments/flow.yaml supplies it. Set the experiment input, and dispatch flow separately from other skills."
exit 1
fi
if [ "$SIM_N" -gt 0 ] && [ "$EXPERIMENT_YAML" = "experiments/flow.yaml" ]; then
echo "::error::$SIM_N task(s) under uipath-maestro-flow/interactive/ have a simulated user and must not run under experiments/flow.yaml, which states no user is present. Exclude that directory from the glob."
exit 1
fi
# Space-separated; downstream jobs word-split on consumption.
echo "linux_globs=${LINUX[*]:-}" >> "$GITHUB_OUTPUT"
echo "windows_globs=${WINDOWS[*]:-}" >> "$GITHUB_OUTPUT"
Expand Down Expand Up @@ -409,6 +439,7 @@ jobs:
E2E_LONG_PROCESS_KEY: ${{ secrets.E2E_LONG_PROCESS_KEY }}
TASK_GLOBS: ${{ needs.partition.outputs.linux_globs }}
TASK_COUNT: ${{ needs.partition.outputs.linux_count }}
EXPERIMENT_YAML: ${{ inputs.experiment }}
TASK_PARALLELISM: ${{ inputs.parallelism }}
AGENT: ${{ inputs.agent }}
AGENT_MODEL: ${{ inputs.agent_model }}
Expand Down Expand Up @@ -465,7 +496,7 @@ jobs:
# Args appended to `coder-eval run`, ONE PER LINE, each verbatim — no
# word splitting, no pathname expansion. That is what the delegate `-D`
# override needs: to bash its `[...]` list is a character class.
args=(-e experiments/nightly.yaml)
args=(-e "${EXPERIMENT_YAML}")

# Agent selection. codex authenticates via CODEX_API_KEY/CODEX_BASE_URL,
# antigravity via GEMINI_API_KEY (SDK baked into the agent image),
Expand Down
4 changes: 2 additions & 2 deletions skills/uipath-maestro-flow/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -61,7 +61,7 @@ Guide for creating, editing, validating, debugging, publishing, diagnosing, and
> **Tool vocabulary.** `Edit` means in-place replacement, `Write` a full-file write, `Read`/`Glob`/`Grep` file access, `Bash` shell, and a progress list the harness task list. Map them to equivalent tools elsewhere; preserve reviewable diffs and use shell file edits only as a last resort.

1. **Use `--output json`; prefer `--output-filter` for extraction.** Filters are global and run against the `Data` envelope, so expressions start at `Data` without a `Data.` prefix. Registry search returns a flat PascalCase array (`NodeType`, `DisplayName`, `Description`, `AvailableOnTenant`), not `Data.Nodes` or lowercase fields. Example: `uip maestro flow registry search <keyword> --output json --output-filter "[*].{NodeType:NodeType,DisplayName:DisplayName,Description:Description,AvailableOnTenant:AvailableOnTenant}"`. With `--local`, omit `AvailableOnTenant`. Use `python3 -c` or `jq` only after verifying shape and when JMESPath cannot express the transform. See [cli-conventions.md §3](references/shared/cli-conventions.md#3-prefer---output-filter-for-extraction).
2. **Do not run `flow debug` without explicit user consent.** It executes the flow for real (sends emails, posts messages, calls APIs).
2. **`flow debug` consent comes from the mandate.** It executes the flow for real (sends emails, posts messages, calls APIs), so run it only when the request is for a flow that *works*: the user asked for something that does X, or said make it work, get it running, iterate until it passes. Building and validating does not discharge that, and a flow that was never executed is not finished. Ask when the request stops short of a working artifact (review this, add a node, validate only); with nobody to ask, report debug as the step not run. Debug also overwrites the Studio Web solution matching the local `.uipx` `SolutionId`, so never debug a solution this run did not scaffold.
Comment thread
rockymadden marked this conversation as resolved.
3. **Search before creating or declaring resources absent.** For named agents, API workflows, RPA processes, and similar resources: (a) pull and search the tenant registry with `uip maestro flow registry pull --force && uip maestro flow registry search "<name>" --output json`; pull first because the cache expires after 30 minutes, login is required, and only published resources are returned; (b) search locally with `uip maestro flow registry list --local --output json` or `search "<name>" --local` (no login; returns sibling projects in the same `.uipx` solution); an empty keyword search does not prove absence, so confirm with `list --local`; (c) scaffold, mock, or create only when both searches find no match and the user explicitly requests embedding/creation or no published resource satisfies the need.

"Coded" and "low-code" describe implementation style, not inline status. Use `uipath.agent.autonomous` only when explicitly asked to embed/inline/create an agent. Use `core.logic.mock` only when the resource is neither in the solution nor published. See [rpa](references/author/plugins/rpa/impl.md) and [agent](references/author/plugins/agent/impl.md).
Expand All @@ -73,7 +73,7 @@ Guide for creating, editing, validating, debugging, publishing, diagnosing, and
**Two tells that you skipped the search and took the brand-name shortcut — both are build defects, not valid manual-mode HTTP:** (a) you authored a manual-mode `core.action.http.v2` node whose `url` targets a well-known SaaS API domain that has a connector (`slack.com/api/*`, `api.github.com`, `*.salesforce.com`, `graph.microsoft.com`, …); (b) you declared an `in` variable to hold that service's API token or secret (e.g. a `slackToken` holding an `xoxb-…` bot token, an `apiKey`, a bearer token). A connector-backed flow never carries the raw credential — the IS connection does. If you find yourself writing either, **stop**: run `uip maestro flow registry search "<service>"` and `uip is connections list "<connector-key>" --all-folders`, then use the connector activity (or connector-mode HTTP: `authentication:"connector"` + `targetConnector` + a bound `connectionId`/`folderKey`). Manual mode is legitimate only for a service the search proves has no connector.

4. **Never invoke other skills automatically** — when a flow needs an RPA process, agent, or app, identify the gap and provide handoff instructions. Let the user decide when to switch skills. **One exception — IXP extraction with documents in hand:** when the flow needs document extraction, the user supplied sample documents, and `registry search "uipath.ixp"` shows no extractor covering them, invoke the `uipath-ixp` skill to build and deploy the model, then resume the flow ([plugins/ixp/impl.md — If the Model Does Not Exist Yet](references/author/plugins/ixp/impl.md#if-the-model-does-not-exist-yet)). Resolve the target Orchestrator folder for the deployment before invoking — from the user's request when it names one, otherwise per rule #5 (its non-interactive fallback applies) — and pass it in the handoff; the sibling stops rather than guess a folder. There is deliberately no separate consent gate on the tenant writes this creates: the project and folder deployment fulfil the extraction request itself, and the one consequential choice — where the deployment lands (deployments have no delete API) — is exactly the folder decision rule #5 just routed. Do NOT drive `uip ixp` project or deployment commands from this skill instead of invoking it — the sibling's guides carry guardrails this skill does not. If `uipath-ixp` is unavailable in the session, fall back to `core.logic.mock` plus an Open Questions entry, exactly as when no documents were supplied.
5. **Always present finite decisions as a dropdown with a final "Something else" escape hatch.** Whenever the skill needs a decision (which solution, publish vs debug vs deploy, which connector, trigger type, or resource to bind, etc.), ask with the enumerated choices plus **"Something else"** last for free-form input; never ask open-ended in chat when a finite set of sensible defaults exists. If the user picks "Something else", parse their answer and continue. No structured-question facility on the harness → ask in chat as a numbered list with "Something else" last. Non-interactively (CI/headless, no user available) → take the marked recommended option, proceed, and record the decision prominently in the final report; if none is recommended, stop and report the open decision instead of guessing. Consent gates (`flow debug`, destructive operations) are never auto-answered — in non-interactive mode, stop and report the blocked step. These fallbacks define "ask the user" / "confirm with the user" wherever this skill's references require it.
5. **Always present finite decisions as a dropdown with a final "Something else" escape hatch.** Whenever the skill needs a decision (which solution, publish vs debug vs deploy, which connector, trigger type, or resource to bind, etc.), ask with the enumerated choices plus **"Something else"** last for free-form input; never ask open-ended in chat when a finite set of sensible defaults exists. If the user picks "Something else", parse their answer and continue. No structured-question facility on the harness → ask in chat as a numbered list with "Something else" last. Non-interactively (CI/headless, no user available) → take the marked recommended option, proceed, and record the decision prominently in the final report; if none is recommended, stop and report the open decision instead of guessing. Consent gates (destructive operations, tenant writes) are never auto-answered — in non-interactive mode, stop and report the blocked step; `flow debug` is not one of them, and is governed by the mandate rule above. These fallbacks define "ask the user" / "confirm with the user" wherever this skill's references require it.
<!--skill-flavor:user-question-options-extra:start-->
<!--skill-flavor:user-question-options-extra:end-->
<!--skill-flavor:project-creation:start-->
Expand Down
4 changes: 2 additions & 2 deletions skills/uipath-maestro-flow/references/author/brownfield.md
Original file line number Diff line number Diff line change
Expand Up @@ -81,8 +81,8 @@ Authoring ends here. For any selected option, read [operate/CAPABILITY.md](../op

| Option | What it does |
|---|---|
| **Publish to Studio Web** (default) | Push the solution to Studio Web so the user can visualize, edit, and publish from the browser. |
| **Debug the solution** | Execute the flow end-to-end against real systems. Confirm consent first because debug has real side effects (see the consent-before-debug rule in [SKILL.md](../../SKILL.md)). |
| **Publish to Studio Web** | Push the solution to Studio Web so the user can visualize, edit, and publish from the browser. |
| **Debug the solution** | Execute the flow end-to-end against real systems. Consent comes from the mandate, not from this menu — see the `flow debug` rule in [SKILL.md](../../SKILL.md). Selecting it here is the user asking for a run. |
| **Deploy to Orchestrator** | Pack and publish directly to Orchestrator (bypasses Studio Web). Only when explicitly chosen; see [/uipath:uipath-platform](/uipath:uipath-platform). |
| **Something else** | Last option. Accept free-form string input and act on it. |

Expand Down
4 changes: 2 additions & 2 deletions skills/uipath-maestro-flow/references/author/greenfield.md
Original file line number Diff line number Diff line change
Expand Up @@ -373,8 +373,8 @@ Authoring terminates here. Each option below hands off to Operate — read [oper

| Option | What it does |
| --- | --- |
| **Publish to Studio Web** (default) | Push the solution to Studio Web so the user can visualize, edit, and publish from the browser. |
| **Debug the solution** | Execute the flow end-to-end against real systems. Confirm consent first — debug has real side effects (see the consent-before-debug rule in [SKILL.md](../../SKILL.md)). |
| **Publish to Studio Web** | Push the solution to Studio Web so the user can visualize, edit, and publish from the browser. |
| **Debug the solution** | Execute the flow end-to-end against real systems. Consent comes from the mandate, not from this menu — see the `flow debug` rule in [SKILL.md](../../SKILL.md). Selecting it here is the user asking for a run. |
| **Deploy to Orchestrator** | Pack and publish directly to Orchestrator (bypasses Studio Web). Only when explicitly chosen — see [/uipath:uipath-platform](/uipath:uipath-platform). |
| **Something else** | Last option. Accept free-form string input and act on it (e.g., "just leave it", "pack but don't publish", "upload to a different tenant"). |

Expand Down
2 changes: 1 addition & 1 deletion skills/uipath-maestro-flow/references/operate/run.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,7 @@ Execute a flow on demand and monitor progress. Three modes: **debug** (controlle

## Debug — controlled end-to-end run

> **Confirm consent first.** `flow debug` executes the flow for real — sends emails, posts messages, calls APIs. See the consent-before-debug rule in [SKILL.md](../../SKILL.md). Do not run without explicit user authorization.
> **Consent comes from the mandate.** `flow debug` executes the flow for real — sends emails, posts messages, calls APIs. Run it when the request is for a flow that works; ask when the request stops at build or validate. Never debug a solution this run did not scaffold: debug overwrites the Studio Web solution matching the local `.uipx` `SolutionId`. See the `flow debug` consent rule in [SKILL.md](../../SKILL.md).

```bash
UIP_LOG_LEVEL=info uip maestro flow debug <path-to-project-dir> --output json
Expand Down
9 changes: 8 additions & 1 deletion tests/Makefile
Original file line number Diff line number Diff line change
@@ -1,9 +1,13 @@
.PHONY: help install all smoke smoke_rpa e2e tags
.PHONY: help install all smoke smoke_rpa e2e tags flow

SKILLS_REPO_PATH ?= $(shell cd .. && pwd)
VENV := .venv
CODER_EVAL := SKILLS_REPO_PATH=$(SKILLS_REPO_PATH) $(VENV)/bin/coder-eval
TASKS := $(shell find tasks -name '*.yaml' -type f)
# Flow's zero-shot config states the run is headless, which is false for the
# simulated tasks under interactive/. The exclusion is the config's contract,
# so it lives with the target rather than in a comment someone has to find.
FLOW_TASKS := $(shell find tasks/uipath-maestro-flow -name '*.yaml' -type f -not -path '*/interactive/*')
TASK_PARALLELISM ?= 1

help: ## Show available commands
Expand Down Expand Up @@ -38,6 +42,9 @@ smoke_rpa: ## Run all Windows RPA smoke tests (tempdir)
e2e: ## Run all end-to-end tests
$(CODER_EVAL) run $(TASKS) -e experiments/default.yaml --tags e2e -j $(TASK_PARALLELISM) -v

flow: ## Run the flow suite zero-shot (headless system prompt; excludes interactive/)
$(CODER_EVAL) run $(FLOW_TASKS) -e experiments/flow.yaml -j $(TASK_PARALLELISM) -v

tags: ## Run tests matching one or more tags: make tags TAGS="connector-feature" [EXPERIMENT=experiments/default.yaml]
@if [ -z "$(TAGS)" ]; then echo "Usage: make tags TAGS=\"tag1 tag2\" [EXPERIMENT=experiments/<file>.yaml]"; exit 2; fi
@TAGS="$(TAGS)" python3 -c 'import os,re,sys,glob; \
Expand Down
5 changes: 5 additions & 0 deletions tests/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -52,6 +52,10 @@ make smoke_rpa
# Run e2e-tagged tests under the default config
make e2e

# Run the flow suite zero-shot: experiments/flow.yaml states the run is headless,
# and the glob excludes interactive/, whose tasks have a simulated user
make flow

# Run tests matching a combination of tags (AND semantics — tasks must carry all listed tags) (defaults to experiments/default.yaml):
make tags TAGS="integration connector-feature"
# Optionally override the experiment config
Expand Down Expand Up @@ -176,6 +180,7 @@ Run-time caps live under `defaults.run_limits` (see coder_eval `RunLimits`).
| `default.yaml` | tempdir | Devs locally, ad-hoc runs | 200 | 1200s | 900s |
| `nightly.yaml` | docker | Nightly cron (`daily.sh`) | 200 | 1200s | 900s |
| `smoke.yaml` | docker | PR-gate smoke (Linux) | 40 | 900s | 900s |
| `flow.yaml` | docker | `make flow` / flow dispatches — nightly's runtime plus a system prompt stating the run is headless | 200 | 1200s | 900s |
| `smoke-windows.yaml` | tempdir | PR-gate smoke (Windows RPA only) | 40 | 900s | 900s |
| `activation.yaml` | tempdir | Skill activation classifier (benchmark) | 3 + early-stop | 360s | 120s |
| `same-ground-headtohead.yaml` | docker | Campaign-only local comparison arm | 200 | 1200s | 900s |
Expand Down
Loading
Loading