Skip to content

[AIGTWY-4882] Add fresh-launch journeys and dedicated CUJ7 coverage - #930

Merged
tt-le merged 16 commits into
mainfrom
tt-le/aigtwy-4882-integ-tests
Oct 10, 2026
Merged

tt-le merged 16 commits into
mainfrom
tt-le/aigtwy-4882-integ-tests

Conversation

@tt-le

@tt-le tt-le commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Add fresh-launch journeys and dedicated CUJ7 model discovery/inference coverage. Tracks AIGTWY-4882.

  • Keep fresh --workspace file tasks and --provider launch checks in the Full agent suites. Provider cases consume existing services and do not perform inference.
  • Add five CUJ7 journeys / eight collected cases: configured Claude picker discovery, configured Codex app-server discovery, fresh scoped file tasks for both agents, and preservation of preexisting Claude OS-managed family defaults during explicit model selection.
  • Exercise Claude's UG-owned --model before -- and native argument forwarding after it. Codex covers model flags before the separator and within native exec. Require completed file answers and model attribution, not just startup or routing banners.
  • Rebase onto [AIGTWY-4972] Consolidate CUJ helpers, fix flakes, parallelize CUJs #1040 and reuse its shared bearer, cleanup, served-response checks, and absolute-path Claude prompts. test_cuj7_model_discovery.py defines CUJ_NAME, so automatic CI discovery runs it as E2E CUJs · CUJ 7 · Unmanaged model discovery without workflow changes.
  • Keep CUJ7 sessions function-scoped for fresh-launch coverage; the shared class fixture requires a published config, while this workspace must remain unmanaged. Preserve existing test names. No production changes, fixture provisioning, or remote configuration mutations.

Workspace prerequisites

CUJ7 is pinned to https://dbc-14e376e8-6541.cloud.databricks.com (workspace 7474651332480761). It must publish no CodingAgentConfig, expose discoverable system.ai models, and retain ug_e2e.models.claude_haiku, ug_e2e.models.claude_sonnet, and ug_e2e.models.gpt_luna.

The shared UG_CUJ_SP_CLIENT_ID / UG_CUJ_SP_CLIENT_SECRET identity needs workspace access, read access to config/model discovery, and use access to those services. Only a short-lived bearer reaches agent processes. Each case checks that the workspace remains unmanaged. The preservation cases seed distinct Sonnet and other-family defaults on a clean disposable runner, assert Sonnet inference and unchanged defaults, then remove that local input.

Regression verification: before and after #1037

Ran the same post-fix component test files (tests/test_cli.py, tests/test_agent_claude.py, and tests/test_agent_codex.py) against isolated source snapshots of the merged #1037 fix and its immediate parent. Confirmed imports came from the intended snapshot in each run.

Source revision Identical updated component suite
Before fix: abf85db12d39c6c78a30b503cc0c7632d26245c6 957 passed, 6 failed
After fix (#1037): de4b119d9e8678131b0e6bd15addeb136164d72f 963 passed

All four test_unmanaged_claude_model_reaches_native_launcher regressions (Claude/GPT IDs × separate/equals spellings) fail before the fix specifically because the native command omits --model MODEL; all four pass after it. The pre-fix revision's original component suite passes 959 tests, demonstrating that the previous assertions missed this boundary.

The older live explicit-model test only placed --model after --, bypassing UG's model-selection path. Coverage now includes the normal UG-owned option before the separator as well as supplemental native forwarding. The before/after comparison is component evidence, not a live CUJ or real-inference comparison.

Current validation

  • After rebasing onto [AIGTWY-4972] Consolidate CUJ helpers, fix flakes, parallelize CUJs #1040 and migrating helpers: 1,192 component/helper/contract tests passed.
  • All eight CUJ7 cases collect; the CI discovery contract recognizes the file and its unique shard name.
  • Repository-wide Ruff checks and formatting, source type checks, and git diff --check: passed.
  • Live inference was not run locally: this host has /etc/claude-code/managed-settings.json, which can override isolated agent settings and model selection. The suite requires a clean disposable runner; CI must validate live inference and SP permissions.

@github-actions

github-actions Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

UG review

The visible changes cover unmanaged model discovery, explicit model forwarding, scoped inference evidence, and cleanup of seeded OS-managed settings. Review is limited to the shown changes because the supplied diff is truncated.

No actionable findings.

Some patches were unavailable or truncated to fit the review context. Treat this as a partial review.

Automated advisory review of 1fbc5d7955fb using Lilly's UG review rubric. It does not approve or block this PR.

Comment thread scripts/run_integration.py Outdated
@tt-le
tt-le force-pushed the tt-le/aigtwy-4882-integ-tests branch from 9e39037 to ccafeff Compare October 6, 2026 14:17
@tt-le tt-le changed the title [AIGTWY-4882] Add fresh launch integration journeys [AIGTWY-4882] Add fresh-launch journeys and dedicated CUJ7 coverage Oct 6, 2026
@tt-le
tt-le force-pushed the tt-le/aigtwy-4882-integ-tests branch from ccafeff to 10a5e5c Compare October 6, 2026 19:33

@david-siqi-liu david-siqi-liu left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed against main's tests/e2e_cuj/AGENTS.md and the CUJ 7 plan in the UG Configure E2E doc. Blocking items only:

P0

  1. E2E CUJs is red: all four CUJ 7 cases error at setup with Existing machine-wide agent settings; use a clean disposable runner (https://github.com/databricks/unity-gateway/actions/runs/37523961500/job/112476241517). That run predates #993, which added the terminal ug revert to the shared cuj teardown, so the leftover files likely came from CUJ 4 earlier in the job. A rebase onto current main (it also conflicts in tests/README.md) and a green E2E CUJs run would confirm. Separately, live_session is per test and only asserts MANAGED_PATHS are absent at setup; test_case_07 launches ug claude in a terminal with no revert afterwards, so if that launch writes machine-wide files, the next case fails the same way.
  2. Coverage: the plan's "preexisting claude family defaults in managed settings are preserved when running ug claude and do NOT pick up ug's default" case isn't in the diff (no ANTHROPIC_DEFAULT_* setup or assertion anywhere).

@tt-le
tt-le force-pushed the tt-le/aigtwy-4882-integ-tests branch 2 times, most recently from d6e0435 to 8485f49 Compare October 7, 2026 13:36
@tt-le
tt-le requested a review from david-siqi-liu October 7, 2026 14:38
Comment thread tests/test_integration_contract.py Outdated
@tt-le
tt-le force-pushed the tt-le/aigtwy-4882-integ-tests branch from 5306eb8 to b08b283 Compare October 7, 2026 17:50
@tt-le
tt-le force-pushed the tt-le/aigtwy-4882-integ-tests branch from b08b283 to 1852d7d Compare October 8, 2026 14:21
@tt-le
tt-le requested a review from david-siqi-liu October 8, 2026 15:02
@tt-le
tt-le force-pushed the tt-le/aigtwy-4882-integ-tests branch 2 times, most recently from 3007eba to 986d455 Compare October 8, 2026 19:53
sunishsheth2009 pushed a commit to erifyc1/unity-gateway that referenced this pull request Oct 8, 2026
…atabricks#1040)

The e2e CUJs grew one copy of each helper per journey. This consolidates
the shared helpers so CUJ 7 (databricks#930) and databricks#1014 / databricks#1015 can build on one
version, and fixes three recurring CUJ flakes found in recent CI runs.

Shared helpers (behavior preserving; every assertion, timeout and
threshold is unchanged unless listed under behavior changes):
- `UserSession.revert_machine_wide` replaces the class teardown revert
and CUJ 3's two per-test copies (same terminal names and messages).
- `base.make_workspace_client` and `base.bearer` replace the per-CUJ
client and bearer code. CUJ 6 now honours `CLIENT_ID_ENV` /
`CLIENT_SECRET_ENV`.
- `Terminal` drives bare `ug` (`agent=`), and `Terminal.wait_until` is
the one permission-rejecting wait for CUJ 2, 3 and 6. CUJ 1 keeps
`AgentTerminal.wait_for_task`, which approves the read-only `find`
prompt Claude emits in CI.
- `evidence.assert_models` / `assert_served` replace CUJ 1's
native-model check and the CUJ 2 / CUJ 3 served-response checks. Per-CUJ
thresholds and matchers stay local.
- `Workspace.agent_configs` splits a published config by agent for CUJ
4, 5 and 6. CUJ 5 reads the config through `workspace.config()` instead
of a raw GET.
- `helpers/poll.py` is the one bounded poll for the CUJ 1 and CUJ 6
trace waits (same budgets) and the `ug mcp list` retry. Trace SQL stays
on `tests/integration/utils/sql.py`.
- `helpers/mcp.py` (`registered_name`), `evidence.claude_file_task`, and
shared model and header constants in `helpers/constants.py`.

Reliability fixes:
- MCP registration (Claude), failing 3 of 13 recent runs: Claude
returned correct receipts but keyed them `mcp__server__tool`, and on
this PR's first run echoed each raw `{"result": "<sha>"}` tool result.
The prompt names the exact keys, and the check accepts exactly those two
native shapes. The receipt values and key set must still match exactly.
- `ug mcp list` health row (`claude:failed`, 1 of 13): Claude's
single-shot probe raced the proxy cold start. The test polls up to 90s
until the rows match, then runs the same exact assertion.
- CUJ 4 Claude tasks use the absolute-path prompt CUJ 3 already used, so
Claude no longer searches with `find` and hits a permission prompt.
- Teardown leaks: a class that leaves `/etc` agent files still fails,
and later classes now name that class and those paths instead of a
generic dirty-runner error.
- The CUJ bearer is refreshed per test (classes such as catalog
discovery can outlive the M2M token). Every bearer a session installs
stays redacted in artifacts.
- ug review: `max_output_tokens` 4,000 to 16,000. High-effort reasoning
consumed the old budget, leaving "returned no output text". The error
now includes `incomplete_details.reason`.

CI:
- E2E CUJs run one runner per file, checked as `E2E CUJs · <CUJ_NAME>`
(for example `E2E CUJs · CUJ 3 · UC model discovery`, names from the CUJ
plan doc). CUJs share machine-wide `/etc` agent settings, so they cannot
run concurrently on one runner. Each file still runs against its own
read-only workspace, with `fail-fast: false`, and `All integration
tests` requires every shard.
- A small plan job discovers every `tests/e2e_cuj/test_*.py` and reads
its module-level `CUJ_NAME`, so adding a CUJ needs no workflow edit (and
no `.github/` owner approval). A file without `CUJ_NAME` fails the plan
job; a unit contract test runs the same discovery script and also
enforces unique names and the `test_cuj<N>_<topic>.py` / `CUJ <N> ·`
convention.
- The CUJ runners set `UV_HTTP_TIMEOUT=120` and `UV_HTTP_RETRIES=6`. All
of them install from PyPI at once, and one shard already failed on a
metadata fetch timeout.
- CUJ files are renamed to `test_cuj<N>_<topic>.py`:
- CUJ 1: `test_cuj1_explicit_models.py` (was
`test_ug_managed_sentinel.py`)
  - CUJ 2: `test_cuj2_mps_mcp.py`, `test_cuj2_skills.py`
- CUJ 3: `test_cuj3_models.py` (was `test_catalog_discovery.py`),
`test_cuj3_mcp.py` (was `test_mcp_registration.py`),
`test_cuj3_skills.py`
- CUJ 4: `test_cuj4_smart_routing.py`; CUJ 5:
`test_cuj5_budget_defaults.py`; CUJ 6: `test_cuj6_repeated_config.py`

Behavior changes worth a look:
- CUJ 2 Claude now rejects "Do you want to proceed?" like every other
CUJ (it previously waited out the timeout).
- The CUJ 6 Codex span poll's last query lands at the 300s deadline
rather than up to 15s after it.

Testing:
- `tests/test_e2e_cuj_helpers.py` adds offline tests for the new
helpers. Terminal tests run the real driver on a PTY against a fake
`ug`.
- Full unit suite and `--collect-only` (78 CUJ tests) pass locally. The
E2E CUJs lane is the live check.

CI timing, measured from job start and end times (main run 37821776097
vs this PR's run 37822258143):
- E2E CUJs: 19.5 min as one job, now 7.8 min wall across 9 parallel
shards. Total runner time is about the same (21.9 min).
- Longest shard: CUJ 6 at 7.8 min. All others take 1.2 to 3.0 min.
- Whole CI run: 20.0 min, now 18.7 min. `Full journeys · Claude` (about
13 to 15 min) is now the critical path, so cutting CUJ time further
won't shorten CI.

Not included: the workspace URL normalization in `src/ucode` from
AIGTWY-4972 (a product behavior change, separate PR).

This pull request and its description were written by Isaac.

Co-authored-by: Isaac <no-reply@databricks.com>
tt-le added a commit that referenced this pull request Oct 9, 2026
…defaults

Resolve the tests/README.md scenario-table conflict by keeping both the
CUJ7 TUI/Codex preservation rows and main's `ug agents` rows (#896).

Co-authored-by: Isaac <no-reply@databricks.com>
@tt-le
tt-le force-pushed the tt-le/aigtwy-4882-integ-tests branch from 5894fc3 to 11fbb2b Compare October 9, 2026 14:32
@tt-le
tt-le force-pushed the tt-le/aigtwy-4882-integ-tests branch from 11fbb2b to 3a1440d Compare October 9, 2026 18:49
tt-le and others added 16 commits October 9, 2026 20:03
read_jsonl used Path.read_text() with the platform default encoding.
On the Windows integration lane that is cp1252, which cannot decode
UTF-8 model text such as a closing curly quote (0xE2 0x80 0x9D), so
every Codex headless case that reads its session transcript through
assert_completed_task_model failed with UnicodeDecodeError.

Read transcripts as UTF-8 and add a reader regression test with
non-ASCII content; it fails under a non-UTF-8 default without the fix.

Co-authored-by: Isaac <no-reply@databricks.com>
`codex exec` cannot prompt for a Windows sandbox on first run and rejects
every shell command until one is configured, so on the Windows lane the
fresh-workspace task could not read its file ("file-system access is
blocked by the current sandbox/policy"). Call
choose_codex_windows_sandbox() as the other Codex headless journeys do;
it is a no-op elsewhere.

Co-authored-by: Isaac <no-reply@databricks.com>
Recount from collection on top of #1051/#1062/#1039: 85 live cases
(39 Claude, 46 Codex), 11 marked TUI, OpenCode adding one live headless
and two managed_fixture cases, and 129 executions (126 with Claude and
Codex selected).

Co-authored-by: Isaac <no-reply@databricks.com>
@tt-le
tt-le force-pushed the tt-le/aigtwy-4882-integ-tests branch from d844ef6 to 1fbc5d7 Compare October 9, 2026 20:04
@tt-le
tt-le enabled auto-merge October 10, 2026 02:33
@tt-le
tt-le added this pull request to the merge queue Oct 10, 2026
Merged via the queue into main with commit 3754d13 Oct 10, 2026
190 of 202 checks passed
@tt-le
tt-le deleted the tt-le/aigtwy-4882-integ-tests branch October 10, 2026 02:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ug-review Run the automated UG review

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants