Repository navigation
[AIGTWY-4882] Add fresh-launch journeys and dedicated CUJ7 coverage - #930
Merged
Merged
Conversation
tt-le
requested review from
AarushiShah-db,
lilly-luo and
rohita5l
as code owners
October 1, 2026 16:07
This was referenced Oct 1, 2026
UG reviewThe visible changes cover unmanaged model discovery, explicit model forwarding, scoped inference evidence, and cleanup of seeded OS-managed settings. Review is limited to the shown changes because the supplied diff is truncated. No actionable findings.
Automated advisory review of |
tt-le
force-pushed
the
tt-le/aigtwy-4882-integ-tests
branch
from
October 6, 2026 14:17
9e39037 to
ccafeff
Compare
tt-le
force-pushed
the
tt-le/aigtwy-4882-integ-tests
branch
from
October 6, 2026 19:33
ccafeff to
10a5e5c
Compare
david-siqi-liu
left a comment
Collaborator
There was a problem hiding this comment.
Reviewed against main's tests/e2e_cuj/AGENTS.md and the CUJ 7 plan in the UG Configure E2E doc. Blocking items only:
P0
E2E CUJsis red: all four CUJ 7 cases error at setup withExisting machine-wide agent settings; use a clean disposable runner(https://github.com/databricks/unity-gateway/actions/runs/37523961500/job/112476241517). That run predates #993, which added the terminalug revertto the sharedcujteardown, so the leftover files likely came from CUJ 4 earlier in the job. A rebase onto current main (it also conflicts intests/README.md) and a greenE2E CUJsrun would confirm. Separately,live_sessionis per test and only assertsMANAGED_PATHSare absent at setup;test_case_07launchesug claudein a terminal with no revert afterwards, so if that launch writes machine-wide files, the next case fails the same way.- Coverage: the plan's "preexisting claude family defaults in managed settings are preserved when running
ug claudeand do NOT pick up ug's default" case isn't in the diff (noANTHROPIC_DEFAULT_*setup or assertion anywhere).
tt-le
force-pushed
the
tt-le/aigtwy-4882-integ-tests
branch
2 times, most recently
from
October 7, 2026 13:36
d6e0435 to
8485f49
Compare
tt-le
force-pushed
the
tt-le/aigtwy-4882-integ-tests
branch
from
October 7, 2026 17:50
5306eb8 to
b08b283
Compare
tt-le
force-pushed
the
tt-le/aigtwy-4882-integ-tests
branch
from
October 8, 2026 14:21
b08b283 to
1852d7d
Compare
tt-le
force-pushed
the
tt-le/aigtwy-4882-integ-tests
branch
2 times, most recently
from
October 8, 2026 19:53
3007eba to
986d455
Compare
sunishsheth2009
pushed a commit
to erifyc1/unity-gateway
that referenced
this pull request
Oct 8, 2026
…atabricks#1040) The e2e CUJs grew one copy of each helper per journey. This consolidates the shared helpers so CUJ 7 (databricks#930) and databricks#1014 / databricks#1015 can build on one version, and fixes three recurring CUJ flakes found in recent CI runs. Shared helpers (behavior preserving; every assertion, timeout and threshold is unchanged unless listed under behavior changes): - `UserSession.revert_machine_wide` replaces the class teardown revert and CUJ 3's two per-test copies (same terminal names and messages). - `base.make_workspace_client` and `base.bearer` replace the per-CUJ client and bearer code. CUJ 6 now honours `CLIENT_ID_ENV` / `CLIENT_SECRET_ENV`. - `Terminal` drives bare `ug` (`agent=`), and `Terminal.wait_until` is the one permission-rejecting wait for CUJ 2, 3 and 6. CUJ 1 keeps `AgentTerminal.wait_for_task`, which approves the read-only `find` prompt Claude emits in CI. - `evidence.assert_models` / `assert_served` replace CUJ 1's native-model check and the CUJ 2 / CUJ 3 served-response checks. Per-CUJ thresholds and matchers stay local. - `Workspace.agent_configs` splits a published config by agent for CUJ 4, 5 and 6. CUJ 5 reads the config through `workspace.config()` instead of a raw GET. - `helpers/poll.py` is the one bounded poll for the CUJ 1 and CUJ 6 trace waits (same budgets) and the `ug mcp list` retry. Trace SQL stays on `tests/integration/utils/sql.py`. - `helpers/mcp.py` (`registered_name`), `evidence.claude_file_task`, and shared model and header constants in `helpers/constants.py`. Reliability fixes: - MCP registration (Claude), failing 3 of 13 recent runs: Claude returned correct receipts but keyed them `mcp__server__tool`, and on this PR's first run echoed each raw `{"result": "<sha>"}` tool result. The prompt names the exact keys, and the check accepts exactly those two native shapes. The receipt values and key set must still match exactly. - `ug mcp list` health row (`claude:failed`, 1 of 13): Claude's single-shot probe raced the proxy cold start. The test polls up to 90s until the rows match, then runs the same exact assertion. - CUJ 4 Claude tasks use the absolute-path prompt CUJ 3 already used, so Claude no longer searches with `find` and hits a permission prompt. - Teardown leaks: a class that leaves `/etc` agent files still fails, and later classes now name that class and those paths instead of a generic dirty-runner error. - The CUJ bearer is refreshed per test (classes such as catalog discovery can outlive the M2M token). Every bearer a session installs stays redacted in artifacts. - ug review: `max_output_tokens` 4,000 to 16,000. High-effort reasoning consumed the old budget, leaving "returned no output text". The error now includes `incomplete_details.reason`. CI: - E2E CUJs run one runner per file, checked as `E2E CUJs · <CUJ_NAME>` (for example `E2E CUJs · CUJ 3 · UC model discovery`, names from the CUJ plan doc). CUJs share machine-wide `/etc` agent settings, so they cannot run concurrently on one runner. Each file still runs against its own read-only workspace, with `fail-fast: false`, and `All integration tests` requires every shard. - A small plan job discovers every `tests/e2e_cuj/test_*.py` and reads its module-level `CUJ_NAME`, so adding a CUJ needs no workflow edit (and no `.github/` owner approval). A file without `CUJ_NAME` fails the plan job; a unit contract test runs the same discovery script and also enforces unique names and the `test_cuj<N>_<topic>.py` / `CUJ <N> ·` convention. - The CUJ runners set `UV_HTTP_TIMEOUT=120` and `UV_HTTP_RETRIES=6`. All of them install from PyPI at once, and one shard already failed on a metadata fetch timeout. - CUJ files are renamed to `test_cuj<N>_<topic>.py`: - CUJ 1: `test_cuj1_explicit_models.py` (was `test_ug_managed_sentinel.py`) - CUJ 2: `test_cuj2_mps_mcp.py`, `test_cuj2_skills.py` - CUJ 3: `test_cuj3_models.py` (was `test_catalog_discovery.py`), `test_cuj3_mcp.py` (was `test_mcp_registration.py`), `test_cuj3_skills.py` - CUJ 4: `test_cuj4_smart_routing.py`; CUJ 5: `test_cuj5_budget_defaults.py`; CUJ 6: `test_cuj6_repeated_config.py` Behavior changes worth a look: - CUJ 2 Claude now rejects "Do you want to proceed?" like every other CUJ (it previously waited out the timeout). - The CUJ 6 Codex span poll's last query lands at the 300s deadline rather than up to 15s after it. Testing: - `tests/test_e2e_cuj_helpers.py` adds offline tests for the new helpers. Terminal tests run the real driver on a PTY against a fake `ug`. - Full unit suite and `--collect-only` (78 CUJ tests) pass locally. The E2E CUJs lane is the live check. CI timing, measured from job start and end times (main run 37821776097 vs this PR's run 37822258143): - E2E CUJs: 19.5 min as one job, now 7.8 min wall across 9 parallel shards. Total runner time is about the same (21.9 min). - Longest shard: CUJ 6 at 7.8 min. All others take 1.2 to 3.0 min. - Whole CI run: 20.0 min, now 18.7 min. `Full journeys · Claude` (about 13 to 15 min) is now the critical path, so cutting CUJ time further won't shorten CI. Not included: the workspace URL normalization in `src/ucode` from AIGTWY-4972 (a product behavior change, separate PR). This pull request and its description were written by Isaac. Co-authored-by: Isaac <no-reply@databricks.com>
tt-le
added a commit
that referenced
this pull request
Oct 9, 2026
…defaults Resolve the tests/README.md scenario-table conflict by keeping both the CUJ7 TUI/Codex preservation rows and main's `ug agents` rows (#896). Co-authored-by: Isaac <no-reply@databricks.com>
tt-le
force-pushed
the
tt-le/aigtwy-4882-integ-tests
branch
from
October 9, 2026 14:32
5894fc3 to
11fbb2b
Compare
david-siqi-liu
approved these changes
Oct 9, 2026
tt-le
force-pushed
the
tt-le/aigtwy-4882-integ-tests
branch
from
October 9, 2026 18:49
11fbb2b to
3a1440d
Compare
read_jsonl used Path.read_text() with the platform default encoding. On the Windows integration lane that is cp1252, which cannot decode UTF-8 model text such as a closing curly quote (0xE2 0x80 0x9D), so every Codex headless case that reads its session transcript through assert_completed_task_model failed with UnicodeDecodeError. Read transcripts as UTF-8 and add a reader regression test with non-ASCII content; it fails under a non-UTF-8 default without the fix. Co-authored-by: Isaac <no-reply@databricks.com>
`codex exec` cannot prompt for a Windows sandbox on first run and rejects
every shell command until one is configured, so on the Windows lane the
fresh-workspace task could not read its file ("file-system access is
blocked by the current sandbox/policy"). Call
choose_codex_windows_sandbox() as the other Codex headless journeys do;
it is a no-op elsewhere.
Co-authored-by: Isaac <no-reply@databricks.com>
tt-le
force-pushed
the
tt-le/aigtwy-4882-integ-tests
branch
from
October 9, 2026 20:04
d844ef6 to
1fbc5d7
Compare
tt-le
enabled auto-merge
October 10, 2026 02:33
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Add fresh-launch journeys and dedicated CUJ7 model discovery/inference coverage. Tracks AIGTWY-4882.
--workspacefile tasks and--providerlaunch checks in the Full agent suites. Provider cases consume existing services and do not perform inference.--modelbefore--and native argument forwarding after it. Codex covers model flags before the separator and within nativeexec. Require completed file answers and model attribution, not just startup or routing banners.test_cuj7_model_discovery.pydefinesCUJ_NAME, so automatic CI discovery runs it asE2E CUJs · CUJ 7 · Unmanaged model discoverywithout workflow changes.Workspace prerequisites
CUJ7 is pinned to
https://dbc-14e376e8-6541.cloud.databricks.com(workspace7474651332480761). It must publish no CodingAgentConfig, expose discoverablesystem.aimodels, and retainug_e2e.models.claude_haiku,ug_e2e.models.claude_sonnet, andug_e2e.models.gpt_luna.The shared
UG_CUJ_SP_CLIENT_ID/UG_CUJ_SP_CLIENT_SECRETidentity needs workspace access, read access to config/model discovery, and use access to those services. Only a short-lived bearer reaches agent processes. Each case checks that the workspace remains unmanaged. The preservation cases seed distinct Sonnet and other-family defaults on a clean disposable runner, assert Sonnet inference and unchanged defaults, then remove that local input.Regression verification: before and after #1037
Ran the same post-fix component test files (
tests/test_cli.py,tests/test_agent_claude.py, andtests/test_agent_codex.py) against isolated source snapshots of the merged #1037 fix and its immediate parent. Confirmed imports came from the intended snapshot in each run.abf85db12d39c6c78a30b503cc0c7632d26245c6de4b119d9e8678131b0e6bd15addeb136164d72fAll four
test_unmanaged_claude_model_reaches_native_launcherregressions (Claude/GPT IDs × separate/equals spellings) fail before the fix specifically because the native command omits--model MODEL; all four pass after it. The pre-fix revision's original component suite passes 959 tests, demonstrating that the previous assertions missed this boundary.The older live explicit-model test only placed
--modelafter--, bypassing UG's model-selection path. Coverage now includes the normal UG-owned option before the separator as well as supplemental native forwarding. The before/after comparison is component evidence, not a live CUJ or real-inference comparison.Current validation
git diff --check: passed./etc/claude-code/managed-settings.json, which can override isolated agent settings and model selection. The suite requires a clean disposable runner; CI must validate live inference and SP permissions.