Skip to content

Add end-to-end Claude and Codex integration CUJs - #568

Closed
rohita5l wants to merge 7 commits into
mainfrom
rohit/integration-tui-tests
Closed

rohita5l wants to merge 7 commits into
mainfrom
rohit/integration-tui-tests

Conversation

@rohita5l

@rohita5l rohita5l commented Sep 11, 2026 •

Copy link
Copy Markdown
Collaborator

What this adds

The existing e2e suite patches substantial setup, which misses problems in fresh consumer installs and actual Claude/Codex launches. This adds a separate integration suite that builds a fresh ug wheel (or installs an exact release), installs it in a new application environment, and drives the real CLI and agent binaries against the existing e2e workspace.

All CUJs are descriptive top-level tests with Scenario: and Expected: docstrings and visible configure/launch/assertion steps. There are 45 live cases, including 10 interactive TUI journeys, plus 3 installation checks:

  • Databricks Hosted and Anthropic/OpenAI MPS configuration, followed by completed TUI file tasks.
  • First-prompt and subagent smart routing, explicit-model bypass, normal exit and reopen.
  • Headless prompt/model arguments, caller settings/hooks, help and parser-error forwarding, and real Codex app-server initialization.
  • Repeated configuration, preservation of user settings, revert, and rejected real-workspace credentials.

Success requires observable results: unpredictable file contents in completed assistant answers, real routing records/child transcripts, exit statuses, preserved configuration, or valid protocol responses. No mocks, monkeypatching, fake executables/services, application imports, seeded ug/onboarding state, or production changes. MCP/skills functionality and broader configure options remain deferred.

Shared process, terminal and evidence helpers plus Docker files live in tests/integration/utils/. The coverage/gaps matrix is in tests/README.md; tests/AGENTS.md and tests/CLAUDE.md define how agents must add, modify and remove tests.

CI and reproduction

The Integration workflow has installation checks, a workspace guard, and one CUJs job using fresh consumer dependency resolution. It reuses UCODE_TEST_WORKSPACE and DATABRICKS_BEARER, and requires https://eng-ml-inference-team-us-east-1.cloud.databricks.com. There is no separate CI model-discovery job. Real ug configure discovers workspace configuration; explicit-model cases record their chosen system.ai model.

Agent versions are pinned to Claude 2.1.268 and Codex 0.154.0. Manual runs accept ug/agent versions and an optional dependency constraint. Evidence includes wheel/version hashes, dependency graphs, JUnit results, redacted commands, terminal screens/actions and native agent transcripts. The README documents local replay using the archived wheel and dependency graph.

CI runs directly on fresh Ubuntu 22.04 VMs. Local Colima/Docker is optional; it is not the current CI execution environment. The real agent sandbox remains enabled.

Validation

The prior CI run installed the fresh wheel successfully and exposed harness cleanup and Codex child-evidence issues addressed here. The new run will also determine whether the OpenAI MPS default-model failure persists after managed settings are restored through real interactive ug revert. Strict app-server stdout and generated-config cleanup assertions remain in the suite.

@rohita5l rohita5l changed the title Add real installed-agent integration tests and TUI boot coverage Add end-to-end Claude and Codex integration CUJs Sep 11, 2026
@rohita5l
rohita5l marked this pull request as ready for review September 11, 2026 20:30
rohita5l and others added 3 commits September 14, 2026 11:36
Two product bugs surfaced by the integration suite:

- `ug codex app-server --listen stdio://` printed status lines like
  "Databricks auth already available" to stdout, corrupting the JSON-RPC
  stream. ug's Rich output now moves to stderr for that launch; fd 1
  stays untouched for the exec'd agent.
- A second configure or a launch-time model-preference clear backed up
  the config file ucode itself generated, so `ug revert` restored that
  snapshot instead of deleting it. Backups now only capture files that
  predate ucode's management of the tool, so revert removes
  `~/.claude/ucode-settings.json` and `~/.codex/ucode.config.toml`.

Co-authored-by: Isaac <no-reply@databricks.com>
The configure MPS CUJs can exceed a 180s task wait on a slow gateway
round trip, and the live CI job's 45-minute timeout fired before the
runner's own 3600s pytest deadline could produce its report.

Co-authored-by: Isaac <no-reply@databricks.com>
Resolved conflicts by keeping both sides: the app-server stdout redirect
and generated-config backup guards alongside main's --parent validation,
MPS model discovery, and get_provider_service moves.

Co-authored-by: Isaac <no-reply@databricks.com>
@rohita5l

Copy link
Copy Markdown
Collaborator Author

Superseded by native GitHub stack #607, split into four PRs in merge order:

  1. Fix app-server stdout and generated-config cleanup #603 — Fix app-server stdout and generated-config cleanup
  2. Add an isolated installed-product test harness #604 — Add an isolated installed-product test harness
  3. Add complete Claude and Codex user journeys #605 — Add complete Claude and Codex user journeys
  4. Parallelize smoke, full journeys, and agent CI #606 — Parallelize smoke, full journeys, and agent CI

The combined tracked contents match this PR plus the requested smoke/full CI parallelization and check-name changes exactly. Full integration coverage remains enabled on PRs. Merge from the bottom up after #606's full CI passes. The original branch is preserved for reference.

@rohita5l rohita5l closed this Sep 14, 2026
lilly-luo pushed a commit that referenced this pull request Sep 14, 2026
This isolates the production fixes from #568. Codex app-server launches
send ug status output to stderr so stdout remains a valid JSON-RPC
stream. Repeated configure and launch-time preference cleanup no longer
back up ug-generated config files, allowing revert to delete them while
preserving pre-existing user files.

Includes regression tests for both behaviors.

Validation: the 557 agent/CLI unit tests pass from this layer's
unchanged product files; the combined stack passes 561 focused tests.
Live gateway tests are exercised by the final stack layer's CI.

Layer 1 of 4. Merge the stack from bottom to top after the final layer's
full CI passes.
rohita5l added a commit that referenced this pull request Sep 15, 2026
Runs smoke and the full installed-product suite on relevant
same-repository PRs and main pushes. Smoke covers four Hosted
configure/TUI and headless cases in two agent jobs. All 45 full journeys
run across eight independent runners. The existing e2e suite is split
into seven gateway/agent jobs, with each agent job installing its
required CLI.

Both matrices disable fail-fast and report aggregate results. Check
names are Unit tests, Gateway API tests, Agent launch tests · Agent,
Smoke journeys · Agent, and Full journeys · Agent · Group. Full coverage
does not depend on a label or manual request. Small `test` and `e2e`
compatibility gates preserve the exact contexts required by current
repository rules, without changing those rules.

Validation:
- 561 focused unit tests pass with the caller's smart-routing
environment flag removed.
- Real collection proves all 45 integration cases and all 39 existing
e2e cases are covered exactly once within each full suite.
- Smoke selects four cases; manual TUI selection covers all ten TUI
cases.
- Aggregate success/failure/skip/cancellation checks pass for every
dispatch mode.
- actionlint, Ruff, formatting, and git diff --check pass.

The final tracked tree matches #568 plus the requested CI changes
byte-for-byte across all 39 changed files. CI will supply live gateway
results; none are claimed from local collection. The existing tracing
test retains its pre-existing skip.

Live CI at `5f5c074`:
- Unit tests, Gateway API tests, all six Agent launch jobs, both Smoke
journey jobs, Installation tests, and the required-check compatibility
gates pass.
- [Full integration
run](https://github.com/databricks/unity-gateway/actions/runs/34869281749):
**44 passed, 1 failed** across the eight full shards. The run finished
in roughly 3 minutes; the slowest full shard took 2m33s.
- The remaining failure is `test_ug_configure_codex_openai_mps`: the
existing CI workspace returns HTTP 404 / `ENDPOINT_NOT_FOUND` for Codex
model discovery and says `codex/v1/models is not enabled for this
workspace`. No test retry, skip, or weaker assertion was added.
- Layers #604 and #605 still have failed earlier e2e checks with Codex
429 rate-limit responses. This stack is not ready to merge until those
checks and the full integration gate pass.

Layer 4 of 4. Merge the stack from bottom to top after this layer's full
CI passes. Required-check compatibility is preserved for the lower stack
layers and other active PRs.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant