Skip to content

feat(codex): guide recovery when automatic credential renewal becomes impossible #5783

Description

@coygeek

Summary

Turn terminal Codex OAuth renewal failures into an actionable recovery workflow before a still-usable access token expires. Preserve working inference, make the affected account and deadline visible through supported operator surfaces, and guide the minimum required human authorization through to verified runtime use of the refreshed credential.

Problem to solve

A two-account deployment continued serving requests while one account's proactive refresh repeatedly failed with refresh_token_reused. Logs explicitly said the access token was being retained because it had not expired. The operator had to inspect private credential metadata and repeated warnings over SSH, determine the expiry time, identify the correct proxy login command, complete browser authorization, inspect a watcher event, and correlate a streaming request with the selected credential. Ordinary successful inference masked the future renewal failure.

The inspected implementation already retains unexpired access tokens after refresh failure. CLI and management OAuth login mechanisms also exist. The missing outcome is a connected, actionable recovery lifecycle that distinguishes current usability from future renewability; this is not a request to add those primitives again or infer that a fresh login proves a future refresh cycle.

Proposed behavior

When the provider explicitly rejects a refresh grant as terminal, record that independent sign-in is required while tracking whether the existing access token can still serve requests. Expose the account's local instance, safe identity, last failure, access-expiry time, and appropriate next action. Use structured state and a deduplicated event that official CLI/TUI or management integrations can present; do not require an always-running desktop GUI.

Stop repeatedly submitting an unchanged grant known to be terminally rejected. Continue bounded retries for genuinely transient failures, and resume normal renewal when a verified credential change supplies a new grant. Preserve usable access and the rest of the pool rather than prematurely evicting a working credential.

Provide an account-targeted recovery action using the supported browser or device flow on the affected instance. Explain that a login on another machine or the client's separate direct login does not repair this instance. The user completes passwords, MFA, and provider consent. After credential persistence, confirm runtime adoption and distinguish authentication saved, loaded, inference verified, quota blocked, and renewal later verified. A small inference check should be explicit or permitted by an operator-configured validation policy, attributed to the intended credential, and consume no unrelated task context.

Acceptance criteria

  • With a valid access token and a mocked terminal refresh rejection, inference remains usable, renewal is marked as requiring sign-in, the expiry deadline is visible, and repeated unchanged failures do not produce a notification or refresh-request storm.
  • A transient refresh failure follows bounded retry behavior and does not incorrectly require permanent reauthentication.
  • At access expiry, the unrenewable credential becomes unavailable while other eligible accounts continue serving requests. The operator can identify the required action without reading raw credential JSON or searching logs.
  • Recovery on one instance targets the intended provider/account. Cancelling the login or authenticating a different account does not falsely mark the affected credential repaired.
  • After a fresh grant is saved, the running service confirms adoption. A permitted streaming validation counts as verified only when the intended credential completes the response; another account's success, HTTP 200 without a terminal completion, or a model-list response is insufficient.
  • Successful sign-in clears superseded authentication failures, but does not blindly clear genuine quota cooldowns or claim that upstream allowance was reset. A fresh grant awaiting its first automatic renewal remains distinguishable from one whose renewal was actually verified.
  • Recovery state and notifications contain no tokens, authorization codes, callback parameters, or reusable browser-login URLs, and work through the deployment's existing access boundaries.

Affected area

Codex OAuth renewal state, operator-visible account health, existing authentication entry points, and post-login runtime verification. This feature is independently useful when credential replacement is in place and does not require a new file-deduplication mechanism.

Non-goals

  • Silently performing passwords, MFA, consent, or other provider-required human approval.
  • Making copied rotating refresh grants safe to share across machines or extracting credentials from another application.
  • Quota-aware selection, cooldown revalidation, global service restarts, or automatic purchases of additional allowance.
  • Replacing the existing behavior that preserves a usable access token after proactive renewal fails.

Supporting context

Local verification on September 12, 2026 observed terminal refresh warnings beginning at 17:56:42 PDT, with access expiry scheduled for the following day. Independent browser sign-in saved a fresh grant at 21:02:53 PDT; the watcher loaded it, and a request attributed to the repaired account produced the exact expected text and terminal response.completed at 21:04:30 PDT. Those manual steps define the requested recovery outcome. The initial cause of the rejected grant was not isolated; prior copying of rotating credentials between machines is relevant context, not a proven proxy defect.

At inspected commit ac02da6c05e18f465aa7e3ed5b0a65a2f060917d, conductor_refresh.go already implements proactive refresh and retention of still-valid access. Management routes already expose refresh, Codex login, auth-status, and cancellation operations. Compose these supported capabilities rather than treating login itself as missing.

Related work includes #5095 and #5311 for retaining unexpired access, #4363 for terminal failure quarantine, and #5766 for stopping duplicate retries within one refresh operation. The latter does not by itself provide an operator recovery workflow across later scheduled attempts. Coordinate with those changes; this request does not assert they are absent or duplicate their fixes.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    pendingWaiting for research

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions