Skip to content

Upgrade Sub2API to v0.2.11 with preserved product contracts - #60

Merged
XiaoSiKe merged 129 commits into
mainfrom
codex/upgrade-v0.2.11
Oct 1, 2026
Merged

XiaoSiKe merged 129 commits into
mainfrom
codex/upgrade-v0.2.11

Conversation

@XiaoSiKe

@XiaoSiKe XiaoSiKe commented Oct 1, 2026

Copy link
Copy Markdown
Owner

Update Zero One API from the pinned Sub2API v0.2.8 baseline to stable v0.2.11, retaining the product pricing, saved upstream cost, authentication, redemption and active-monitoring contracts. The upstream merge adds model support, balance inflight reservations, API Key creation limits and gateway fixes. Existing keys and historical balances are preserved; there are no new SQL migrations.

The integration reuses the existing Anthropic response/usage owner, retains dashboard refresh bypass, avoids reviving retired Codex/passive-monitor code, and keeps published frontend bytes unchanged. Source compatibility builds fingerprint immutable assets instead of requiring evolved API source to reproduce historical output. Tests retain the customer group-and-account long-context gate and isolate the process-wide memory probe without changing its workload or budget. Axios, DOMPurify and brace-expansion receive the current security patches, with no audit exceptions.

The reviewed UI tag changes 19 version-badge snapshots only in meaningful content. Upstream reset-credit controls and remote Codex presentation remain deferred under the Approved UI Snapshot; their API/backend support is available. The deterministic change map and v0.2.11 release contract record the integration and data boundaries.

Validation:

  • Complete local make test (ordinary, unit, integration, backend lint, Console and Landing).
  • 135 policy regressions; UI/upstream/recorded-upgrade guards.
  • Console: 305 files, 2,215 tests; Landing: 18 files, 134 tests; current source builds and frozen adapters.
  • Full Chromium: 290 passed, 78 conditional skips.
  • Go reachability scan and patched frontend dependency audit.
  • Release maintenance, routing, Compose, Docker context, safe Edge rollback and shell checks.

Production deployment follows successful PR/current-main CI and immutable paired image publication. It requires encrypted off-host backups, an actual populated restore and old/new application rehearsal, stable-boundary fingerprints, Backend-first cutover, dedicated limited-key probes, and at least 30 minutes of observation. No production runtime has been switched by this PR.

Achordchan and others added 30 commits September 20, 2026 10:46
…ery instead of hiding mapped models

A single passthrough account made GetAvailableModels return nil for the whole
group, so every other account's model_mapping was dropped from /v1/models and
the Codex catalog. Treat passthrough accounts like unmapped ones: skip their
(possibly stale) mapping and let supplementUnmappedOpenAIModels add the
default set, while ordinary accounts' mappings still reach the list.

Refs #6647, #7517
#7395 made Codex (and Grok) imports store a base URL ending in /v1, but the
usage script still appends a hard-coded /v1/usage, so CC Switch requests
/v1/v1/usage, gets a 404 and shows "query failed" for every newly imported
Codex/Grok provider. Move the script into ccswitchImport.ts and strip an
existing trailing /v1 from {{baseUrl}} before appending /v1/usage, which
covers both the imported value and base URLs users edit either way later.
…tion

OpenAI candidate filtering reads the Redis metadata projection, but
filterSchedulerExtra dropped auto_reset_credit_enabled, the 5h/7d
reset-credit thresholds and codex_auto_reset_credit_state. The scheduler
therefore always saw automatic reset credits as disabled, so accounts
were paused at the ops quota auto-pause threshold even when a fresh
state showed an available reset credit. The "keep scheduling until the
reset-credit threshold" branch never took effect, and accounts with
credits stayed idle until their window reset naturally.

- Keep the four keys in the scheduler projection; the test round-trips
  the real JSON payload and resolves the credit config from it.
- Once visible, every scheduling pass notifies the auto-reset worker for
  paused credit-enabled accounts, and each notification reloads the
  account. Add a 30s per-account cooldown for the scheduler hot path
  only; other event-driven notifications are unchanged and the
  per-minute scan remains the fallback.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…pstream

- OpenCode Zen / Go 把 deepseek-* 模型转发给 DeepSeek,thinking mode 的
  reasoning_content 约束原样生效,上游 400 原文一字不差地回吐
- 原判定只认 platform=deepseek 与 base_url 指向 api.deepseek.com,
  opencode_go 账号漏出,Responses→CC 回退在缓存未命中时必踩 400
- 按「官方 OpenCode 上游 + deepseek-* 模型」补一条判定,不放宽成任意上游的
  deepseek-* 模型名,避免对无此约束的上游无谓改写请求体

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…odel

isReasoningModel matched the literal "gpt-5" prefix plus the GPT-6 Sol/Luna
spellings, so gpt-6 and gpt-6-astra (both already in the model catalog and
price list) fell into the sampling branch. Anthropic Messages and Chat
Completions requests converted to the Responses API then forwarded
temperature/top_p, which the Responses API rejects for these models with
"Unsupported parameter".

Key the check on the parsed GPT generation number instead, keeping the
existing Sol/Luna spelling match, so later families are covered without
another edit. Non-generation GPT families (gpt-image-1, gpt-audio) and
gpt-4.x are unaffected.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s empty

BufferedResponseAccumulator.SupplementResponseOutput rebuilt the output only
when the terminal response had an empty output array. Upstreams can also end
with a non-empty output whose message carries no (or only blank) output_text.
The buffered Chat Completions and Anthropic Messages paths then returned an
empty reply although the text had streamed, while usage still billed the
terminal output_tokens.

When the terminal output has no usable message text, fill the first empty
output_text part (or append one, or append a message item) from the
accumulated deltas. Non-empty terminal text stays authoritative, and nothing
is synthesized when no text streamed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The canonical Anthropic stream leaves tool_use input empty on
content_block_start and streams the arguments as input_json_delta.
Anthropic-compatible relays may instead put the complete arguments on
content_block_start and send no delta. The Anthropic-to-Responses stream
converter ignored that input, so function_call_arguments.done carried ""
and the closed item "{}", and clients executed the tool with no arguments.

Keep the inline input as a seed. A real input_json_delta supersedes it, so
canonical streams keep their exact event sequence and two JSON documents are
never concatenated. When no delta arrived by content_block_stop, emit the
seed as a single arguments delta before the done event, which still repeats
exactly what the deltas streamed. Empty, "{}" and null inputs keep the
existing "{}" fallback.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…hem unset

When a channel pricing entry left image_output_price empty,
GetModelPricingWithChannel and the pricing resolver zeroed the image
output price and marked it explicit, so image output tokens were billed
at $0. image_input_price followed the same "unset means zero" rule and
fell back to the text input price instead of the catalog image input
price. A channel that lists an image model such as gpt-image-2 in token
mode without filling in prices therefore generated images without
charging for image output tokens, while the available-channels view
(#2475) displayed the catalog image price for the same entry.

#2930 wanted an explicit 0 on the channel to stop falling back to the
catalog; the change applied the same treatment to unset values as well.

Treat the two image fields like every other token price field:
- Set on the channel (including 0): override the catalog. An explicit
  image_output_price keeps ImageOutputPriceExplicit, so 0 stays free.
- Left empty: keep the catalog price (LiteLLM or built-in fallback). If
  the catalog has none, computeTokenBreakdown still falls back to the
  text input/output price.

The admin UI saves an empty field as null and only an entered 0 as 0,
so the two cases stay distinguishable. The flat, interval and
GetModelPricingWithChannel paths now share applyChannelImagePriceOverrides.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…issing

The API-key Responses capability probe treated every 404 as conclusive and
stored responses_supported=false. The probe model is the first mapped
upstream model in lexical order, so on third-party accounts it is often an
auxiliary entry such as codex-auto-review; a model_not_found answer then
permanently routed the account away from /v1/responses although the
endpoint works.

A 400/404 whose error code/type or message says the probe model does not
exist or is unavailable is now inconclusive, so the stored capability is not
overwritten. Plain 404/405 still mean the endpoint is missing. The probe
also prefers a general gpt-* text model (not an image model) among the
mapped models before falling back to the first one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…search fallback

The personal-access-token /v1/alpha/search fallback turned any upstream
Responses SSE body into a 200 search result and billed one web search call.
Error, response.failed and response.incomplete events, a response.completed
with a non-success status or without its response, a truncated stream and a
bare [DONE] were all accepted, returning partial text as a successful and
billable search.

Parse the stream strictly: those shapes now return an error, so nothing is
written to the client and no WebSearchCalls result is produced. A successful
response.completed keeps the existing output and citation handling. The
existing PAT metadata test now feeds a Responses SSE stream, which is what
that upstream path returns.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
分组开启 claude_code_only 并配置降级分组时,/v1/messages 会在选号入口的
checkClaudeCodeRestriction 中把非 Claude Code 请求解析到降级分组;但
GatewayHandler.ChatCompletions / Responses 在 handler 层只要分组开了
claude_code_only 就直接返回 403,不看降级分组,请求到不了调度层:只会发
Chat Completions / Responses 的第三方客户端拿这类分组的 Key 一律 403。

- handler 层提前拒绝只在未配置降级分组时生效,拒绝的状态码与文案不变
- 配置了降级分组时请求继续,由 SelectAccountWithLoadAwareness 解析到降级分组
  选号,再经已有的 Chat Completions / Responses ⇄ Anthropic Messages 转换转发;
  service 层无改动
- 与 /v1/messages 的降级语义一致:选号与渠道限制预检按降级分组,计费与用量
  记录仍按 API Key 所属分组

新增 handler 单测覆盖两个入口在有/无降级分组时的行为。
Chat Completions file parts already become Responses input_file, but the
Antigravity path dropped them when converting to Anthropic and then to
Gemini. Map data-URI input_file to a document block and emit it as
inlineData so Gemini receives the PDF.

Fixes #7583
…ck off after query failures

With the credit config now visible to the scheduler, every scheduling pass
notifies the auto-reset worker for paused credit-enabled accounts.
evaluateAccount forced an upstream query whenever the reset threshold was
reached, regardless of the last result. Accounts with no credits left and
7d >= 95% were therefore polled on /wham/usage every ~15-20s across
instances (previously once a minute by the scanner), and during upstream
outages the failed state caused the same storm.

- At the reset threshold, skip the forced query while a no_credit result is
  still within the state TTL (10 min); re-query once it expires so newly
  granted credits are still picked up.
- After a query-stage failure (RESET_CREDIT_QUERY_FAILED,
  USAGE_SNAPSHOT_WRITE_FAILED, RESET_CREDIT_DETAILS_UNAVAILABLE), wait one
  minute before querying again, including snapshot-staleness triggers.
- Redemption-stage failures are unaffected and still retry immediately with
  the same idempotency key; accounts with available credits still query and
  redeem immediately.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Wei-Shaw and others added 29 commits September 29, 2026 15:43
…ing-trend

feat(dashboard): toggle recent usage between tokens and spending
…at-keepalive

fix(antigravity): keep compat streams alive before first content
…flict

fix(accounts): warn when whitelist model conflicts with mapping
…ting

fix(gateway): resolve composite routes for websocket aliases
…sage

fix(gateway): forward received chat stream usage
…dits-ui

feat(claude): 查询 Claude 重置次数(只读),UI 复用 Codex 次数按钮 / Read-only Claude reset credit count, reusing Codex count UI
fix: 移除 OpenAI 分组 Codex 配置中的模型目录依赖
feat: 新增可信用户风控白名单,保留审计并免除本地自动处罚
feat: Claude Code 专用分组仅展示受支持的客户端入口
POST /admin/accounts/:id/claude/reset-credits/redeem (Idempotency-Key
required, same middleware chain as the Codex reset-quota route). The server
re-queries eligibility and claims only the upstream next_grant_id when it is
redeemable; grant/org IDs never reach the client. Claims are serialized by
account and organization Redis leases, and a durable per-organization fence
in the idempotency table blocks further redemption after an unconfirmed
outcome (settles after 24h). Ported from upstream PR #7591.

Co-authored-by: korkin25 <eugeny.kuzakov@gmail.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Orange 「重置」 button mirrors the Codex cell: enabled only after a query
shows a redeemable credit, danger confirm, one idempotency key per
confirmation reused on retry, outcome-specific messages, then refreshes
credits and the usage row. Ported from upstream PR #7591.

Co-authored-by: korkin25 <eugeny.kuzakov@gmail.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- R1-1: an explicit upstream "unavailable" answer now fences the organization
  for 15 minutes (CLAUDE_RESET_UPSTREAM_UNAVAILABLE) instead of 24h; ambiguous
  outcomes keep the 24h fence. Frontend tells the operator to retry later.
- R1-2: rotate the idempotency key after definite pre-claim refusals (busy,
  not available, unresolved, upstream unavailable); map IDEMPOTENCY_IN_PROGRESS
  and IDEMPOTENCY_RETRY_BACKOFF to clear zh/en messages.
- R1-3: redeem reason and cleared windows pass only through explicit
  allowlists, so upstream values can never echo a grant ID to clients.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… R1-4)

- Map five_hour/seven_day/seven_day_overage_included to 5h/7d/7d overage
  (raw key fallback) in the confirm message, credit tooltip and success message.
- Align zh/en confirm copy with the Codex dialog; {count} is the count
  remaining after this use.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…enai-compat

fix(gateway): Claude Code 限制分组的 OpenAI 兼容入口改走降级分组,不再直接 403
- 新增 api_key_create.max_active_per_user(默认 200)与
  api_key_create.max_per_user_per_hour(默认 20),0 表示不限制
- 创建频率按固定一小时窗口统计,自定义与自动生成的 Key 统一计入
- 创建计数独立于 Key 生命周期,删除或修改 Key 不再重置计数
- 移除不再使用的 DeleteCreateAttemptCount 缓存方法
feat(api-key): 增加 API Key 创建数量与频率限制
…reservation

fix(billing): Redis 在途预留,防止余额模式并发透支
feat(codex): 支持 model_catalog_url 远程模型目录配置
feat(openai): support GPT-6.1 Sol, Codex plans and Astra Ultrafast
feat(claude): 手动兑换 Claude 重置次数 / Manually redeem Claude limit resets
@XiaoSiKe
XiaoSiKe merged commit 6a540fc into main Oct 1, 2026
12 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.