Old behavior: any (user_id, device_id) duplicate returned 409 and the
per-user device cap was checked BEFORE the dup-check. Combined effect:
a client that already paired once but lost the Manager row (or just
wants to re-confirm on every startup) hit 409 or 403 forever, with no
way to recover except an admin DELETE.
New behavior:
- Same (user_id, device_id, pubkey) tuple → 200 with reused:true.
Lets bootstrap call pair on every login as an idempotent liveness
probe.
- Same (user_id, device_id) but different pubkey → 409 with explicit
"already paired with different key" message. Client treats this as
a signal to clear local identity and regenerate.
- Cap check moved AFTER the dup check so re-pair of an existing
device is never blocked by "device limit reached".
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
After 67225fd fixed the V2 chat 403 caused by missing SetupContextForToken,
the next probe call surfaced a new 403:
"token quota is not enough, token remain quota: \$0.000000,
need quota: \$0.001590"
Root cause: PairDevice initialised the new tokens row with
UnlimitedQuota=false and didn't set RemainQuota, so it defaulted to 0.
Every subsequent V2 chat then failed at pre-consume since the token had
no spendable budget — even though the user's actual User.Quota was
positive.
Device tokens aren't a billing boundary in our model; they're the
Ed25519 binding for a single client install. Quota belongs on the User
row. Flip UnlimitedQuota=true so the relay path consumes from
User.Quota directly, matching exactly what the legacy sk- bearer was
already doing (legacy tokens in this deployment are unlimited too).
Verified end-to-end via /tmp/v2_probe2.js after deploy: POST
/v1/messages with full V2 envelope returns HTTP 200 with the model's
reply.
V2 chat returned HTTP 403 with body
{"error":{"type":"new_api_error","message":"record not found ..."}}
even after Manager body_decrypt and Ed25519 verify both passed and the
device row was found. Root cause: the V2 dispatch in TokenAuth set
`id` + `token_id` directly via VerifyV2DeviceSignedRequest, called
applyTokenPolicyAndContext for IP/user/group checks, then jumped to
c.Next() — skipping SetupContextForToken entirely.
SetupContextForToken populates eight more keys the downstream relay
and billing pipeline expect:
token_key, token_name, token_unlimited_quota, token_quota,
token_model_limit_enabled, token_model_limit,
ContextKeyTokenGroup, ContextKeyTokenCrossGroupRetry
Without them, channel distribute / pre-consume / log_consume
silently misroute and a generic "record not found" leaks out as 403.
The legacy bearer path didn't have this bug because it always
finished with SetupContextForToken before c.Next().
Verified with /tmp/v2_probe2.js (after this deploys): POST
/v1/messages with full V2 envelope returns HTTP 200 with the model's
reply body.
Real test in C:\temp\v2_probe.js shows POST /api/devices/pair returning
HTTP 200 with body {"success":false, "message":"Unauthorized, invalid
access token"} when called with a Bearer sk- — the same sk- the OAuth
callback hands the client. Root cause: the route group used UserAuth(),
which only accepts a session cookie or a user JWT in Authorization, not
a relay-tier sk- bearer.
The OAuth-redirect flow (Heicode default) never produces a JWT — it
just hands cc-haha a sk-. So in production the pair call after
"一键登录" always 401'd, device-binding never activated, and V2
encryptedFetch silently fell back to legacy bearer for every request.
Fix: split the /devices route into two groups.
- /devices/* (list, rename, revoke): still UserAuth(). A sk- must
NOT be allowed to enumerate or revoke another device — that
would let an attacker with a stolen sk- delete the legitimate
owner's device binding.
- /devices/pair: TokenOrUserAuth(). Pair is the bootstrap step, by
definition no device key exists yet, so sk- IS the only credential
available on the OAuth-redirect flow.
TokenOrUserAuth calls c.Set("id", token.UserId) via its TokenAuth
fallback, so the PairDevice controller's c.GetInt("id") keeps working.
Verified by re-running v2_probe.js after deploy: pair returns
HTTP 200 success:true.
Constraint: Cloudflare can terminate long silent origin waits before upstream emits response bytes.\nRejected: Only pinging after upstream stream scanning starts | does not protect the first-response wait window.\nConfidence: high\nScope-risk: moderate\nDirective: Preserve per-channel ping disablement while keeping default stream relay heartbeats enabled.\nTested: go test ./relay/helper; go test ./relay/channel\nNot-tested: production Cloudflare edge rollout before deployment
Eliminate sk- bearer from the client wire entirely. V2 requests
authenticate via Ed25519 device signature (over a canonical that
binds method/path/timestamp/nonce/fingerprint/eph-pubkey/plaintext-
body-hash) and encrypt the request body with X25519 ECDH +
ChaCha20-Poly1305-AEAD. Server-issued sk- tokens still exist for
legacy callers during a 30-day deadline window; after the deadline
bare-bearer sk- on /v1/* is rejected.
What's new server-side:
- model/server_key.go + service/server_keys.go: long-lived X25519
keypair persisted in DB. Private half is AES-256-GCM-sealed with a
key derived from CRYPTO_SECRET so a SQL dump alone doesn't leak it.
Generated on first launch by main.go::EnsureServerECDHKey.
- common/crypto.go: SealWithCryptoSecret / UnsealWithCryptoSecret
helpers (AES-GCM); SafeWipe defense-in-depth zero-out.
- controller/server_pubkey.go + GET /api/server-pubkey: public
endpoint clients fetch at startup to obtain the ECDH pubkey.
- middleware/body_decrypt.go: ChaCha20-Poly1305 decrypt of V2 bodies.
AD binds device_id/timestamp/nonce/method/path so tampering any
fails AEAD verify. Replaces c.Request.Body with plaintext for
downstream relay handlers to consume unchanged.
- middleware/device_signature.go: new VerifyV2DeviceSignedRequest()
looks up token by device_id (not bearer) and verifies an extended
canonical that includes the ephemeral pubkey + plaintext body hash.
- middleware/auth.go::TokenAuth: dispatch on Content-Encoding header.
V2 path skips ValidateUserToken entirely. Legacy path adds a 30-day
/v1/* deadline knob.
- model/token.go::FindTokenByDeviceId: V2 lookup helper.
- controller/device.go::PairDevice: stops returning the sk in
responses. Client identifies itself by device_id + signature from
now on, no bearer needed.
- setting/operation_setting/device_binding_setting.go: new
LegacySkV1DeadlineMs knob (0 = disabled until operator sets it).
Backward compatibility: V1 device-signed tokens (those issued by
the earlier PairDevice that DID return a sk-) keep working through
the legacy bearer path; the existing V1 signature middleware still
runs for them. The 30-day deadline is opt-in until ops sets it.
Tests: V1 regression suite passes (middleware + common).
V2-specific tests come in a follow-up commit alongside the client
encryptedFetch wiring; deferring lets us land the server-side
plumbing first without coupling.