Verification-Gated Routing¶
Verification-Gated Routing (VGR) generates a candidate on a local tier, gathers bounded evidence about that exact candidate, and serves it only when the policy licenses a local commit. All other outcomes use the cloud tier.
Tool-calling turns continue on the local tier without being mistaken for a terminal answer. Missing context, unavailable verifiers, malformed verdicts, and expired decision deadlines fail closed to cloud.
Configure a route¶
Declare local, cloud, and optional verifier targets using the normal deployment schema, then reference their target names:
[routes.assistant]
id = "assistant"
type = "vgr"
local_target = "local"
cloud_target = "cloud"
judge_target = "local"
cloud_judge_target = "cloud"
mode = "shadow"
deadline_seconds = 30
task_typing = true
breaker_threshold = 5
breaker_cooldown_seconds = 30
confirmed_recovery_min_clean_tail = 1
judge_target defaults to local_target. cloud_judge_target is optional.
The local and cloud targets must resolve to distinct model IDs. When the local
backend reports its live context capacity, VGR republishes it through
/v1/models. confirmed_recovery_min_clean_tail, when set, lets an agentic
run that recovered from tool errors commit locally once that many trailing tool
results are clean and the cloud judge confirms the evidence.
Agentic handoff¶
These settings are off by default:
agentic_handoff tells the cloud tier, once per user turn, that it is taking
over unverified tool-using work, including the local tier's unchecked final
message. compact_handoff also condenses the local tier's history into a
digest and reapplies that same digest to every later request of the user turn,
which keeps the cloud tier's prompt cache warm. local_turn_budget_seconds
escalates a user turn once the local tier has spent that much wall-clock time
on it.
Serving modes¶
offskips candidate generation and serves cloud. This is the default.shadowcomputes a decision but serves cloud.evaluateserves the policy decision for isolated evaluation.activeserves the policy decision in production and requiresactive_approval = "prospective-validation-and-canary-approved".
VGR buffers a local candidate before releasing it. A streamed candidate that is licensed is replayed as a stream; a rejected candidate is not sent to the client.
x-model-router-selected-model on each response names the tier that served it.