Loading page
Please wait while the page loads.
Please wait while the page loads.
Connect, understand, execute, verify, meter: not isolated features but one engineering pipeline you can observe, control and reuse.
Access
One base URL and one key bring your everyday coding terminals onto the same path without changing how you use them.
Chat Completions and Responses; existing SDKs and scripts switch by changing one base_url.
Terminals such as Claude Code connect directly through ANTHROPIC_BASE_URL.
Pick Codex CLI, Claude Code, Cursor, Continue and more; copy a config with the key already filled in.
The setup guide waits for the key's first real request and shows success, latency, model and protocol.
SSE streaming, tool calls and multi-turn messages behave like the native API.
Context
Explain something once per project; afterwards any terminal can retrieve and reuse it.
Stable facts in five scopes: stack, structure, conventions, pitfalls, stage summary. Each entry can be disabled.
Five kinds of entries (facts, decisions, preferences, pitfalls, structure) with state, hit count and source; confirm, disable or delete.
Candidate memories are extracted after a task and only injected once confirmed.
The “tell the assistant” form writes a rule that takes effect immediately.
⌘K searches projects, tasks, memories and requests, narrowable by project.
Control profile injection and memory enhancement per project and compare the effect any time.
Agents
Split one big change into parallel, traceable executions with usage and results collected in one place.
In-flight concurrency, burst ceiling and a live pulse per key; overruns return a readable reason.
Name a repository, branch, project and allowed actions; an agent engine executes in the cloud.
Turn-by-turn requests and replies next to a work-state panel of progress, decisions, pitfalls and to-dos.
Ordered steps, each with instructions, allowed actions and a verify command; save as a template and rerun.
Set an equivalent-token ceiling per run; it stops at the limit instead of silently burning allowance.
Routing
Route by stage between the fast and the deep model, wasting nothing on small things and settling for nothing on hard ones.
Retrieval, rewriting and summaries go fast; architecture and complex fixes go deep.
High-difficulty turns are escalated automatically, with the reason noted in the request details.
Each key is limited to the models it may use, preventing accidental use of expensive ones.
Models convert at their ratio into one allowance number.
Upstream failures retry by credential weight and health, with every reason recorded.
Caching
System prompts, project profiles and long context injected repeatedly for the same repository are reused, so repeats are not billed at full price.
System prompts, profiles and tool definitions are recognised as a stable prefix and cached with no request changes.
A task's history is cached incrementally by turn; long sessions hit more the further they go.
Every request splits cache-hit and cache-miss tokens so the bill shows what was saved.
Set TTL and minimum cache length per project or key; switch it off for sensitive projects.
Planned: not in the current release; the copy describes the target shape.
When a profile or memory changes the matching cache is invalidated so stale context is never reused.
Planned: not in the current release; the copy describes the target shape.
Hit rate, tokens saved, money saved and TTFT improvement, by day and by key.
Planned: not in the current release; the copy describes the target shape.
Keys
Split keys by purpose; each has its own limits, usage and project.
Create, rotate, enable, disable and revoke; the secret is shown once, at creation.
Model allowlist, IP allowlist, concurrency, RPM and a per-period usage cap.
A key bound to a project carries that project's profile and memories automatically.
Concurrency and burst curves, daily usage bars, recent requests and last use.
Metering
Every call lands on one line of the ledger; allowance and spend can be checked any time.
Input, output, cache-hit and cache-miss tokens plus cost and duration.
Today / 7 days / 30 days by day, model or key, with CSV export.
Five plans, 30-day periods, an allowance bar and days remaining at a glance.
Top-ups, subscription charges and failed-payment refunds, line by line.
A reminder at 80% so a task does not stop halfway.
Observability
No guessing when something breaks: request body, parameters and failure reason are kept together.
Filter by key, model, status and time; infinite scroll; open the detail drawer.
Parameters, usage and cost, peak marker, failure reason code and message body; copy as curl.
Pick a key and a model and chat; tokens, cost, TTFT and duration update live.
In-flight concurrency, requests this second and this minute against their limits, refreshed every 5 s.
TTFT P50 / P95 and total duration, split by model and key.
Governance
Clear account, data and permission boundaries so teams can trust it.
Phone binding and rebinding, WeChat binding, passwords, session list and remote sign-out.
Choose not to store message bodies and keep only metering and status.
Cloud agents only run allowed actions and never exceed the terminal's own permission model.
Concurrency and allowance are checked live; anomalies leave replay and reconciliation records.
Key, plan, top-up and permission changes are logged line by line.
Create a key, copy a terminal config, verify the connection. No plan purchase first, no automatic charge.