Every AI model. One platform.
One gateway, one balance, 143 AI models. Lower prices than buying direct from the official providers.
ModelOpenAI · gpt 5.6 sol−50%
Prepaid · no subscription · multi-endpoint
POST /v1/chat/completions
"model": "claude-fable-5"
200 OK · 412 ms
143 models · 1 endpoint · 1 balance
New key created for Marketing
This month's bill 1 invoice
One key per team · one bill
Quota for tk-sup-••••2b7 is used up · request refused
Quota, limits, stop — per key
One endpoint · 143 AI models · pay for what you use
The problem
AI costs climb quietly, and nobody can point to the cause.
Your teams use AI every day. The bills arrive from four directions at once.
Scattered keys
API keys sit in chats, repos and personal laptops.
Leak riskOne card per vendor
Every vendor wants a credit card and sends its own dollar invoice.
Split billingCosts with no owner
Bills arrive with no sign of which team ran them up.
Untracked spendNo stop button
One runaway agent can drain the balance before anyone notices.
Balance leaksThe solution
Every model behind one gateway.
Route for availability, price or TTFT. One base URL, a quota per key.
POST /v1/chat/completions
5-Layer AI Platform
The Zework AI platform architecture
Toko Token AI covers layers 1–2: the gateway, quotas and a model group per key. The layers above are platform expansion stages.
- Ready to use on the desktop
- Works with local files and the web
- Runs on Windows, macOS and Linux
- Share skills across the team
- Manage agents from the cloud
- Ready for the whole team
- Keeps long-term memory
- Supports private deployment
- Ready-made AI agents for each industry
- A marketplace for internal skills
- Automates work across files and the web
- Watch operations live
- Share quota from headquarters down to each employee
- Every level has its own quota and reports
- Each person gets their own API key and model group
- When a quota runs out, that key stops and the others keep working
- One gateway for 143 AI models
- Set a model group for every key
- Bring your own API key (BYOK)
- When the main model is down, the gateway reroutes the request
- Track usage and cost in the dashboard
- Logs record time, key, model, tokens and cost
Layers 1–2 provide model access, quotas and key control. The result: lower costs and usage you can control.
One platform for all your work.
For everyone, not only technical teams. Ideas and templates are ready, so you never start a prompt from scratch.
The latest release is on GitHub Releases. Pick the installer for your operating system.
You start from an empty box, find the idea and write the prompt yourself.
Pick from many templates. The idea is already there, so prompting gets easier.
Integration
Change two lines of code.
You're connected.
Keep the OpenAI SDK. Just change the base URL and API key.
const client = new OpenAI({baseURL: "https://api.openai.com/v1",baseURL: "https://api.tokotokenai.com/v1",apiKey: process.env.OPENAI_API_KEY,apiKey: process.env.TOKO_TOKEN_API_KEY,});Streaming, tool calling and images work on the models that support them.
The CLI sets the API key, base URL and default model in your AI tools.
curl https://api.tokotokenai.com/v1/chat/completions \
-H "Authorization: Bearer $TOKO_TOKEN_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-fable-5",
"stream": true,
"messages": [{"role": "user",
"content": "What is an AI gateway?"}]
}'
An AI gateway is one door to many AI models: one base URL, one balance and a quota per key. Your code keeps the OpenAI format.
- Install
@anthropic-ai/claude-code - Set
ANTHROPIC_BASE_URLandANTHROPIC_AUTH_TOKEN - Run
claude, then switch models with/model
- Install
@openai/codex - Add an OpenAI-compatible provider to
config.toml - Run
codex, then switch models with/model
- Settings → Models
- Turn on OpenAI API Key and
Override Base URL - Add models with Add Custom Model, then pick one in the chat panel
- Set
urlto the full endpoint/v1/chat/completions - Set
apiKeythrough an environment variable - Add models to
availableModels, then restart
- Run
openclaw onboard - Add a provider whose
baseUrlends in/v1 - Point
model.primaryat the model you want
- Install with the official script, then run
hermes model - Choose
Custom Endpoint, then enter the base URL and API key - Run
hermes, then switch models with/model
Savings
Lower prices,
set by model group.
Every model sits in one price group. The group sets the discount, and Creator has no discount.
An example spend split. Discounts follow the price group.
Monthly savings against the providers' official rates.
Example monthly savings against official provider rates. Creator (Seedance) has no discount.
See every model's final rate in Model SquareControl
Give each person one key.
Set the quota and model group.
Match the model group to each user's needs. Set the quota before you hand out the key.
Catalog
Reach 143 AI models
through one endpoint.
Use the same base URL for text, images, video, audio and embeddings.
chat, agent, tool calling
Enterprise group · pay as you go
Seedance 2.5 and 2.0 · Creator, no discount
TTS, STT, embedding
Security
Your prompts are not stored.
Your keys have layered protection.
The gateway works as a pass-through. Conversation content is never kept as stored data.
Request content lives in memory only for the duration of the call, then is released once forwarded. Logs record only the model, tokens and latency.
pass-throughlog = meter onlyAll traffic is encrypted, from your app all the way to the model provider.
TLSExpiry, spend caps, IP and model allowlists, and instant revocation. One leaked key can't reach any other.
IP allowlistinstant revokeMonitoring never stops. Unusual patterns are blocked and quarantined automatically, without waiting for a manual report.
auto-blockquarantineTOTP 2FA, passkeys and step-up verification for sensitive actions. Runs on infrastructure from an ISO 27001-certified operator.
2FA · PasskeyISO 27001 · platform operator