Save up to 50% compared with buying direct from official providers

Every AI model. One platform.

One gateway, one balance, 143 AI models. Lower prices than buying direct from the official providers.

ModelOpenAI · gpt 5.6 sol−50%

Prepaid · no subscription · multi-endpoint

143 AI modelsOne platform for all of them
Multi-endpointChat · Anthropic · Gemini · Video
Automatic failoverA model goes down, requests keep going
Pay as you goNo subscription

One endpoint · 143 AI models · pay for what you use

OpenAI Anthropic Google DeepSeek Alibaba xAI Moonshot ByteDance Zhipu MiniMax
gpt 5.6 sol claude fable 5 gemini 3.1 pro kimi k3 glm 5.2 seedance 2.5 gpt 5.6 luna claude opus 5 gemini 3.5 flash seedance 2.0 fast

The problem

AI costs climb quietly, and nobody can point to the cause.

Your teams use AI every day. The bills arrive from four directions at once.

Scattered keys

API keys sit in chats, repos and personal laptops.

Leak risk

One card per vendor

Every vendor wants a credit card and sends its own dollar invoice.

Split billing

Costs with no owner

Bills arrive with no sign of which team ran them up.

Untracked spend

No stop button

One runaway agent can drain the balance before anyone notices.

Balance leaks

The solution

Every model behind one gateway.

Route for availability, price or TTFT. One base URL, a quota per key.

The main model is down, so requests are rerouted. Your code doesn't change.

Your app

POST /v1/chat/completions

Provider down?
enterprise:anthropicclaude opus 5 Down
primary
enterprise:moonshotkimi k3
fallback
api.tokotokenai.com/v1 balance logged log 10:24:18 200 OK

5-Layer AI Platform

The Zework AI platform architecture

Toko Token AI covers layers 1–2: the gateway, quotas and a model group per key. The layers above are platform expansion stages.

Unified & secure
Lean & controlled
Ready & flexible
Growing together
Dashboard · see activity, usage and cost Reports · automatic operations and ROI reports Alerts · raised when unusual activity is detected Handover · a structured handoff to the operations team
AI training · by role, with certification and monthly reviews Custom agents · built around each team's needs Internal skills · create and share skills inside the company Guidance · from training to deployment
AI CoachTop Sales Recruitment AssistantIT Support
Zework AI Workforce · Desktop
  • Ready to use on the desktop
  • Works with local files and the web
  • Runs on Windows, macOS and Linux
  • Share skills across the team
Zework AI Workforce · Cloud
  • Manage agents from the cloud
  • Ready for the whole team
  • Keeps long-term memory
  • Supports private deployment
  • Ready-made AI agents for each industry
  • A marketplace for internal skills
  • Automates work across files and the web
  • Watch operations live
Four-Tier Governance Model Tiered Usage-Based Billing Token Quality Governance ROI Measurement Security & Compliance Localization
  • Share quota from headquarters down to each employee
  • Every level has its own quota and reports
  • Each person gets their own API key and model group
  • When a quota runs out, that key stops and the others keep working
OpenAI Anthropic Google DeepSeek ByteDance Alibaba Moonshot Zhipu MiniMax xAI …
  • One gateway for 143 AI models
  • Set a model group for every key
  • Bring your own API key (BYOK)
  • When the main model is down, the gateway reroutes the request
  • Track usage and cost in the dashboard
  • Logs record time, key, model, tokens and cost

Layers 1–2 provide model access, quotas and key control. The result: lower costs and usage you can control.

Nultron Beta Desktop app

One platform for all your work.

For everyone, not only technical teams. Ideas and templates are ready, so you never start a prompt from scratch.

Download Nultron Windows · macOS preview

The latest release is on GitHub Releases. Pick the installer for your operating system.

A typical AI chat

You start from an empty box, find the idea and write the prompt yourself.

Nultron

Pick from many templates. The idea is already there, so prompting gets easier.

Integration

Change two lines of code.
You're connected.

Keep the OpenAI SDK. Just change the base URL and API key.

client.js
const client = new OpenAI({
−baseURL: "https://api.openai.com/v1",
+baseURL: "https://api.tokotokenai.com/v1",
−apiKey: process.env.OPENAI_API_KEY,
+apiKey: process.env.TOKO_TOKEN_API_KEY,
});

Streaming, tool calling and images work on the models that support them.

Or use one CLI command
macOS / Linux
curl -fsSL https://api.tokotokenai.com/cli/install.sh | sh
Windows PowerShell
irm https://api.tokotokenai.com/cli/install.ps1 | iex

The CLI sets the API key, base URL and default model in your AI tools.

request · streamingexample · illustration
curl https://api.tokotokenai.com/v1/chat/completions \
  -H "Authorization: Bearer $TOKO_TOKEN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-fable-5",
    "stream": true,
    "messages": [{"role": "user",
      "content": "What is an AI gateway?"}]
  }'
response200 OK

text/event-streamclaude-fable-5

assistant

An AI gateway is one door to many AI models: one base URL, one balance and a quota per key. Your code keeps the OpenAI format.

Claude Code
  1. Install @anthropic-ai/claude-code
  2. Set ANTHROPIC_BASE_URL and ANTHROPIC_AUTH_TOKEN
  3. Run claude, then switch models with /model
Codex CLI
  1. Install @openai/codex
  2. Add an OpenAI-compatible provider to config.toml
  3. Run codex, then switch models with /model
Cursor
  1. Settings → Models
  2. Turn on OpenAI API Key and Override Base URL
  3. Add models with Add Custom Model, then pick one in the chat panel
Needs Cursor Pro or a higher plan. Open docs
CodeBuddy
  1. Set url to the full endpoint /v1/chat/completions
  2. Set apiKey through an environment variable
  3. Add models to availableModels, then restart
OpenClaw
  1. Run openclaw onboard
  2. Add a provider whose baseUrl ends in /v1
  3. Point model.primary at the model you want
Hermes Agents
  1. Install with the official script, then run hermes model
  2. Choose Custom Endpoint, then enter the base URL and API key
  3. Run hermes, then switch models with /model

Savings

Lower prices,
set by model group.

Every model sits in one price group. The group sets the discount, and Creator has no discount.

Discount by price group5 groups · 143 models
Valuegpt 5.6 luna
−50%save $500
Standardclaude opus 5, gemini
−40%save $400
Enterpriseclaude fable 5, gpt 5.6 sol
−10%save $100
ChinaLLMkimi k3, glm 5.2
−15%save $150
Creatorseedance 2.5, 2.0
no discount
Old billsUSD
Vendor A · card 1$400
Vendor B · card 2$300
Vendor C · card 3$200
Vendor D · card 4$100
4 bills$1,000
Via gatewayUSD
Value −50%$200
Standard −40%$180
Enterprise −10%$180
ChinaLLM −15%$85
1 bill$645

An example spend split. Discounts follow the price group.

Worked example

Monthly savings against the providers' official rates.

official ratevia gateway
Official $500 · Valuesave $250
Official $625 · Standardsave $250
Official $1,000 · Enterprisesave $100
Official $2,000 · ChinaLLMsave $300
Official $500 · Creatorno discount

Example monthly savings against official provider rates. Creator (Seedance) has no discount.

See every model's final rate in Model Square

Control

Give each person one key.
Set the quota and model group.

Match the model group to each user's needs. Set the quota before you hand out the key.

console · Toko Token AIexample · illustration
Spend this month$512.50
Requests124,000
Active models18
sk-••••••••••••••7f2K$112.50 / $125
sk-••••••••••••••m9Qx$40 / $62.50
sk-••••••••••••••3a8V$30 / $31.25
sk-••••••••••••••Zc51$13.13 / $62.50
Run one request. Budi's quota runs out while the other keys stay active.
10:24:16sk-••••••••••••••7f2K · claude opus 5200 OK
10:24:17sk-••••••••••••••Zc51 · kimi k3200 OK
10:24:18sk-••••••••••••••3a8V · gemini 3.5 flash200 OK
10:24:19sk-••••••••••••••m9Qx · gpt 5.6 sol200 OK
EmployeeMain modelTokens usedToolsKey status
AR
Andi Rahmansk-••••••••••••••7f2K
claude opus 5
1.24M
code · chat
Active
SP
Siti Putrisk-••••••••••••••m9Qx
gpt 5.6 sol
820k
writer · vision
Active
BW
Budi Wibowosk-••••••••••••••3a8V
gemini 3.5 flash
2.10M
agent · tools
Active
DN
Dewi Nursk-••••••••••••••Zc51
kimi k3
410k
RAG · file
Active

Catalog

Reach 143 AI models
through one endpoint.

Use the same base URL for text, images, video, audio and embeddings.

Textsave up to 50%
74

chat, agent, tool calling

Images−10%
31

Enterprise group · pay as you go

Videono discount
21

Seedance 2.5 and 2.0 · Creator, no discount

Audio & embeddings−15%
17

TTS, STT, embedding

https://api.tokotokenai.com/v1

TextImagesVideoAudio = 143
143
models available74 + 31 + 21 + 17
Open Model Square

Security

Your prompts are not stored.
Your keys have layered protection.

The gateway works as a pass-through. Conversation content is never kept as stored data.

Prompts are not stored

Request content lives in memory only for the duration of the call, then is released once forwarded. Logs record only the model, tokens and latency.

pass-throughlog = meter only
Incoming request
"Summarize this contract…"
model: claude-opus-5
GWpass-through
Forwarded to the provider
"Summarize this contract…"
TLS · not stored
What the log keeps "Summarize this contract…" claude-opus-5 · 1,284 tok · 412 ms
End-to-end TLS encryption

All traffic is encrypted, from your app all the way to the model provider.

TLS
A leak stays inside one key

Expiry, spend caps, IP and model allowlists, and instant revocation. One leaked key can't reach any other.

IP allowlistinstant revoke
Unusual usage is blocked

Monitoring never stops. Unusual patterns are blocked and quarantined automatically, without waiting for a manual report.

auto-blockquarantine
Layered account security

TOTP 2FA, passkeys and step-up verification for sensitive actions. Runs on infrastructure from an ISO 27001-certified operator.

2FA · PasskeyISO 27001 · platform operator