How tokens are counted & credits are charged — per model
This document explains, for the Rayu‑hosted (paid) models, exactly how token usage is turned into a credit charge, and the per‑model rate for every model currently offered.
TL;DR
- 1,000 credits = 1,000,000 tokens (i.e. 1 credit = 1,000 tokens).
- Every model has a rate = credits per 1,000,000 tokens (its
creditMultiplier).- You are charged:
credits = tokens_used × rate ÷ 1,000, wheretokens_usedis the actual input + output tokens the model reported.- A failed or unavailable request costs 0 credits.
For the underlying mechanism (reserve/settle, cache‑aware billing, the Redis counter), see
credits-and-limits.md.
1. The charge formula
Charging happens in the gateway, off the model's real usage — not an estimate. The billing formula prices each token bucket independently:
billable_tokens = (fresh_input × input_rate)
+ (cache_hit × cache_read_rate)
+ (cache_write × cache_write_rate)
+ (output × output_rate)
credits_charged = billable_tokens ÷ 1,000
Each model has four admin-configurable rates (credits per 1M tokens):
creditMultiplier— input (fresh/cache-miss) tokensoutputCreditMultiplier— output (completion) tokenscacheReadCreditMultiplier— cache-hit tokens (a ratio of the input rate)cacheWriteCreditMultiplier— cache-creation tokens
For flat-rate models (input = output, no caching), this simplifies to:
credits = (input + output) × rate ÷ 1,000
rate = 1.0→ 1,000 tokens costs 1 credit; 1M tokens costs 1,000 creditsrate = 2.5→ 1,000 tokens costs 2.5 credits; 1M tokens costs 2,500 creditsrate = 0.33→ 1,000 tokens costs 0.33 credits; 1M tokens costs 330 credits
Only tokens the provider actually reports are counted; the display in /usage
shows the real tokens used, not a rounded‑up number.
2. Per‑model rates
Rates below are the current defaults. All of them are admin‑editable in the dashboard (Providers → per‑model charges); nothing here is hard‑coded into the charge.
| Model (id) | Served via | Input | Output | Cache Read | Cache Write |
|---|---|---|---|---|---|
DeepSeek V4 Flash (deepseek-v4-flash) | DeepSeek API | 0.33 | 0.33 | 10% of input | 0.33 |
DeepSeek V4 Pro (deepseek-v4-pro) | DeepSeek API | 1.0 | 1.0 | 10% of input | 1.0 |
LongCat 2.0 (longcat-2) | LongCat | 0.5 | 2.0 | 2% of input | 0.5 |
GLM‑5.2 (glm-5.2) | Ollama Cloud | 2.5 | 2.5 | 10% of input | 2.5 |
Kimi K2.7 (kimi-k2.7) | Ollama Cloud | 2.5 | 2.5 | 10% of input | 2.5 |
MiniMax M3 (minimax-m3) | Ollama Cloud | 2.5 | 2.5 | 10% of input | 2.5 |
Llama 4 (llama-4) | Ollama Cloud | 1.0 | 1.0 | 10% of input | 1.0 |
GPT‑OSS 120B (gpt-oss-120b) | Ollama Cloud | 0.75 | 0.75 | 10% of input | 0.75 |
Qwen3.5 397B (qwen3.5-397b) | Ollama Cloud | 0.75 | 0.75 | 10% of input | 0.75 |
Qwen3.5 122B (qwen3.5-122b) | Ollama Cloud | 0.75 | 0.75 | 10% of input | 0.75 |
Notes:
- Rates are credits per 1M tokens. The
cacheReadModefor all models isratio, meaning cache reads cost a percentage of the input rate (not an absolute value). - Ollama Cloud does not do prompt caching today, so the cache-read rate is latent for those models — it will apply automatically if/when caching appears.
- LongCat 2.0 has a higher output rate (2.0 vs 0.5 input) because the provider charges ~4× more for output tokens ($2.95 vs $0.75 per 1M).
Rate tiers at a glance (input rate)
- 0.33 — DeepSeek V4 Flash (cheapest)
- 0.5 — LongCat 2.0
- 0.75 — GPT‑OSS 120B, Qwen3.5 397B, Qwen3.5 122B
- 1.0 — DeepSeek V4 Pro (reference), Llama 4
- 2.5 — GLM‑5.2, Kimi K2.7, MiniMax M3
3. Worked examples
At the current baseline (1,000 credits = 1,000,000 tokens, i.e.
tokensPerCredit = 1,000):
| Model | Tokens used | Charge |
|---|---|---|
| DeepSeek V4 Pro (1.0) | 1,000,000 | 1,000 credits |
| DeepSeek V4 Pro (1.0) | 100,000 (a typical turn) | 100 credits |
| DeepSeek V4 Flash (0.33) | 100,000 | 33 credits |
| GPT‑OSS 120B (0.75) | 100,000 | 75 credits |
| GLM‑5.2 (2.5) | 100,000 | 250 credits |
| GLM‑5.2 (2.5) | 10,000 (a small turn) | 25 credits |
A trivial message (e.g. a few thousand tokens) costs a handful of credits — it is not rounded up to a thousand credits.
4. What a plan buys
A plan grants a credit allowance per billing period (creditsPerPeriod).
Because 1 credit = 1,000 tokens at the reference model, the allowance converts
to tokens per model by dividing by the model's rate.
Example — a plan with 50,000 credits / period (Pro):
| If you only used… | You'd get about… |
|---|---|
| DeepSeek V4 Pro (1.0) | 50,000,000 tokens (50,000 ÷ 1.0 × 1,000) |
| DeepSeek V4 Flash (0.33) | ~151,000,000 tokens (50,000 ÷ 0.33 × 1,000) |
| GPT‑OSS 120B (0.75) | ~66,000,000 tokens (50,000 ÷ 0.75 × 1,000) |
| GLM‑5.2 (2.5) | 20,000,000 tokens (50,000 ÷ 2.5 × 1,000) |
Credits deplete over the period and reset at renewal. When the balance is
exhausted, further hosted requests return a credit‑limit error (or draw from
top‑up credits if the plan enables top‑up). Plan credit amounts and prices are
admin‑configured — see credits-and-limits.md.
5. Edge cases (what you are NOT charged for)
- Failed / errored request — if the upstream fails, usage settles to the real amount (0 for a failed turn), so a failed request costs 0 credits.
- Disabled provider — a request for a model whose provider is turned off
returns
503before any reserve, so it charges 0 credits and does not consume a daily turn. - Rate‑limited key — the gateway rotates/fails over across the provider's keys automatically; you are only charged for the request that actually succeeds, at the model's rate.
6. Where the numbers live (admin‑editable)
| What | Where |
|---|---|
| Per‑model rates (input, output, cache-read, cache-write) | Admin → Providers (per model) |
baselineCreditsPer1M (1000 → 1 credit = 1,000 tokens) | Admin → Plans & Credits → Global limits |
Plan credit allowance (creditsPerPeriod) | Admin → Plans |
| Which providers are active | Admin → Providers (providers.enabled) |
| Where/how a provider is called (base URL, wire format, auth) | Admin → Providers (providers table) |
| Per‑model capabilities (thinking, image input) | Admin → Providers (per model flags) |
The gateway computes every charge from these values at request time — changing a rate in the dashboard changes the charge with no code change or redeploy of the CLI. The CLI is display‑only; the gateway is the single source of truth for billing.