How tokens are counted & credits are charged — per model

This document explains, for the Rayu‑hosted (paid) models, exactly how token usage is turned into a credit charge, and the per‑model rate for every model currently offered.

TL;DR

  • 1,000 credits = 1,000,000 tokens (i.e. 1 credit = 1,000 tokens).
  • Every model has a rate = credits per 1,000,000 tokens (its creditMultiplier).
  • You are charged: credits = tokens_used × rate ÷ 1,000, where tokens_used is the actual input + output tokens the model reported.
  • A failed or unavailable request costs 0 credits.

For the underlying mechanism (reserve/settle, cache‑aware billing, the Redis counter), see credits-and-limits.md.


1. The charge formula

Charging happens in the gateway, off the model's real usage — not an estimate. The billing formula prices each token bucket independently:

billable_tokens = (fresh_input × input_rate)
               + (cache_hit   × cache_read_rate)
               + (cache_write × cache_write_rate)
               + (output      × output_rate)

credits_charged = billable_tokens ÷ 1,000

Each model has four admin-configurable rates (credits per 1M tokens):

  • creditMultiplier — input (fresh/cache-miss) tokens
  • outputCreditMultiplier — output (completion) tokens
  • cacheReadCreditMultiplier — cache-hit tokens (a ratio of the input rate)
  • cacheWriteCreditMultiplier — cache-creation tokens

For flat-rate models (input = output, no caching), this simplifies to:

credits = (input + output) × rate ÷ 1,000
  • rate = 1.0 → 1,000 tokens costs 1 credit; 1M tokens costs 1,000 credits
  • rate = 2.5 → 1,000 tokens costs 2.5 credits; 1M tokens costs 2,500 credits
  • rate = 0.33 → 1,000 tokens costs 0.33 credits; 1M tokens costs 330 credits

Only tokens the provider actually reports are counted; the display in /usage shows the real tokens used, not a rounded‑up number.


2. Per‑model rates

Rates below are the current defaults. All of them are admin‑editable in the dashboard (Providers → per‑model charges); nothing here is hard‑coded into the charge.

Model (id)Served viaInputOutputCache ReadCache Write
DeepSeek V4 Flash (deepseek-v4-flash)DeepSeek API0.330.3310% of input0.33
DeepSeek V4 Pro (deepseek-v4-pro)DeepSeek API1.01.010% of input1.0
LongCat 2.0 (longcat-2)LongCat0.52.02% of input0.5
GLM‑5.2 (glm-5.2)Ollama Cloud2.52.510% of input2.5
Kimi K2.7 (kimi-k2.7)Ollama Cloud2.52.510% of input2.5
MiniMax M3 (minimax-m3)Ollama Cloud2.52.510% of input2.5
Llama 4 (llama-4)Ollama Cloud1.01.010% of input1.0
GPT‑OSS 120B (gpt-oss-120b)Ollama Cloud0.750.7510% of input0.75
Qwen3.5 397B (qwen3.5-397b)Ollama Cloud0.750.7510% of input0.75
Qwen3.5 122B (qwen3.5-122b)Ollama Cloud0.750.7510% of input0.75

Notes:

  • Rates are credits per 1M tokens. The cacheReadMode for all models is ratio, meaning cache reads cost a percentage of the input rate (not an absolute value).
  • Ollama Cloud does not do prompt caching today, so the cache-read rate is latent for those models — it will apply automatically if/when caching appears.
  • LongCat 2.0 has a higher output rate (2.0 vs 0.5 input) because the provider charges ~4× more for output tokens ($2.95 vs $0.75 per 1M).

Rate tiers at a glance (input rate)

  • 0.33 — DeepSeek V4 Flash (cheapest)
  • 0.5 — LongCat 2.0
  • 0.75 — GPT‑OSS 120B, Qwen3.5 397B, Qwen3.5 122B
  • 1.0 — DeepSeek V4 Pro (reference), Llama 4
  • 2.5 — GLM‑5.2, Kimi K2.7, MiniMax M3

3. Worked examples

At the current baseline (1,000 credits = 1,000,000 tokens, i.e. tokensPerCredit = 1,000):

ModelTokens usedCharge
DeepSeek V4 Pro (1.0)1,000,0001,000 credits
DeepSeek V4 Pro (1.0)100,000 (a typical turn)100 credits
DeepSeek V4 Flash (0.33)100,00033 credits
GPT‑OSS 120B (0.75)100,00075 credits
GLM‑5.2 (2.5)100,000250 credits
GLM‑5.2 (2.5)10,000 (a small turn)25 credits

A trivial message (e.g. a few thousand tokens) costs a handful of credits — it is not rounded up to a thousand credits.


4. What a plan buys

A plan grants a credit allowance per billing period (creditsPerPeriod). Because 1 credit = 1,000 tokens at the reference model, the allowance converts to tokens per model by dividing by the model's rate.

Example — a plan with 50,000 credits / period (Pro):

If you only used…You'd get about…
DeepSeek V4 Pro (1.0)50,000,000 tokens (50,000 ÷ 1.0 × 1,000)
DeepSeek V4 Flash (0.33)~151,000,000 tokens (50,000 ÷ 0.33 × 1,000)
GPT‑OSS 120B (0.75)~66,000,000 tokens (50,000 ÷ 0.75 × 1,000)
GLM‑5.2 (2.5)20,000,000 tokens (50,000 ÷ 2.5 × 1,000)

Credits deplete over the period and reset at renewal. When the balance is exhausted, further hosted requests return a credit‑limit error (or draw from top‑up credits if the plan enables top‑up). Plan credit amounts and prices are admin‑configured — see credits-and-limits.md.


5. Edge cases (what you are NOT charged for)

  • Failed / errored request — if the upstream fails, usage settles to the real amount (0 for a failed turn), so a failed request costs 0 credits.
  • Disabled provider — a request for a model whose provider is turned off returns 503 before any reserve, so it charges 0 credits and does not consume a daily turn.
  • Rate‑limited key — the gateway rotates/fails over across the provider's keys automatically; you are only charged for the request that actually succeeds, at the model's rate.

6. Where the numbers live (admin‑editable)

WhatWhere
Per‑model rates (input, output, cache-read, cache-write)Admin → Providers (per model)
baselineCreditsPer1M (1000 → 1 credit = 1,000 tokens)Admin → Plans & Credits → Global limits
Plan credit allowance (creditsPerPeriod)Admin → Plans
Which providers are activeAdmin → Providers (providers.enabled)
Where/how a provider is called (base URL, wire format, auth)Admin → Providers (providers table)
Per‑model capabilities (thinking, image input)Admin → Providers (per model flags)

The gateway computes every charge from these values at request time — changing a rate in the dashboard changes the charge with no code change or redeploy of the CLI. The CLI is display‑only; the gateway is the single source of truth for billing.