API Keys
Generate API keys to access Rayu's hosted LLM models programmatically — from your own applications, CI pipelines, agent frameworks, or any tool that speaks the OpenAI or Anthropic API format.
Getting a Key
- Sign in to rayucode.com/dashboard/api-keys
- Click Create API Key and give it a name
- Copy the key immediately — it is shown only once and cannot be recovered
API keys require a Pro plan or higher. Free and Basic plans do not include API access.
Base URL
https://gateway.rayucode.com/v1
All endpoints are served under this base URL. Point your SDK's base_url here and use your Rayu API key as the api_key.
Authentication
OpenAI-compatible (Authorization header)
Authorization: Bearer rayu_sk_live_...
Anthropic-compatible (x-api-key header)
x-api-key: rayu_sk_live_...
Both headers are accepted on all endpoints.
Endpoints
| Method | Path | Format | Description |
|---|---|---|---|
| POST | /v1/chat/completions | OpenAI | Chat completions (streaming + non-streaming) |
| POST | /v1/messages | Anthropic | Anthropic Messages API |
| POST | /v1/messages/count_tokens | Anthropic | Token counting (free, no credits charged) |
| GET | /v1/models | OpenAI list | Available models for your plan |
| GET | /v1/credits | Rayu | Credit balance and usage |
Quick Start — OpenAI Format
Python
from openai import OpenAI client = OpenAI( api_key="rayu_sk_live_...", base_url="https://gateway.rayucode.com/v1" ) response = client.chat.completions.create( model="", messages=[{"role": "user", "content": "Hello!"}], stream=True ) for chunk in response: if chunk.choices[0].delta.content: print(chunk.choices[0].delta.content, end="")
TypeScript
import OpenAI from 'openai'; const client = new OpenAI({ apiKey: 'rayu_sk_live_...', baseURL: 'https://gateway.rayucode.com/v1', }); const stream = await client.chat.completions.create({ model: '', messages: [{ role: 'user', content: 'Hello!' }], stream: true, }); for await (const chunk of stream) { process.stdout.write(chunk.choices[0]?.delta?.content || ''); }
curl
curl https://gateway.rayucode.com/v1/chat/completions \ -H "Authorization: Bearer rayu_sk_live_..." \ -H "Content-Type: application/json" \ -d '{ "model": "deepseek-v3", "messages": [{"role": "user", "content": "Hello!"}], "stream": true }'
Quick Start — Anthropic Format
Python
import anthropic client = anthropic.Anthropic( api_key="rayu_sk_live_...", base_url="https://gateway.rayucode.com/v1" ) message = client.messages.create( model="deepseek-v3", max_tokens=1024, messages=[{"role": "user", "content": "Hello!"}] ) print(message.content[0].text)
Streaming (Anthropic)
with client.messages.stream( model="deepseek-v3", max_tokens=1024, messages=[{"role": "user", "content": "Explain quantum computing"}] ) as stream: for text in stream.text_stream: print(text, end="")
Streaming
Both formats support SSE streaming. Set "stream": true in the request body.
- OpenAI format: Emits
data: {"id":...,"choices":[{"delta":...}]}chunks followed bydata: [DONE] - Anthropic format: Emits standard Anthropic SSE events (
message_start,content_block_delta, etc.)
Streams are unbuffered — tokens arrive as soon as the model generates them. Long-running streams (several minutes for large outputs) are fully supported.
Tools & Function Calling
OpenAI Format
response = client.chat.completions.create( model="deepseek-v3", messages=[{"role": "user", "content": "What's the weather in London?"}], tools=[{ "type": "function", "function": { "name": "get_weather", "description": "Get current weather", "parameters": { "type": "object", "properties": {"city": {"type": "string"}}, "required": ["city"] } } }], tool_choice="auto" )
Tool calls are returned in choices[0].message.tool_calls and streamed as argument deltas.
Anthropic Format
Tools work exactly as documented in the Anthropic API — tools array with input_schema, tool_use content blocks in responses, tool_result in follow-up messages.
Vision (Image Input)
Models with image support accept images via base64 data URIs:
response = client.chat.completions.create( model="claude-sonnet-4", # must support images messages=[{ "role": "user", "content": [ {"type": "text", "text": "What's in this image?"}, {"type": "image_url", "image_url": { "url": "data:image/png;base64,iVBOR..." }} ] }] )
Remote URLs (https://...) are not supported — use base64 data URIs only. This avoids SSRF risks and provider-dependent behavior.
Credit Headers
Every response includes credit usage information in headers:
| Header | Description |
|---|---|
x-rayu-credits-used | Billable tokens consumed this period |
x-rayu-credits-remaining | Tokens remaining in the period allowance |
x-rayu-topup-balance | Top-up credit balance |
x-rayu-limit | Total period allowance |
Per-Key Controls
Each API key can have optional limits set in the dashboard:
- Credit cap: Maximum credits this key can spend per billing period
- Model allowlist: Restrict which models this key can access
- Rate limit (RPM): Maximum requests per minute
- Expiry date: Key automatically stops working after this date
Rate Limits & Errors
| Status | Reason | Retry? |
|---|---|---|
| 401 | Invalid, revoked, or expired API key | No — check or regenerate your key |
| 403 | Model not available on your plan, or API access not enabled | No — upgrade plan or check model name |
| 429 | Per-key RPM limit, per-key credit cap, or plan period limit reached | Yes — after Retry-After seconds |
| 502 | Upstream provider temporarily unavailable | Yes — retry with exponential backoff |
| 503 | Gateway at capacity | Yes — after Retry-After seconds |
Error responses follow the format of the endpoint you called:
- OpenAI endpoints:
{"error": {"message": "...", "type": "...", "code": null}} - Anthropic endpoints:
{"type": "error", "error": {"type": "...", "message": "..."}}
Supported Parameters
OpenAI /v1/chat/completions
| Parameter | Supported | Notes |
|---|---|---|
model | Yes | Rayu model code |
messages | Yes | system, user, assistant, tool roles |
max_tokens / max_completion_tokens | Yes | Default 4096 if omitted |
temperature | Yes | |
top_p | Yes | |
stop | Yes | String or array |
stream | Yes | |
stream_options.include_usage | Yes | Final usage chunk |
tools | Yes | Function calling |
tool_choice | Yes | auto, none, required, specific function |
n | No | Only n=1 supported |
logprobs / top_logprobs | No | Rejected with 400 |
seed | No | Rejected with 400 |
logit_bias | No | Rejected with 400 |
frequency_penalty / presence_penalty | Ignored | |
response_format | Not yet |
Anthropic /v1/messages
The full Anthropic Messages API is supported, including:
system,messages,max_tokens,temperature,top_p,stop_sequencestools,tool_choice,tool_use/tool_resultcontent blocksstreamwith all event types- Image content blocks (base64)
- Extended thinking (
thinkingcontent blocks) - Prompt caching (
cache_controlblocks) — charged at the model's cache read/write multipliers
Security Best Practices
- Never expose API keys in client-side code (browsers, mobile apps). Keys are for server-to-server use.
- Use per-key credit caps to limit blast radius if a key is leaked.
- Set an expiry for keys used in temporary environments (CI, demos).
- Use the model allowlist to prevent a leaked key from accessing expensive models.
- Revoke immediately if you suspect a key has been compromised — revocation takes effect within seconds.
- Rotate keys periodically as a hygiene practice.
Model List
To see which models are available on your plan:
curl https://gateway.rayucode.com/v1/models \ -H "Authorization: Bearer rayu_sk_live_..."
Returns an OpenAI-compatible model list with capabilities (supportsReasoning, supportsImage, supportsTools, contextWindow).
How Model Fetch Works
This section explains how Rayu discovers, filters, and refreshes the model catalog — both for API key callers (this document) and for Rayu OAuth (Auth) callers. The two paths authenticate differently but enforce the same filtering rules and ordering.
The two Rayu providers
| Rayu Auth (OAuth) | Rayu API Key | |
|---|---|---|
| Provider id | rayu-hosted | rayu |
| Kind | rayu-hosted | anthropic-compatible |
| Credential | Rayu account JWT (from /login) | rayu_sk_live_… API key |
| Chat endpoint | {gateway}/anthropic/v1/messages (JWT-injecting fetch) | {gateway}/anthropic/v1/messages (key as Bearer) |
| Model catalog endpoint | GET {backend}/me/entitlements | GET {gateway}/v1/models |
| Registered by | Auto on /login | /connect wizard or RAYU_API_KEY env |
| Display name | Rayu | Rayu API Key |
They are deliberately separate providers so a user can have both configured
without either clobbering the other's credential — the same reasoning as
anthropic vs claude-subscription.
Path 1 — Rayu Auth (OAuth / JWT)
Authentication flow
User → Google OAuth (rayu-web) → rayu-backend /api/auth/oauth/google
→ issues Rayu JWT (signed with RAYU_JWT_SECRET)
→ CLI stores JWT in ~/.rayu/rayu-auth.json
Model fetch flow
CLI → GET {backend}/me/entitlements (Authorization: Bearer <JWT>)
→ backend resolves user's active plan
→ backend queries hosted_models WHERE enabled=true AND provider.enabled=true
→ returns TWO lists:
• allowedModels — plan-filtered subset (drives entitlement/gating)
• hostedModels — full enabled catalog (shown to ALL signed-in users)
Backend query (NestJS + Prisma)
// All enabled models (the catalog shown to every signed-in user) findEnabled(): Promise<HostedModelWithProvider[]> { return this.prisma.hostedModel.findMany({ where: { enabled: true, provider: { enabled: true } }, orderBy: [{ sortOrder: 'asc' }, { id: 'asc' }], include: WITH_PROVIDER, }) } // Plan-allowed subset (drives entitlement) async findAllowedForPlan(planCode: string) { const all = await this.findEnabled() return all.filter((m) => this.allowedCodes(m).includes(planCode)) }
Entitlements response shape
{ "plan": { "code": "pro", "name": "pro", "priceCents": 2900 }, "allowedModels": [ { "code": "deepseek-v4-pro", "label": "DeepSeek V4 Pro", "contextWindow": 131072, "supportsReasoning": true, "supportsImage": false, "supportsTools": true, "creditMultiplier": 1.0 } ], "hostedModels": [ { "code": "deepseek-v4-pro", "label": "DeepSeek V4 Pro", ... }, { "code": "deepseek-v4-flash", "label": "DeepSeek V4 Flash", ... } ] }
How the CLI uses it
// Visibility uses the full catalog; usability uses the entitled subset. const catalog = ent?.hostedModels ?? ent?.allowedModels ?? [] const entitled = ent?.allowedModels ?? [] const models = catalog.map((m) => m.code)
hostedModelspresent → shows ALL enabled models (Free users see them but are gated on use; a model is usable iff it also appears inallowedModels).hostedModelsabsent (older backend) → falls back toallowedModels(plan-filtered only).
Model ordering
Models are returned in ORDER BY sortOrder ASC, id ASC. The admin dashboard's
reorder UI sets sortOrder = index × 10. The CLI preserves this order exactly —
no client-side sorting.
Auto-refresh
- On login:
syncRayuHostedProvider()is called with the fresh entitlements. - Background:
getCachedEntitlements()kicks a rate-limited (30s cooldown) background refresh on every read. - On
/modelopen: the model picker callsrefreshHostedCatalog()which re-fetches entitlements and re-renders only if the catalog changed.
Path 2 — Rayu API Key (this path)
Authentication flow
User → rayucode.com/dashboard/api-keys → creates rayu_sk_live_… key
→ pastes key into /connect → CLI sends it as Bearer to gateway
Model fetch flow
CLI → GET {gateway}/v1/models (Authorization: Bearer <key>)
→ gateway resolves user's identity from the key
→ gateway resolves user's active plan
→ gateway filters: enabled model + enabled provider + plan-allowed + key-allowlist
→ returns OpenAI list shape with capabilities
Gateway query (Rust + SQLx)
-- Loads ALL hosted_models rows (no WHERE — filtering happens in Rust) SELECT m.*, p.* FROM hosted_models m JOIN providers p ON p.id = m.provider_id ORDER BY m.sortOrder, m.id
Gateway filtering (Rust)
pub fn allowed_models(models: &[HostedModel], plan_code: &str) -> Vec<HostedModel> { models .iter() .filter(|m| { m.enabled && m.provider.enabled // matches backend's findEnabled() && m.allowed_plan_codes.iter().any(|pc| pc == plan_code) }) .cloned() .collect() }
Then the API key's own allowlist is applied on top:
fn visible_chat_models(ent: &Entitlement, api_key: Option<&ApiKeyContext>) -> Vec<&HostedModel> { ent.allowed_models .iter() .filter(|m| api_key.is_none_or(|ak| ak.allows_model(&m.code))) .collect() }
/v1/models response shape
{ "object": "list", "data": [ { "id": "deepseek-v4-pro", "object": "model", "created": 1700000000, "owned_by": "rayu", "label": "DeepSeek V4 Pro", "supportsReasoning": true, "supportsImage": false, "supportsTools": true, "contextWindow": 131072 } ] }
How the CLI uses it
export function parseRayuCatalog(payload: unknown) { // Sanitize every id, dedupe, preserve gateway order (no .sort()) return { models: entries.map(e => e.code), modelLabels: hostedModelLabels(entries), modelContextWindows: hostedContextWindows(entries), } }
Model ordering
The gateway returns models in ORDER BY m.sortOrder, m.id. The CLI preserves
this order — no client-side sorting. This matches the Auth path exactly.
Auto-refresh
- On connect: the
/connectwizard fetches the catalog and persists it. - On
/modelopen: the model picker callsrefreshRayuApiKeyCatalog()which re-fetches from the gateway and re-renders only if the catalog changed. - Background:
refreshActiveProviderModels()callsrefreshRayuApiKeyCatalog()when the provider is active.
Filtering rules (both paths)
A model appears in the catalog only if ALL of these are true:
| Rule | Auth path | API key path |
|---|---|---|
Model enabled = true | Backend findEnabled() | Gateway allowed_models() |
Provider enabled = true | Backend findEnabled() | Gateway m.provider.enabled |
Model's allowedPlanCodes includes user's plan | Backend findAllowedForPlan() | Gateway allowed_models() |
API key's allowed_models includes the model | N/A (no key) | Gateway visible_chat_models() |
An empty allowedPlanCodes means NOBODY for chat models — a model must
be explicitly granted to a plan. (The opposite rule applies to media models,
where empty means EVERY plan.)
Per-key controls
An API key can further narrow the catalog via the dashboard:
| Control | Effect on /v1/models |
|---|---|
| Model allowlist | Only listed models appear (intersected with plan) |
| Empty allowlist | No restriction — full plan catalog |
| Credit cap | Doesn't affect listing; enforced on request path |
| Rate limit (RPM) | Doesn't affect listing; enforced on request path |
A key allowlist is a narrowing, never a grant: it cannot add a model the plan doesn't include.
Stale-default pruning
Both paths prune the user's chosen default/small model if the admin removes it from the catalog. Holding on to a removed code would send every request to a model the gateway now rejects (403 "model not available"), which reads like a CLI bug rather than a catalog change.
// Auth path const inCatalog = (code?: string): boolean => !!code && models.includes(code) defaultModel: inCatalog(existing?.defaultModel) ? existing?.defaultModel : preferredCode // API key path if (!cur.defaultModel || !result.models.includes(cur.defaultModel)) { cur.defaultModel = fallback.defaultModel }
Config persistence
Both providers store their catalog in ~/.rayu/providers.json:
{ "id": "rayu", "kind": "anthropic-compatible", "baseURL": "https://gateway.rayucode.com/anthropic", "apiKey": "rayu_sk_live_...", "models": ["deepseek-v4-pro", "deepseek-v4-flash"], "fetchedModels": ["deepseek-v4-pro", "deepseek-v4-flash"], "modelLabels": { "deepseek-v4-pro": "DeepSeek V4 Pro" }, "modelContextWindows": { "deepseek-v4-pro": 131072 }, "defaultModel": "deepseek-v4-pro", "smallFastModel": "deepseek-v4-flash" }
models— the catalog in display order (what/modelshows).fetchedModels— same list, tracked separately for refresh detection.modelLabels— admin display names, keyed by model id.modelContextWindows— admin context windows in tokens, keyed by model id.
Error handling
| Failure | Auth path | API key path |
|---|---|---|
| Network error | Keep cached catalog, log diagnostic | Keep cached catalog, log diagnostic |
| 401 (bad JWT/key) | Remove provider, prompt re-login | Remove provider, prompt re-enter key |
| 403 (account inactive) | Keep cache, surface error | Keep cache, surface error |
| 503 (gateway DB down) | Keep cache, don't blame user | Keep cache, don't blame user |
| Empty catalog | Remove provider (no models to show) | Keep cache, let /model refresh later |
Both paths never throw and never empty the cache on a transient failure — an offline launch or a gateway blip must not wipe the model list.
Quick reference — API endpoints
Backend (rayu-backend, NestJS)
| Method | Path | Auth | Description |
|---|---|---|---|
| GET | /api/me/entitlements | Rayu JWT | Plan, features, allowedModels, hostedModels |
Gateway (rayu-gateway-rust)
| Method | Path | Auth | Description |
|---|---|---|---|
| POST | /anthropic/v1/messages | JWT or API key | Chat (Anthropic Messages format) |
| POST | /v1/chat/completions | API key | Chat (OpenAI format) |
| GET | /v1/models | API key | Available chat models for the caller's plan |
| GET | /v1/models?media=image | API key | Available image-generation models |
| GET | /v1/models?media=video | API key | Available video-generation models |
| GET | /v1/credits | API key | Credit balance and usage |