API Keys

Generate API keys to access Rayu's hosted LLM models programmatically — from your own applications, CI pipelines, agent frameworks, or any tool that speaks the OpenAI or Anthropic API format.

Getting a Key

  1. Sign in to rayucode.com/dashboard/api-keys
  2. Click Create API Key and give it a name
  3. Copy the key immediately — it is shown only once and cannot be recovered

API keys require a Pro plan or higher. Free and Basic plans do not include API access.

Base URL

https://gateway.rayucode.com/v1

All endpoints are served under this base URL. Point your SDK's base_url here and use your Rayu API key as the api_key.

Authentication

OpenAI-compatible (Authorization header)

Authorization: Bearer rayu_sk_live_...

Anthropic-compatible (x-api-key header)

x-api-key: rayu_sk_live_...

Both headers are accepted on all endpoints.

Endpoints

MethodPathFormatDescription
POST/v1/chat/completionsOpenAIChat completions (streaming + non-streaming)
POST/v1/messagesAnthropicAnthropic Messages API
POST/v1/messages/count_tokensAnthropicToken counting (free, no credits charged)
GET/v1/modelsOpenAI listAvailable models for your plan
GET/v1/creditsRayuCredit balance and usage

Quick Start — OpenAI Format

Python

from openai import OpenAI

client = OpenAI(
    api_key="rayu_sk_live_...",
    base_url="https://gateway.rayucode.com/v1"
)

response = client.chat.completions.create(
    model="",
    messages=[{"role": "user", "content": "Hello!"}],
    stream=True
)

for chunk in response:
    if chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="")

TypeScript

import OpenAI from 'openai';

const client = new OpenAI({
  apiKey: 'rayu_sk_live_...',
  baseURL: 'https://gateway.rayucode.com/v1',
});

const stream = await client.chat.completions.create({
  model: '',
  messages: [{ role: 'user', content: 'Hello!' }],
  stream: true,
});

for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content || '');
}

curl

curl https://gateway.rayucode.com/v1/chat/completions \
  -H "Authorization: Bearer rayu_sk_live_..." \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-v3",
    "messages": [{"role": "user", "content": "Hello!"}],
    "stream": true
  }'

Quick Start — Anthropic Format

Python

import anthropic

client = anthropic.Anthropic(
    api_key="rayu_sk_live_...",
    base_url="https://gateway.rayucode.com/v1"
)

message = client.messages.create(
    model="deepseek-v3",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Hello!"}]
)
print(message.content[0].text)

Streaming (Anthropic)

with client.messages.stream(
    model="deepseek-v3",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Explain quantum computing"}]
) as stream:
    for text in stream.text_stream:
        print(text, end="")

Streaming

Both formats support SSE streaming. Set "stream": true in the request body.

  • OpenAI format: Emits data: {"id":...,"choices":[{"delta":...}]} chunks followed by data: [DONE]
  • Anthropic format: Emits standard Anthropic SSE events (message_start, content_block_delta, etc.)

Streams are unbuffered — tokens arrive as soon as the model generates them. Long-running streams (several minutes for large outputs) are fully supported.

Tools & Function Calling

OpenAI Format

response = client.chat.completions.create(
    model="deepseek-v3",
    messages=[{"role": "user", "content": "What's the weather in London?"}],
    tools=[{
        "type": "function",
        "function": {
            "name": "get_weather",
            "description": "Get current weather",
            "parameters": {
                "type": "object",
                "properties": {"city": {"type": "string"}},
                "required": ["city"]
            }
        }
    }],
    tool_choice="auto"
)

Tool calls are returned in choices[0].message.tool_calls and streamed as argument deltas.

Anthropic Format

Tools work exactly as documented in the Anthropic API — tools array with input_schema, tool_use content blocks in responses, tool_result in follow-up messages.

Vision (Image Input)

Models with image support accept images via base64 data URIs:

response = client.chat.completions.create(
    model="claude-sonnet-4",  # must support images
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "What's in this image?"},
            {"type": "image_url", "image_url": {
                "url": "data:image/png;base64,iVBOR..."
            }}
        ]
    }]
)

Remote URLs (https://...) are not supported — use base64 data URIs only. This avoids SSRF risks and provider-dependent behavior.

Credit Headers

Every response includes credit usage information in headers:

HeaderDescription
x-rayu-credits-usedBillable tokens consumed this period
x-rayu-credits-remainingTokens remaining in the period allowance
x-rayu-topup-balanceTop-up credit balance
x-rayu-limitTotal period allowance

Per-Key Controls

Each API key can have optional limits set in the dashboard:

  • Credit cap: Maximum credits this key can spend per billing period
  • Model allowlist: Restrict which models this key can access
  • Rate limit (RPM): Maximum requests per minute
  • Expiry date: Key automatically stops working after this date

Rate Limits & Errors

StatusReasonRetry?
401Invalid, revoked, or expired API keyNo — check or regenerate your key
403Model not available on your plan, or API access not enabledNo — upgrade plan or check model name
429Per-key RPM limit, per-key credit cap, or plan period limit reachedYes — after Retry-After seconds
502Upstream provider temporarily unavailableYes — retry with exponential backoff
503Gateway at capacityYes — after Retry-After seconds

Error responses follow the format of the endpoint you called:

  • OpenAI endpoints: {"error": {"message": "...", "type": "...", "code": null}}
  • Anthropic endpoints: {"type": "error", "error": {"type": "...", "message": "..."}}

Supported Parameters

OpenAI /v1/chat/completions

ParameterSupportedNotes
modelYesRayu model code
messagesYessystem, user, assistant, tool roles
max_tokens / max_completion_tokensYesDefault 4096 if omitted
temperatureYes
top_pYes
stopYesString or array
streamYes
stream_options.include_usageYesFinal usage chunk
toolsYesFunction calling
tool_choiceYesauto, none, required, specific function
nNoOnly n=1 supported
logprobs / top_logprobsNoRejected with 400
seedNoRejected with 400
logit_biasNoRejected with 400
frequency_penalty / presence_penaltyIgnored
response_formatNot yet

Anthropic /v1/messages

The full Anthropic Messages API is supported, including:

  • system, messages, max_tokens, temperature, top_p, stop_sequences
  • tools, tool_choice, tool_use / tool_result content blocks
  • stream with all event types
  • Image content blocks (base64)
  • Extended thinking (thinking content blocks)
  • Prompt caching (cache_control blocks) — charged at the model's cache read/write multipliers

Security Best Practices

  • Never expose API keys in client-side code (browsers, mobile apps). Keys are for server-to-server use.
  • Use per-key credit caps to limit blast radius if a key is leaked.
  • Set an expiry for keys used in temporary environments (CI, demos).
  • Use the model allowlist to prevent a leaked key from accessing expensive models.
  • Revoke immediately if you suspect a key has been compromised — revocation takes effect within seconds.
  • Rotate keys periodically as a hygiene practice.

Model List

To see which models are available on your plan:

curl https://gateway.rayucode.com/v1/models \
  -H "Authorization: Bearer rayu_sk_live_..."

Returns an OpenAI-compatible model list with capabilities (supportsReasoning, supportsImage, supportsTools, contextWindow).


How Model Fetch Works

This section explains how Rayu discovers, filters, and refreshes the model catalog — both for API key callers (this document) and for Rayu OAuth (Auth) callers. The two paths authenticate differently but enforce the same filtering rules and ordering.

The two Rayu providers

Rayu Auth (OAuth)Rayu API Key
Provider idrayu-hostedrayu
Kindrayu-hostedanthropic-compatible
CredentialRayu account JWT (from /login)rayu_sk_live_… API key
Chat endpoint{gateway}/anthropic/v1/messages (JWT-injecting fetch){gateway}/anthropic/v1/messages (key as Bearer)
Model catalog endpointGET {backend}/me/entitlementsGET {gateway}/v1/models
Registered byAuto on /login/connect wizard or RAYU_API_KEY env
Display nameRayuRayu API Key

They are deliberately separate providers so a user can have both configured without either clobbering the other's credential — the same reasoning as anthropic vs claude-subscription.


Path 1 — Rayu Auth (OAuth / JWT)

Authentication flow

User → Google OAuth (rayu-web) → rayu-backend /api/auth/oauth/google
  → issues Rayu JWT (signed with RAYU_JWT_SECRET)
  → CLI stores JWT in ~/.rayu/rayu-auth.json

Model fetch flow

CLI → GET {backend}/me/entitlements  (Authorization: Bearer <JWT>)
  → backend resolves user's active plan
  → backend queries hosted_models WHERE enabled=true AND provider.enabled=true
  → returns TWO lists:
      • allowedModels  — plan-filtered subset (drives entitlement/gating)
      • hostedModels   — full enabled catalog (shown to ALL signed-in users)

Backend query (NestJS + Prisma)

// All enabled models (the catalog shown to every signed-in user)
findEnabled(): Promise<HostedModelWithProvider[]> {
  return this.prisma.hostedModel.findMany({
    where: { enabled: true, provider: { enabled: true } },
    orderBy: [{ sortOrder: 'asc' }, { id: 'asc' }],
    include: WITH_PROVIDER,
  })
}

// Plan-allowed subset (drives entitlement)
async findAllowedForPlan(planCode: string) {
  const all = await this.findEnabled()
  return all.filter((m) => this.allowedCodes(m).includes(planCode))
}

Entitlements response shape

{
  "plan": { "code": "pro", "name": "pro", "priceCents": 2900 },
  "allowedModels": [
    { "code": "deepseek-v4-pro", "label": "DeepSeek V4 Pro",
      "contextWindow": 131072, "supportsReasoning": true, "supportsImage": false,
      "supportsTools": true, "creditMultiplier": 1.0 }
  ],
  "hostedModels": [
    { "code": "deepseek-v4-pro", "label": "DeepSeek V4 Pro", ... },
    { "code": "deepseek-v4-flash", "label": "DeepSeek V4 Flash", ... }
  ]
}

How the CLI uses it

// Visibility uses the full catalog; usability uses the entitled subset.
const catalog = ent?.hostedModels ?? ent?.allowedModels ?? []
const entitled = ent?.allowedModels ?? []
const models = catalog.map((m) => m.code)
  • hostedModels present → shows ALL enabled models (Free users see them but are gated on use; a model is usable iff it also appears in allowedModels).
  • hostedModels absent (older backend) → falls back to allowedModels (plan-filtered only).

Model ordering

Models are returned in ORDER BY sortOrder ASC, id ASC. The admin dashboard's reorder UI sets sortOrder = index × 10. The CLI preserves this order exactly — no client-side sorting.

Auto-refresh

  • On login: syncRayuHostedProvider() is called with the fresh entitlements.
  • Background: getCachedEntitlements() kicks a rate-limited (30s cooldown) background refresh on every read.
  • On /model open: the model picker calls refreshHostedCatalog() which re-fetches entitlements and re-renders only if the catalog changed.

Path 2 — Rayu API Key (this path)

Authentication flow

User → rayucode.com/dashboard/api-keys → creates rayu_sk_live_… key
  → pastes key into /connect → CLI sends it as Bearer to gateway

Model fetch flow

CLI → GET {gateway}/v1/models  (Authorization: Bearer <key>)
  → gateway resolves user's identity from the key
  → gateway resolves user's active plan
  → gateway filters: enabled model + enabled provider + plan-allowed + key-allowlist
  → returns OpenAI list shape with capabilities

Gateway query (Rust + SQLx)

-- Loads ALL hosted_models rows (no WHERE — filtering happens in Rust)
SELECT m.*, p.*
FROM hosted_models m
JOIN providers p ON p.id = m.provider_id
ORDER BY m.sortOrder, m.id

Gateway filtering (Rust)

pub fn allowed_models(models: &[HostedModel], plan_code: &str) -> Vec<HostedModel> {
    models
        .iter()
        .filter(|m| {
            m.enabled
                && m.provider.enabled    // matches backend's findEnabled()
                && m.allowed_plan_codes.iter().any(|pc| pc == plan_code)
        })
        .cloned()
        .collect()
}

Then the API key's own allowlist is applied on top:

fn visible_chat_models(ent: &Entitlement, api_key: Option<&ApiKeyContext>) -> Vec<&HostedModel> {
    ent.allowed_models
        .iter()
        .filter(|m| api_key.is_none_or(|ak| ak.allows_model(&m.code)))
        .collect()
}

/v1/models response shape

{
  "object": "list",
  "data": [
    {
      "id": "deepseek-v4-pro",
      "object": "model",
      "created": 1700000000,
      "owned_by": "rayu",
      "label": "DeepSeek V4 Pro",
      "supportsReasoning": true,
      "supportsImage": false,
      "supportsTools": true,
      "contextWindow": 131072
    }
  ]
}

How the CLI uses it

export function parseRayuCatalog(payload: unknown) {
  // Sanitize every id, dedupe, preserve gateway order (no .sort())
  return {
    models: entries.map(e => e.code),
    modelLabels: hostedModelLabels(entries),
    modelContextWindows: hostedContextWindows(entries),
  }
}

Model ordering

The gateway returns models in ORDER BY m.sortOrder, m.id. The CLI preserves this order — no client-side sorting. This matches the Auth path exactly.

Auto-refresh

  • On connect: the /connect wizard fetches the catalog and persists it.
  • On /model open: the model picker calls refreshRayuApiKeyCatalog() which re-fetches from the gateway and re-renders only if the catalog changed.
  • Background: refreshActiveProviderModels() calls refreshRayuApiKeyCatalog() when the provider is active.

Filtering rules (both paths)

A model appears in the catalog only if ALL of these are true:

RuleAuth pathAPI key path
Model enabled = trueBackend findEnabled()Gateway allowed_models()
Provider enabled = trueBackend findEnabled()Gateway m.provider.enabled
Model's allowedPlanCodes includes user's planBackend findAllowedForPlan()Gateway allowed_models()
API key's allowed_models includes the modelN/A (no key)Gateway visible_chat_models()

An empty allowedPlanCodes means NOBODY for chat models — a model must be explicitly granted to a plan. (The opposite rule applies to media models, where empty means EVERY plan.)


Per-key controls

An API key can further narrow the catalog via the dashboard:

ControlEffect on /v1/models
Model allowlistOnly listed models appear (intersected with plan)
Empty allowlistNo restriction — full plan catalog
Credit capDoesn't affect listing; enforced on request path
Rate limit (RPM)Doesn't affect listing; enforced on request path

A key allowlist is a narrowing, never a grant: it cannot add a model the plan doesn't include.


Stale-default pruning

Both paths prune the user's chosen default/small model if the admin removes it from the catalog. Holding on to a removed code would send every request to a model the gateway now rejects (403 "model not available"), which reads like a CLI bug rather than a catalog change.

// Auth path
const inCatalog = (code?: string): boolean => !!code && models.includes(code)
defaultModel: inCatalog(existing?.defaultModel) ? existing?.defaultModel : preferredCode

// API key path
if (!cur.defaultModel || !result.models.includes(cur.defaultModel)) {
  cur.defaultModel = fallback.defaultModel
}

Config persistence

Both providers store their catalog in ~/.rayu/providers.json:

{
  "id": "rayu",
  "kind": "anthropic-compatible",
  "baseURL": "https://gateway.rayucode.com/anthropic",
  "apiKey": "rayu_sk_live_...",
  "models": ["deepseek-v4-pro", "deepseek-v4-flash"],
  "fetchedModels": ["deepseek-v4-pro", "deepseek-v4-flash"],
  "modelLabels": { "deepseek-v4-pro": "DeepSeek V4 Pro" },
  "modelContextWindows": { "deepseek-v4-pro": 131072 },
  "defaultModel": "deepseek-v4-pro",
  "smallFastModel": "deepseek-v4-flash"
}
  • models — the catalog in display order (what /model shows).
  • fetchedModels — same list, tracked separately for refresh detection.
  • modelLabels — admin display names, keyed by model id.
  • modelContextWindows — admin context windows in tokens, keyed by model id.

Error handling

FailureAuth pathAPI key path
Network errorKeep cached catalog, log diagnosticKeep cached catalog, log diagnostic
401 (bad JWT/key)Remove provider, prompt re-loginRemove provider, prompt re-enter key
403 (account inactive)Keep cache, surface errorKeep cache, surface error
503 (gateway DB down)Keep cache, don't blame userKeep cache, don't blame user
Empty catalogRemove provider (no models to show)Keep cache, let /model refresh later

Both paths never throw and never empty the cache on a transient failure — an offline launch or a gateway blip must not wipe the model list.


Quick reference — API endpoints

Backend (rayu-backend, NestJS)

MethodPathAuthDescription
GET/api/me/entitlementsRayu JWTPlan, features, allowedModels, hostedModels

Gateway (rayu-gateway-rust)

MethodPathAuthDescription
POST/anthropic/v1/messagesJWT or API keyChat (Anthropic Messages format)
POST/v1/chat/completionsAPI keyChat (OpenAI format)
GET/v1/modelsAPI keyAvailable chat models for the caller's plan
GET/v1/models?media=imageAPI keyAvailable image-generation models
GET/v1/models?media=videoAPI keyAvailable video-generation models
GET/v1/creditsAPI keyCredit balance and usage