Documentation
Calling the RealRelay.ai API
Base URL, authentication, model IDs, channel routing, the endpoint list and the error codes the API returns.
Quickstart
Three steps. If you already have an OpenAI integration, only the second one applies to you.
- 1
Install an official SDK
No proprietary client. Use the openai or anthropic package you already know. - 2
Point it at RealRelay.ai
Set base_url to https://www.realrelay.ai/v1 and use your RealRelay.ai key. - 3
Choose a model
Set model to an ID from the catalog. Nothing else in your code changes.
1pip install openai1npm install openai1from openai import OpenAI23client = OpenAI(4 base_url="https://www.realrelay.ai/v1",5 api_key="sk-***",6)78response = client.chat.completions.create(9 model="jd/glm-5.2",10 messages=[{"role": "user", "content": "Hello"}],11)1213print(response.choices[0].message.content)1import OpenAI from "openai";23const client = new OpenAI({4 baseURL: "https://www.realrelay.ai/v1",5 apiKey: process.env.REALRELAY_API_KEY,6});78const response = await client.chat.completions.create({9 model: "jd/glm-5.2",10 messages: [{ role: "user", content: "Hello" }],11});1213console.log(response.choices[0].message.content);1curl https://www.realrelay.ai/v1/chat/completions \2 -H "Authorization: Bearer $REALRELAY_API_KEY" \3 -H "Content-Type: application/json" \4 -d '{5 "model": "jd/glm-5.2",6 "messages": [{ "role": "user", "content": "Hello" }]7 }'Authentication
Create a key in the console. The default is bearer authentication in the
Authorization header. Two alternatives are
accepted so the vendor SDKs work unmodified:
x-api-key on
/v1/messages, and
x-goog-api-key — or a
?key= query parameter — on
/v1beta. They are equivalent; use whichever your
SDK sends.
1Authorization: Bearer sk-***2Content-Type: application/jsonNever ship a key in frontend code or commit it to a repository. Browser-side features should call your own backend, which holds the key and forwards the request.
Model IDs
A model ID has two forms, and both are real addresses you can send. A bare
canonical name such as glm-5.2 is routed to
whichever channel is serving that model. Adding a channel prefix, as in
jd/glm-5.2, pins the request to that one channel.
1# Routed: any channel serving this model2model="glm-5.2"34# Pinned: exactly this channel5model="jd/glm-5.2"The catalog lists the prefixed form, because that is the one that names a
single price. For the exact set of IDs your key accepts, call
GET /v1/models.
Channel routing
When you send a bare canonical name, a routing strategy picks the channel. Three are available, selected per account in the console:
lowest_price— the cheapest channel serving the model. This is the default when nothing is set.lowest_latency— the channel with the best recent response time.default— the platform order: channel priority first, then weight.
A prefixed model ID does not consult the strategy. Pinning a channel and asking for the cheapest one are two different requests, and the prefix wins.
Endpoints
The base URL is https://www.realrelay.ai. Which endpoints a model
supports varies; the catalog lists them per model. One of these is
asynchronous — see the poll step on its row.
/v1/chat/completionsChat completions. OpenAI-compatible and the default entry point for most integrations./v1/responsesResponses API, for tool calling and structured output./v1/messagesAnthropic-compatible shape, so the Anthropic SDK works unchanged./v1beta/models/{model}:generateContentGemini-native shape, served through the protocol compatibility layer./v1/images/generationsImage generation and editing./v1/video/generationsGET /v1/video/generations/{task_id}Video generation. Asynchronous: submit, then poll the task ID until it finishes./v1/modelsList the models your key can call.Errors
Errors use standard HTTP status codes with a JSON body. Branch on the status
and on error.code:
error.type says where the failure came from —
new_api_error for the platform, or the provider's
own type when the model provider is what failed — so it is too coarse to
branch on.
1{2"error": {3 "message": "...",4 "type": "new_api_error",5 "code": "insufficient_user_quota"6}7}bad_request_bodyThe request body is malformed or missing a required field.—The API key is missing, malformed, disabled or unknown. This failure carries no error code.access_deniedThe key exists but is not allowed to make this request — for example the caller is outside its IP allowlist.insufficient_user_quotaThe account balance or subscription allowance is exhausted.—Too many requests in the current window. Retry with exponential backoff; no Retry-After header is sent.model_not_foundNo enabled channel currently serves this model ID.do_request_failedThe request reached the model provider but did not complete. Safe to retry.Rate limits
Limits apply per API key and depend on your plan; current values are shown
in the console. Exceeding one returns 429 with the
error body above. There are no
X-RateLimit-* headers and no
Retry-After — the status code is the whole signal,
so use exponential backoff with your own schedule rather than waiting for a
hint that will not arrive.
Still stuck?
Check live usage and request logs in the console, or reach out for integration help.
