Documentation
Text generation
Calling text models over the OpenAI, Anthropic and Gemini wire protocols, with parameters, streaming and the limits of protocol compatibility.
Your first request
Text models answer synchronously: one request, one response. Point an
official SDK at https://www.realrelay.ai/v1 and send a model ID
from the catalog.
1from openai import OpenAI23client = OpenAI(4 base_url="https://www.realrelay.ai/v1",5 api_key="sk-***",6)78response = client.chat.completions.create(9 model="jd/glm-5.2",10 messages=[{"role": "user", "content": "Hello"}],11)1213print(response.choices[0].message.content)1import OpenAI from "openai";23const client = new OpenAI({4 baseURL: "https://www.realrelay.ai/v1",5 apiKey: process.env.REALRELAY_API_KEY,6});78const response = await client.chat.completions.create({9 model: "jd/glm-5.2",10 messages: [{ role: "user", content: "Hello" }],11});1213console.log(response.choices[0].message.content);1curl https://www.realrelay.ai/v1/chat/completions \2 -H "Authorization: Bearer $REALRELAY_API_KEY" \3 -H "Content-Type: application/json" \4 -d '{5 "model": "jd/glm-5.2",6 "messages": [{ "role": "user", "content": "Hello" }]7 }'Three wire protocols
The same text model can be addressed in three request shapes. Pick the one your existing code already speaks — there is no performance or pricing difference between them, only the shape of the JSON.
OpenAI — /v1/chat/completions and
/v1/responses
Chat completions is the default, and nearly every text model in the catalog accepts it — reach for it unless you have a reason not to.
A few models are served only on the newer Responses shape, so a model
page may show this sample instead of the chat one. It takes
input in place of
messages, and caps the reply with
max_output_tokens rather than
max_tokens.
1from openai import OpenAI23client = OpenAI(4 base_url="https://www.realrelay.ai/v1",5 api_key="sk-***",6)78response = client.responses.create(9 model="openai/gpt-5.5-pro",10 input="Hello",11 max_output_tokens=1024,12)1314print(response.output_text)1import OpenAI from "openai";23const client = new OpenAI({4 baseURL: "https://www.realrelay.ai/v1",5 apiKey: process.env.REALRELAY_API_KEY,6});78const response = await client.responses.create({9 model: "openai/gpt-5.5-pro",10 input: "Hello",11 max_output_tokens: 1024,12});1314console.log(response.output_text);1curl https://www.realrelay.ai/v1/responses \2 -H "Authorization: Bearer $REALRELAY_API_KEY" \3 -H "Content-Type: application/json" \4 -d '{5 "model": "openai/gpt-5.5-pro",6 "input": "Hello",7 "max_output_tokens": 10248 }'Anthropic — /v1/messages
Accepts the Anthropic SDK unchanged. max_tokens is
required by this shape; send it explicitly rather than relying on a default,
which can be far smaller than you expect. Also avoid sending
temperature and top_p
together — some models accept only one of the two, and the second is
dropped rather than refused.
1import anthropic23client = anthropic.Anthropic(4 base_url="https://www.realrelay.ai",5 api_key="sk-***",6)78message = client.messages.create(9 model="anthropic/claude-sonnet-4-6",10 max_tokens=1024,11 messages=[{"role": "user", "content": "Hello"}],12)1314print(message.content[0].text)1import Anthropic from "@anthropic-ai/sdk";23const client = new Anthropic({4 baseURL: "https://www.realrelay.ai",5 apiKey: process.env.REALRELAY_API_KEY,6});78const message = await client.messages.create({9 model: "anthropic/claude-sonnet-4-6",10 max_tokens: 1024,11 messages: [{ role: "user", content: "Hello" }],12});1314console.log(message.content[0].text);1curl https://www.realrelay.ai/v1/messages \2 -H "x-api-key: $REALRELAY_API_KEY" \3 -H "anthropic-version: 2023-06-01" \4 -H "Content-Type: application/json" \5 -d '{6 "model": "anthropic/claude-sonnet-4-6",7 "max_tokens": 1024,8 "messages": [{ "role": "user", "content": "Hello" }]9 }'Gemini — /v1beta
Accepts the Gemini SDK and its x-goog-api-key
header. Streaming is selected by the URL — the
:streamGenerateContent action, or
?alt=sse — rather than by a field in the body.
1from google import genai23client = genai.Client(4 api_key="sk-***",5 http_options={"base_url": "https://www.realrelay.ai"},6)78response = client.models.generate_content(9 model="jd/gemini-3.1-flash-image-preview",10 contents="Hello",11)1213print(response.text)1import { GoogleGenAI } from "@google/genai";23const client = new GoogleGenAI({4 apiKey: process.env.REALRELAY_API_KEY,5 httpOptions: { baseUrl: "https://www.realrelay.ai" },6});78const response = await client.models.generateContent({9 model: "jd/gemini-3.1-flash-image-preview",10 contents: "Hello",11});1213console.log(response.text);1curl "https://www.realrelay.ai/v1beta/models/jd/gemini-3.1-flash-image-preview:generateContent" \2 -H "x-goog-api-key: $REALRELAY_API_KEY" \3 -H "Content-Type: application/json" \4 -d '{5 "contents": [{ "parts": [{ "text": "Hello" }] }]6 }'Parameters
Shown for /v1/chat/completions. Semantics match
OpenAI; compatible parameters not listed here are forwarded upstream
unchanged.
| Parameter | Type | Required | Description |
|---|---|---|---|
model | string | Yes | Model ID. A bare name uses smart routing; channel/name pins one channel. See Model IDs. |
messages | array | Yes | Conversation messages. Each item has a role (system, user, assistant) and content. |
stream | boolean | — | Return incremental results over SSE. Defaults to false. |
temperature | number | — | Sampling temperature between 0 and 2. Lower is more deterministic. |
max_tokens | integer | — | Maximum tokens to generate. The ceiling depends on the selected model. |
tools | array | — | Tool definitions the model may call. Requires a model that reports tool use. |
response_format | object | — | Request structured output, for example the json_object type. |
Streaming
Set stream: true to receive server-sent events,
terminated by data: [DONE]. On
/v1beta the URL selects streaming instead, as
above.
1from openai import OpenAI23client = OpenAI(4 base_url="https://www.realrelay.ai/v1",5 api_key="sk-***",6)78stream = client.chat.completions.create(9 model="jd/glm-5.2",10 messages=[{"role": "user", "content": "Explain gradient descent"}],11 stream=True,12)1314for chunk in stream:15 delta = chunk.choices[0].delta.content16 if delta:17 print(delta, end="", flush=True)1import OpenAI from "openai";23const client = new OpenAI({4 baseURL: "https://www.realrelay.ai/v1",5 apiKey: process.env.REALRELAY_API_KEY,6});78const stream = await client.chat.completions.create({9 model: "jd/glm-5.2",10 messages: [{ role: "user", content: "Explain gradient descent" }],11 stream: true,12});1314for await (const chunk of stream) {15 const delta = chunk.choices[0]?.delta?.content;16 if (delta) process.stdout.write(delta);17}1curl -N https://www.realrelay.ai/v1/chat/completions \2 -H "Authorization: Bearer $REALRELAY_API_KEY" \3 -H "Content-Type: application/json" \4 -d '{5 "model": "jd/glm-5.2",6 "messages": [{ "role": "user", "content": "Explain gradient descent" }],7 "stream": true8 }'910# data: {"choices":[{"delta":{"content":"Gradient"}}]}11# data: {"choices":[{"delta":{"content":" descent"}}]}12# data: [DONE]If you proxy these requests through infrastructure of your own, disable response buffering on that route — otherwise the increments arrive batched and the stream stops being a stream.
Protocol compatibility
The catalog records which shape each model advertises natively. Requests in another shape are translated, which is best-effort rather than guaranteed: every conversion passes through the OpenAI shape, so what does not exist there cannot survive the trip.
Converting between OpenAI and either Anthropic or Gemini is well-trodden.
Converting between Anthropic and Gemini is two conversions back to back. We do not recommend it, and we do not promise fidelity for it — if you need one of those two shapes, choose a model that advertises it.
Embeddings, reranking, audio and realtime have no conversion at all. A request in the wrong shape for those is rejected, not translated.
