Documentation

Text generation

Calling text models over the OpenAI, Anthropic and Gemini wire protocols, with parameters, streaming and the limits of protocol compatibility.

Your first request

Text models answer synchronously: one request, one response. Point an official SDK at https://www.realrelay.ai/v1 and send a model ID from the catalog.

1from openai import OpenAI23client = OpenAI(4    base_url="https://www.realrelay.ai/v1",5    api_key="sk-***",6)78response = client.chat.completions.create(9    model="jd/glm-5.2",10    messages=[{"role": "user", "content": "Hello"}],11)1213print(response.choices[0].message.content)
1import OpenAI from "openai";23const client = new OpenAI({4  baseURL: "https://www.realrelay.ai/v1",5  apiKey: process.env.REALRELAY_API_KEY,6});78const response = await client.chat.completions.create({9  model: "jd/glm-5.2",10  messages: [{ role: "user", content: "Hello" }],11});1213console.log(response.choices[0].message.content);
1curl https://www.realrelay.ai/v1/chat/completions \2  -H "Authorization: Bearer $REALRELAY_API_KEY" \3  -H "Content-Type: application/json" \4  -d '{5    "model": "jd/glm-5.2",6    "messages": [{ "role": "user", "content": "Hello" }]7  }'

Three wire protocols

The same text model can be addressed in three request shapes. Pick the one your existing code already speaks — there is no performance or pricing difference between them, only the shape of the JSON.

OpenAI — /v1/chat/completions and /v1/responses

Chat completions is the default, and nearly every text model in the catalog accepts it — reach for it unless you have a reason not to.

A few models are served only on the newer Responses shape, so a model page may show this sample instead of the chat one. It takes input in place of messages, and caps the reply with max_output_tokens rather than max_tokens.

1from openai import OpenAI23client = OpenAI(4    base_url="https://www.realrelay.ai/v1",5    api_key="sk-***",6)78response = client.responses.create(9    model="openai/gpt-5.5-pro",10    input="Hello",11    max_output_tokens=1024,12)1314print(response.output_text)
1import OpenAI from "openai";23const client = new OpenAI({4  baseURL: "https://www.realrelay.ai/v1",5  apiKey: process.env.REALRELAY_API_KEY,6});78const response = await client.responses.create({9  model: "openai/gpt-5.5-pro",10  input: "Hello",11  max_output_tokens: 1024,12});1314console.log(response.output_text);
1curl https://www.realrelay.ai/v1/responses \2  -H "Authorization: Bearer $REALRELAY_API_KEY" \3  -H "Content-Type: application/json" \4  -d '{5    "model": "openai/gpt-5.5-pro",6    "input": "Hello",7    "max_output_tokens": 10248  }'

Anthropic — /v1/messages

Accepts the Anthropic SDK unchanged. max_tokens is required by this shape; send it explicitly rather than relying on a default, which can be far smaller than you expect. Also avoid sending temperature and top_p together — some models accept only one of the two, and the second is dropped rather than refused.

1import anthropic23client = anthropic.Anthropic(4    base_url="https://www.realrelay.ai",5    api_key="sk-***",6)78message = client.messages.create(9    model="anthropic/claude-sonnet-4-6",10    max_tokens=1024,11    messages=[{"role": "user", "content": "Hello"}],12)1314print(message.content[0].text)
1import Anthropic from "@anthropic-ai/sdk";23const client = new Anthropic({4  baseURL: "https://www.realrelay.ai",5  apiKey: process.env.REALRELAY_API_KEY,6});78const message = await client.messages.create({9  model: "anthropic/claude-sonnet-4-6",10  max_tokens: 1024,11  messages: [{ role: "user", content: "Hello" }],12});1314console.log(message.content[0].text);
1curl https://www.realrelay.ai/v1/messages \2  -H "x-api-key: $REALRELAY_API_KEY" \3  -H "anthropic-version: 2023-06-01" \4  -H "Content-Type: application/json" \5  -d '{6    "model": "anthropic/claude-sonnet-4-6",7    "max_tokens": 1024,8    "messages": [{ "role": "user", "content": "Hello" }]9  }'

Gemini — /v1beta

Accepts the Gemini SDK and its x-goog-api-key header. Streaming is selected by the URL — the :streamGenerateContent action, or ?alt=sse — rather than by a field in the body.

1from google import genai23client = genai.Client(4    api_key="sk-***",5    http_options={"base_url": "https://www.realrelay.ai"},6)78response = client.models.generate_content(9    model="jd/gemini-3.1-flash-image-preview",10    contents="Hello",11)1213print(response.text)
1import { GoogleGenAI } from "@google/genai";23const client = new GoogleGenAI({4  apiKey: process.env.REALRELAY_API_KEY,5  httpOptions: { baseUrl: "https://www.realrelay.ai" },6});78const response = await client.models.generateContent({9  model: "jd/gemini-3.1-flash-image-preview",10  contents: "Hello",11});1213console.log(response.text);
1curl "https://www.realrelay.ai/v1beta/models/jd/gemini-3.1-flash-image-preview:generateContent" \2  -H "x-goog-api-key: $REALRELAY_API_KEY" \3  -H "Content-Type: application/json" \4  -d '{5    "contents": [{ "parts": [{ "text": "Hello" }] }]6  }'

Parameters

Shown for /v1/chat/completions. Semantics match OpenAI; compatible parameters not listed here are forwarded upstream unchanged.

ParameterTypeRequiredDescription
modelstringYesModel ID. A bare name uses smart routing; channel/name pins one channel. See Model IDs.
messagesarrayYesConversation messages. Each item has a role (system, user, assistant) and content.
streamboolean—Return incremental results over SSE. Defaults to false.
temperaturenumber—Sampling temperature between 0 and 2. Lower is more deterministic.
max_tokensinteger—Maximum tokens to generate. The ceiling depends on the selected model.
toolsarray—Tool definitions the model may call. Requires a model that reports tool use.
response_formatobject—Request structured output, for example the json_object type.

Streaming

Set stream: true to receive server-sent events, terminated by data: [DONE]. On /v1beta the URL selects streaming instead, as above.

1from openai import OpenAI23client = OpenAI(4    base_url="https://www.realrelay.ai/v1",5    api_key="sk-***",6)78stream = client.chat.completions.create(9    model="jd/glm-5.2",10    messages=[{"role": "user", "content": "Explain gradient descent"}],11    stream=True,12)1314for chunk in stream:15    delta = chunk.choices[0].delta.content16    if delta:17        print(delta, end="", flush=True)
1import OpenAI from "openai";23const client = new OpenAI({4  baseURL: "https://www.realrelay.ai/v1",5  apiKey: process.env.REALRELAY_API_KEY,6});78const stream = await client.chat.completions.create({9  model: "jd/glm-5.2",10  messages: [{ role: "user", content: "Explain gradient descent" }],11  stream: true,12});1314for await (const chunk of stream) {15  const delta = chunk.choices[0]?.delta?.content;16  if (delta) process.stdout.write(delta);17}
1curl -N https://www.realrelay.ai/v1/chat/completions \2  -H "Authorization: Bearer $REALRELAY_API_KEY" \3  -H "Content-Type: application/json" \4  -d '{5    "model": "jd/glm-5.2",6    "messages": [{ "role": "user", "content": "Explain gradient descent" }],7    "stream": true8  }'910# data: {"choices":[{"delta":{"content":"Gradient"}}]}11# data: {"choices":[{"delta":{"content":" descent"}}]}12# data: [DONE]

If you proxy these requests through infrastructure of your own, disable response buffering on that route — otherwise the increments arrive batched and the stream stops being a stream.

Protocol compatibility

The catalog records which shape each model advertises natively. Requests in another shape are translated, which is best-effort rather than guaranteed: every conversion passes through the OpenAI shape, so what does not exist there cannot survive the trip.

  • Converting between OpenAI and either Anthropic or Gemini is well-trodden.

  • Converting between Anthropic and Gemini is two conversions back to back. We do not recommend it, and we do not promise fidelity for it — if you need one of those two shapes, choose a model that advertises it.

  • Embeddings, reranking, audio and realtime have no conversion at all. A request in the wrong shape for those is rejected, not translated.