FAQ
Everything you might want to know
Models, API access, custom enterprise workflows, deployment, and billing, in plain language.
Platform and access
Is RealRelay.ai a model provider?
No. RealRelay.ai is not a model provider. Instead, it connects to upstream model providers to offer reliable, stable, secure, and unified smart routing and model access. It supports multiple API formats compatible with mainstream providers, including OpenAI, Anthropic, and Gemini.
What is RealRelay.ai, in plain English?
RealRelay.ai is a Unified Agentic AI Platform that provides unified, secure model access alongside Agent development, deployment, and execution capabilities. For enterprise users, we also offer private/on-premise deployments and custom Agent implementations tailored to specific business workflows.
Who is RealRelay.ai designed for?
RealRelay.ai caters to international developers, startups, and enterprise clients located outside Mainland China. We offer unified, secure AI model routing alongside end-to-end Agent development and deployment services tailored for global operations.
Do I have to be technical to use it?
No. Developers can use the API and console directly. Business teams can review workflow examples and contact us to discuss a custom implementation without having to design the technical architecture themselves.
What base URL and key do I use?
Point your client at https://www.realrelay.ai/v1 and authenticate with a key from the console Keys page, sent as Authorization: Bearer. Anthropic-style x-api-key and Gemini-style x-goog-api-key headers are accepted on their respective endpoints.
Do I have to rewrite an existing OpenAI integration?
Usually not. Point base_url at https://www.realrelay.ai/v1, swap in a RealRelay.ai key, and pick a model id from the catalog. The request shape stays the same.
Which SDKs can I use?
You can use OpenAI-compatible SDKs such as the official Python and Node.js openai packages, as well as the Anthropic SDK against /v1/messages or the Gemini SDK against the native /v1beta path. The API docs include copyable examples for OpenAI-compatible Python and Node.js clients, the Anthropic SDK, and cURL.
Do you support Anthropic and Gemini endpoints natively?
Yes. /v1/messages speaks the Anthropic request shape and /v1beta/models/* the Gemini shape, so those SDKs keep working with only a base URL and key change.
Models and routing
Where does the catalog data come from?
From a catalog snapshot the operator publishes from the platform — the model page shows the date it was published. Struck-through original prices come from our channel list-price records; what we charge always comes from the platform.
How does the automatic model "failover" work?
When you request a model by its bare name, RealRelay.ai picks the best channel serving it according to your routing strategy (lowest price by default, switchable in the console). If your first-choice channel is slow or down, it instantly reroutes to the best backup in under 10ms, so your users never see an outage.
Can I pin a specific channel?
Yes. Write the bare model name to let the gateway pick the cheapest channel serving it, or prefix it with a channel name to lock a specific one. The routing strategy (default, lowest price, lowest latency) is an account-level setting in the console.
How do I know what a model supports?
The catalog lists reported context length, capabilities such as tools, vision and reasoning, and supported endpoints from the latest published catalog snapshot.
What happens when a model goes offline?
The website reflects the latest published catalog snapshot. After a model is removed from that snapshot, its detail URL returns 404 and it is omitted from the sitemap.
Are streaming and tool calling supported?
Streaming works via the standard stream parameter over SSE. Tool calling depends on the model. The catalog lists which models report it.
Account and billing
How does the unified API key work?
RealRelay.ai acts as a proxy wrapper. Use one API key to call the models that key can access. For a bare model ID, we dynamically route the request across available channels, optimizing for latency or price.
Where do I create an API key?
In the console Keys page. A key can carry a quota, a model allowlist, an expiry date and an IP allowlist, so you can issue separate keys per environment.
How is billing structured?
You prepay credits in the console wallet, and usage is deducted per million tokens at the published rate, at direct provider cost without hidden markups. Image and media models bill per call. The catalog shows the mode alongside the price.
How is billing structured across plans?
Pay-as-you-go is prepaid per million tokens with hard budget caps. Enterprise is invoiced monthly against a credit line with SLAs and SSO. Sovereign Core is a custom annual contract for fully isolated deployments. You can change tiers as you grow; see the pricing page for the current credit tiers.
What happens when my balance runs out?
Requests return 402 insufficient_quota. Top up the wallet in the console to resume service.
How do I verify the detailed bill?
RealRelay.ai provides detailed token usage logs for model requests. Users can query these usage logs in the console to verify granular billing details and fee breakdowns.
Can I request a refund after making payment?
Payments are non-refundable by default. However, to accommodate accidental purchases, users may request a refund within 24 hours of payment.
Privacy and enterprise
Where does my data go? Is it private?
Before any request reaches a public AI model, RealRelay.ai's PII redaction layer masks sensitive details (medical info, card numbers, passwords). Responses are safely un-masked for you. For maximum control, you can run RealRelay.ai entirely on your own infrastructure.
Can I run it on my own servers (on-premise)?
Yes. For organizations with strict compliance needs (government, hospitals, banks), RealRelay.ai can be deployed on your own server arrays or an isolated local cloud. In that setup, you have greater control over your data. The Security page details the sovereign deployment options.
What are the rate limits?
Limits apply per API key and depend on your plan; current values are shown in the console. On 429, retry with exponential backoff.
Will the platform store my model access data?
No. RealRelay.ai does not store or analyze user model access data. It acts strictly as a routing gateway, transparently forwarding data between users and model providers.
Will the platform use my data to optimize its capabilities?
No. RealRelay.ai strictly adheres to user privacy protection principles and will never use any user data to train or optimize platform capabilities.
How does the platform collaborate with model providers?
RealRelay.ai accesses models by establishing direct enterprise contracts with model providers via their official enterprise APIs. Under these enterprise agreements, providers deliver stricter data privacy protections and enhanced API service stability.
Create a key, change one line, and use the models that key can access.
