Qwen
qwen3.8-flash
Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart anal
Aggregated modalities
Channel
Each channel is compared at its lowest input-price tier; output breaks ties. A dash means /api/pricing does not provide that rate. Selecting a channel updates the endpoints and call example below.
ChannelUSD per 1M tokensInputOutputCache readCache write
qwen/qwen3.8-flashInput
$0.15 USD per 1M tokens
Output
$0.47 USD per 1M tokens
Cache read
- USD per 1M tokens
Cache write
- USD per 1M tokens
Endpoints
POST
/v1/chat/completionsChat CompletionsCall it
Using qwen/qwen3.8-flash
1from openai import OpenAI23client = OpenAI(4 base_url="https://www.realrelay.ai/v1",5 api_key="sk-***",6)78response = client.chat.completions.create(9 model="qwen/qwen3.8-flash",10 messages=[{"role": "user", "content": "Hello"}],11)1213print(response.choices[0].message.content)1import OpenAI from "openai";23const client = new OpenAI({4 baseURL: "https://www.realrelay.ai/v1",5 apiKey: process.env.REALRELAY_API_KEY,6});78const response = await client.chat.completions.create({9 model: "qwen/qwen3.8-flash",10 messages: [{ role: "user", content: "Hello" }],11});1213console.log(response.choices[0].message.content);1curl https://www.realrelay.ai/v1/chat/completions \2 -H "Authorization: Bearer $REALRELAY_API_KEY" \3 -H "Content-Type: application/json" \4 -d '{5 "model": "qwen/qwen3.8-flash",6 "messages": [{ "role": "user", "content": "Hello" }]7 }'Prices and availability as published on Oct 4, 2026.
Capabilities
ReasoningStructured outputTool useVision
Input
TextImageVideo
Output
Text
