Qwen-Proxy is a proxy service that converts https://chat.qwen.ai and Qwen Code / Qwen Cli into an OpenAI-compatible API. With this project, you only need one account to use any OpenAI API-compatible client (such as ChatGPT-Next-Web, LobeChat, etc.) to call various models from https://chat.qwen.ai and Qwen Code / Qwen Cli. Models under the /cli endpoint are provided by Qwen Code / Qwen Cli, supporting 256k context and native tools parameter support.
Main Features:
- Compatible with OpenAI API format for seamless integration with various clients
- Compatible with Anthropic Messages API (
/v1/messages), supporting Claude Code, Anthropic SDK, and other clients - Supports Function Calling (OpenAI
tools/ Anthropictools), including streamingargumentsincremental chunks andtool_choice=requiredstrict validation retry - Supports multi-account polling to improve availability
- Supports streaming/non-streaming responses
- Supports multimodal (image recognition, video understanding, image/video generation)
- Supports OpenAI-style resource endpoints:
/v1/images/generations,/v1/images/edits,/v1/videos - Supports advanced features like smart search and deep thinking
- Supports CLI endpoints with 256K context and tool calling capabilities
- Provides web management interface for easy configuration and monitoring
- Batch account addition supports real-time progress display, adjustable login concurrency in system settings
Each account can be configured with its own outbound proxy, allowing multiple accounts to use different IPs simultaneously, avoiding IP-based association bans from chat.qwen.ai.
Priority: account.proxy > Global PROXY_URL > No proxy
Supported Proxy Protocols: HTTP / HTTPS / SOCKS5 / SOCKS5H (consistent with PROXY_URL)
Frontend Configuration (Recommended): Open the management panel → Fill in the "Proxy Address" field when adding accounts, or click the "Modify Proxy" button on existing account cards.
ENV Configuration (DATA_SAVE_MODE=none):
# Old format (backward compatible, per-account proxy left empty)
ACCOUNTS=user1@mail.com:pass1,user2@mail.com:pass2
# New format (separated by | for proxy URLs, can be mixed with old format)
ACCOUNTS=user1@mail.com:pass1|http://10.0.0.1:8080,user2@mail.com:pass2|socks5://10.0.0.2:1080File mode (data/data.json) schema:
{
"accounts": [
{
"email": "user@mail.com",
"password": "...",
"token": "...",
"expires": 1234567890,
"proxy": "http://10.0.0.1:8080"
}
]
}When the proxy field is null or missing, the account falls back to the global PROXY_URL (if configured).
⚠️ Note: Proxy URLs returned by the interface are not sanitized. This project assumes it runs in a trusted local or private network environment used by a single administrator.
- Bun 1.3.14+ (source deployments; Docker pins 1.3.14)
- Node.js 24+ (development regression tests and lint only, not the production image)
- Docker (optional)
- Redis (optional, for data persistence)
Create a .env file and configure the following parameters:
# 🌐 Service Configuration
LISTEN_ADDRESS=localhost # Listen address
SERVICE_PORT=3000 # Service port
# 🔐 Security Configuration
API_KEY=sk-123456,sk-456789 # API key (required, supports multiple keys)
ACCOUNTS= # Account configuration (format: user1:pass1[|proxy_url],user2:pass2[|proxy_url])
# 🔍 Feature Configuration
SEARCH_INFO_MODE=table # Search info display mode (table/text)
OUTPUT_THINK=true # Whether to output thinking process (true/false)
LEGACY_REASONING_IN_CONTENT=false # Reasoning format, false=reasoning_content field, true=legacy <think> inside content (true/false)
SIMPLE_MODEL_MAP=false # Simplify model mapping (true/false)
MODELS_CACHE_TTL=3600 # Model list cache TTL in seconds, 0=never expires
AGENT_TURN_MAX_TOOL_CALLS=24 # Anthropic path: text-channel tool_use cap per agent turn (4-256), upstream cut after it
# 🌐 Proxy and Reverse Proxy Configuration
QWEN_CHAT_PROXY_URL= # Custom Chat API reverse proxy URL (default: https://chat.qwen.ai)
QWEN_CLI_PROXY_URL= # Custom CLI API reverse proxy URL (default: https://portal.qwen.ai)
PROXY_URL= # HTTP/HTTPS/SOCKS5/SOCKS5H proxy address (example: http://127.0.0.1:7890)
# 🗄️ Data Storage
DATA_SAVE_MODE=none # Data save mode (none/file/redis)
REDIS_URL= # Redis connection address (optional, rediss:// for TLS)
BATCH_LOGIN_CONCURRENCY=5 # Login concurrency during batch account addition
# 📸 Cache Configuration
CACHE_MODE=default # Image cache mode (default/file)| Parameter | Description | Example |
|---|---|---|
LISTEN_ADDRESS |
Service listen address | localhost or 0.0.0.0 |
SERVICE_PORT |
Service running port | 3000 |
API_KEY |
API access key, supports multi-key configuration. The first is the admin key (can access frontend management page), others are regular keys (API calls only). Multiple keys separated by commas | sk-admin123,sk-user456,sk-user789 |
SEARCH_INFO_MODE |
Search result display format | table or text |
OUTPUT_THINK |
Whether to show AI thinking process | true or false |
LEGACY_REASONING_IN_CONTENT |
Reasoning output format. Default false = reasoning goes to a separate reasoning_content field; true = legacy behavior (<think> inside content) |
true or false |
SIMPLE_MODEL_MAP |
Simplify model mapping, return basic models without variants only | true or false |
MODEL_MAP |
Incoming model name mapping: alias=qwen-id,...,*=fallback. Exact entry wins (trailing [..] stripped, case-insensitive), existing Qwen ids pass through, everything else uses *; applies to /v1/chat/completions and /v1/messages only. Also editable at runtime in the dashboard (Settings → Model mapping); a dashboard-saved map overrides this variable, see .env.example |
*=qwen3.8-max-thinking |
MODELS_CACHE_TTL |
Model list cache TTL in seconds; after expiry the next request refreshes it from upstream; 0 = never expires |
3600 |
AGENT_TURN_MAX_TOOL_CALLS |
Anthropic path: cap on text-channel tool_use blocks per agent turn (4–256). When the model runs away after a narrated [TOOL CALL] (repeats the same call hundreds of times, hallucinates a whole session), the upstream is cut right after the N-th admitted call and the admitted calls are delivered with stop_reason=tool_use; once a call was admitted in an earlier delta, a duplicate, a rejected call or prose/thinking also cuts the turn |
24 |
AGENT_CONTEXT_FILE_THRESHOLD_BYTES |
Externalize complete Agent tool definitions and history as a Qwen text document when the request body exceeds this size, avoiding the roughly 128 KiB WAF limit | 92160 (90 KiB) |
AGENT_CONTEXT_LIVE_PROMPT_BYTES |
Maximum size of the tool protocol and current turn kept in the live request after context externalization | 49152 (48 KiB) |
QWEN_CHAT_PROXY_URL |
Custom Chat API reverse proxy address | https://your-proxy.com |
QWEN_CLI_PROXY_URL |
Custom CLI API reverse proxy address | https://your-cli-proxy.com |
PROXY_URL |
Outbound request proxy address, supports HTTP/HTTPS/SOCKS5/SOCKS5H | http://127.0.0.1:7890 |
DATA_SAVE_MODE |
Data persistence method | none/file/redis |
REDIS_URL |
Redis database connection address, use rediss:// protocol when using TLS encryption |
redis://localhost:6379 or rediss://xxx.upstash.io |
BATCH_LOGIN_CONCURRENCY |
Login concurrency during batch account addition, can be adjusted dynamically in frontend system settings | 5 |
CACHE_MODE |
Image cache storage method | default/file |
| LOG_LEVEL | Log level | DEBUG/INFO/WARN/ERROR |
ENABLE_FILE_LOG |
Enable file logging | true or false |
LOG_DIR |
Log file directory | ./logs |
MAX_LOG_FILE_SIZE |
Maximum log file size (MB) | 10 |
MAX_LOG_FILES |
Number of log files to retain | 5 |
💡 Tip: You can create a free Redis instance at Upstash, use
rediss://...format when using TLS protocol
The API_KEY environment variable supports configuring multiple API keys to implement access control with different permission levels:
Configuration Format:
# Single key (admin privileges)
API_KEY=sk-admin123
# Multiple keys (first is admin, others are regular users)
API_KEY=sk-admin123,sk-user456,sk-user789Permission Description:
| Key Type | Permission Scope | Function Description |
|---|---|---|
| Admin Key | Full Permissions | • Access frontend management page • Modify system settings • Call all API interfaces • Add/delete regular keys |
| Regular Key | API Call Permissions | • API interface calls only • Cannot access frontend management page • Cannot modify system settings |
Use Cases:
- Team Collaboration: Assign different permission API keys to different team members
- Application Integration: Provide restricted API access permissions for third-party applications
- Security Isolation: Separate management permissions from regular usage permissions
Notes:
- The first API_KEY automatically becomes the admin key with highest privileges
- Admins can dynamically add or delete regular keys through the frontend page
- All keys can normally call API interfaces, permission differences only affect management functions
The CACHE_MODE environment variable controls the storage method of image caching to optimize image upload and processing performance:
| Mode | Description | Use Case |
|---|---|---|
default |
Memory cache mode (default) | Single process deployment, cache lost after restart |
file |
File cache mode | Multi-process deployment, cache persisted to ./caches/ directory |
Recommended Configuration:
- Single Process Deployment: Use
CACHE_MODE=default, best performance - Multi-Process/Cluster Deployment: Use
CACHE_MODE=file, ensure inter-process cache sharing - Docker Deployment: Recommend using
CACHE_MODE=fileand mounting./cachesdirectory
File Cache Directory Structure:
caches/
├── [signature1].txt # Cache file containing image URL
├── [signature2].txt
└── ...
docker run -d \
-p 3000:3000 \
-e API_KEY=sk-admin123,sk-user456,sk-user789 \
-e DATA_SAVE_MODE=none \
-e CACHE_MODE=file \
-e ACCOUNTS= \
-v ./caches:/app/caches \
--name qwen2api \
rfym21/qwen2api:latest# Download configuration file
curl -o docker-compose.yml https://raw.githubusercontent.com/Rfym21/Qwen2API/refs/heads/main/docker/docker-compose.yml
# Start service
docker compose pull && docker compose up -d# Clone project
git clone https://github.com/Rfym21/Qwen2API.git
cd Qwen2API
# Install dependencies
bun install --frozen-lockfile
bun install --cwd public --frozen-lockfile
bun run build:frontend
# Configure environment variables
cp .env.example .env
# Edit .env file
# Start one Bun process
bun run start
# Development mode
bun run devThe service runs in one Bun process; PM2 and automatic cluster startup have been removed.
PM2_INSTANCES and PM2_MAX_MEMORY no longer apply. Compose handles restarts with
restart: always; use mem_limit for container memory limits. Account statistics,
rate limits, and runtime settings include process-local state, so multiple replicas
need coordinated state before sharing persistence.
bun run test retains the Node test runner for the existing regression suite.
bun run test:bun starts a real Bun service with a local mock upstream to verify
login, frontend files, WASM, and OpenAI/Anthropic JSON and SSE responses without real accounts.
Source execution is the default. node src/server.js remains available for temporary comparisons.
Run bun run build:binary to build the frontend and a native executable in pkg_dist/.
For Linux x64, use bun run build:binary --target bun-linux-x64 (use bun-linux-x64-musl
for Alpine). Cross-compilation does not verify execution on the target platform.
Test a Windows build with bun run test:bun --binary pkg_dist/qwen2api-windows-x64.exe.
The executable embeds Bun, the backend, frontend assets, and tiktoken WASM; deployment
does not need Node, Bun, or node_modules. Run it from a writable directory containing
your .env, or pass environment variables. Credentials and .env are not embedded.
The binary writes data/, logs/, and caches/ under the working directory by default;
set QWEN2API_RUNTIME_DIR to override this root. --version prints the version without
starting the server. Docker uses Alpine plus the musl executable, without source files,
node_modules, or a separate Bun CLI.
ci.yml validates pushes and PRs without publishing. verify.yml is shared by CI and
releases: Node 24 runs lint/regressions, Bun is selected from packageManager, the frontend
is built once, and Windows x64/Linux x64/Linux arm64 binaries are built and tested on native
runners. Linux jobs also test the actual Alpine containers before exporting their images.
release.yml publishes only when package.json version changes on main, or through a
manual Run workflow with publish checked on main. Unchecked manual runs only build artifacts.
Other package fields, normal source changes, PRs, tag pushes, and new branches without a
previous commit do not automatically publish.
Releases include three .tar.gz archives (executable, .env.example, instructions), SHA256SUMS, and linux/amd64 + linux/arm64 Docker images. Downloads use glibc; containers use musl. Tested image archives are published without rebuilding. Version tags cannot move to another commit; published releases cannot be replaced. Failed drafts may be retried at the same commit. latest is updated only while main has the same version and the release commit remains in its history; later docs commits do not suppress it. Publish jobs use queue:max (up to 100 pending runs), so later package-only changes do not replace pending releases. Publishing across registries is not atomic: a failed run can leave a draft and partial image tags.
Configure DOCKERHUB_USERNAME and DOCKERHUB_TOKEN secrets and a writable Docker Hub repository.
The publish job alone gets contents:write and uses the release environment, where you can add
main-only restrictions and approvals. Bump the existing version before the first new release.
Local build: docker build -f docker/Dockerfile -t qwen2api:local ..
On Linux, test it with bun run test:bun --docker-image qwen2api:local.
Quickly deploy to Hugging Face Spaces:
Quickly deploy to Vercel:
This platform entry retains Node/Express compatibility; the Bun runtime migration targets local and Docker deployments.
Need to configure environment variables:
ACCOUNTS=email:password
SERVICE_PORT=80
API_KEY=sk-xxx
DATA_SAVE_MODE=none
Qwen2API/
├── README.md
├── README-en.md
├── bun.lock # Bun dependency lockfile
├── package.json
│
├── docker/ # Docker configuration directory
│ ├── Dockerfile
│ ├── docker-compose.yml
│ └── docker-compose-redis.yml
│
├── caches/ # Cache file directory
├── data/ # Data file directory
│ ├── data.json
│ └── data_template.json
├── scripts/ # Script directory
│ └── fingerprint-injector.js # Browser fingerprint injection script
│
├── src/ # Backend source code directory
│ ├── server.js # Main server file
│ ├── config/
│ │ └── index.js # Configuration file
│ ├── controllers/ # Controllers directory
│ │ ├── chat.js # Chat controller
│ │ ├── chat.image.video.js # Image/video generation controller
│ │ ├── cli.chat.js # CLI chat controller
│ │ └── models.js # Models controller
│ ├── middlewares/ # Middlewares directory
│ │ ├── authorization.js # Authorization middleware
│ │ └── chat-middleware.js # Chat middleware
│ ├── models/ # Models directory
│ │ └── models-map.js # Model mapping configuration
│ ├── routes/ # Routes directory
│ │ ├── accounts.js # Account routes
│ │ ├── chat.js # Chat routes
│ │ ├── cli.chat.js # CLI chat routes
│ │ ├── models.js # Model routes
│ │ ├── settings.js # Settings routes
│ │ └── verify.js # Verification routes
│ └── utils/ # Utility functions directory
│ ├── account-rotator.js # Account rotator
│ ├── account.js # Account management
│ ├── chat-helpers.js # Chat helper functions
│ ├── cli.manager.js # CLI manager
│ ├── cookie-generator.js # Cookie generator
│ ├── data-persistence.js # Data persistence
│ ├── fingerprint.js # Browser fingerprint generation
│ ├── img-caches.js # Image cache
│ ├── logger.js # Logging utility
│ ├── precise-tokenizer.js # Precise tokenizer
│ ├── proxy-helper.js # Proxy helper functions
│ ├── redis.js # Redis connection
│ ├── request.js # HTTP request wrapper
│ ├── setting.js # Setting management
│ ├── ssxmod-manager.js # Ssxmod parameter management
│ ├── token-manager.js # Token manager
│ ├── tools.js # Tool calling processing
│ └── upload.js # File upload
│
└── public/ # Frontend project directory
├── dist/ # Compiled frontend files
│ ├── assets/ # Static resources
│ ├── favicon.png
│ └── index.html
├── src/ # Frontend source code
│ ├── App.vue # Main application component
│ ├── main.js # Entry file
│ ├── style.css # Global styles
│ ├── assets/ # Static resources
│ │ └── background.mp4
│ ├── routes/ # Route configuration
│ │ └── index.js
│ └── views/ # Page components
│ ├── auth.vue # Authentication page
│ ├── dashboard.vue # Dashboard page
│ └── settings.vue # Settings page
├── package.json # Frontend dependency configuration
├── package-lock.json
├── index.html # Frontend entry HTML
├── postcss.config.js # PostCSS configuration
├── tailwind.config.js # TailwindCSS configuration
├── vite.config.js # Vite build configuration
└── public/ # Public static resources
└── favicon.png
This API supports multi-key authentication mechanism. All API requests require a valid API key in the request header:
Authorization: Bearer sk-your-api-keySupported Key Types:
- Admin Key: First configured API_KEY, has full permissions
- Regular Key: Other configured API_KEYs, API calls only
Authentication Example:
# Using admin key
curl -H "Authorization: Bearer sk-admin123" http://localhost:3000/v1/models
# Using regular key
curl -H "Authorization: Bearer sk-user456" http://localhost:3000/v1/chat/completionsGet the list of all available AI models.
GET /v1/models
Authorization: Bearer sk-your-api-keyGET /models (no authentication required)Description:
- id: Recommended to use directly as model in requests, prioritizing more readable model names
- name: Upstream original model ID, convenient for official interfaces or logs reference
upstream_id: Upstream model ID without capability suffixdisplay_name: Display name without capability suffix- When
SIMPLE_MODEL_MAP=false, additionally returns capability variants like-thinking,-search,-image,-video,-image-edit, etc.
Response Example:
{
"object": "list",
"data": [
{
"id": "Qwen3-Omni-Flash-image",
"name": "qwen3-omni-flash-2025-12-01-image",
"upstream_id": "qwen3-omni-flash-2025-12-01",
"display_name": "Qwen3-Omni-Flash",
"object": "model",
"created": 1677610602,
"owned_by": "qwen"
}
]
}Send chat messages and get AI responses.
POST /v1/chat/completions
Content-Type: application/json
Authorization: Bearer sk-your-api-keyRequest Body:
{
"model": "Qwen3.6-Plus",
"messages": [
{
"role": "system",
"content": "You are a helpful assistant."
},
{
"role": "user",
"content": "Hello, please introduce yourself."
}
],
"stream": false,
"temperature": 0.7,
"max_tokens": 2000
}Response Example:
{
"id": "chatcmpl-123",
"object": "chat.completion",
"created": 1677652288,
"model": "qwen3.6-plus",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Hello! I am an AI assistant..."
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 20,
"completion_tokens": 50,
"total_tokens": 70
}
}/v1/chat/completions supports the complete OpenAI Function Calling protocol. Even if the upstream web interface doesn't have native tools capability, this service makes its behavior consistent with OpenAI API through prompt injection and streaming state machine parsing:
- Automatically compresses
tools[]into TS-style signatures injected into prompts, saving about 70% token overhead - Streaming output follows OpenAI specification: first sends
function.name + empty argumentsheader block, followed by multipleargumentsslices assistant.tool_callsandrole:"tool"in historical messages automatically fold back in chain,tool_call_idprecisely associatedtool_choiceall four states:"auto"/"required"/{type:"function",function:{name:"..."}}/"none"- When
tool_choice="required"or specifying function, if no tool call triggered initially, automatically appends strong constraint prompt for retry once - Automatically retries once when the upstream returns reasoning only, with no visible text or executable tool call, preventing an empty terminal Agent turn
- Treats a clean HTTP EOF from Qwen Web as normal
stop/tool_calls; only actual transport failures such as connection resets become stream errors - Externalizes complete tool definitions and history through Qwen's official file APIs once the Agent request reaches the safety threshold, while keeping the current turn live to avoid WAF/captcha failures during long tool loops
- Surfaces HTTP-200 WAF/captcha business frames as
upstream_waf_challengeinstead of disguising them as an empty success or a generic 502
For long-running Agents such as Codex, Claude Code, and OpenClaw, keep the default 90 KiB / 48 KiB thresholds. Lower
AGENT_CONTEXT_FILE_THRESHOLD_BYTESif your reverse proxy adds a substantial request body overhead.
Request Example:
{
"model": "qwen3-coder-plus",
"stream": true,
"messages": [
{"role": "user", "content": "Check Beijing's weather"}
],
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get city weather",
"parameters": {
"type": "object",
"properties": { "city": { "type": "string" } },
"required": ["city"]
}
}
}
],
"tool_choice": "required"
}Streaming Response (excerpt):
data: {"choices":[{"delta":{"tool_calls":[{"index":0,"id":"call_xxx","type":"function","function":{"name":"get_weather","arguments":""}}]}}]}
data: {"choices":[{"delta":{"tool_calls":[{"index":0,"function":{"arguments":"{\"city\":\"Beijing\"}"}}]}}]}
data: {"choices":[{"delta":{},"finish_reason":"tool_calls"}]}
data: [DONE]
OpenAI SDK, LangChain, Cline, Continue, and other clients following OpenAI tool protocols can be directly integrated.
Provides an Anthropic-compatible bridge for the /v1/messages endpoint, allowing direct use with common clients such as Claude Code, the Anthropic SDK, and aider.
Note: Qwen2API is a compatibility bridge, not an Anthropic-equivalent backend. For unsupported fields, the current strategy is intentionally permissive: the request is accepted whenever possible, and important unsupported fields are surfaced through response headers and server logs instead of being silently ignored. Fields such as
system, multi-turnmessages,tools,tool_choice, andthinkingare currently approximate compatibility, not native Anthropic semantics.
POST /v1/messages
Content-Type: application/json
Authorization: Bearer sk-your-api-keyCompatibility matrix:
| Field / capability | Status | Current behavior | Notes / client impact |
|---|---|---|---|
model |
Supported | Mapped to a Qwen model name | Any resolvable Qwen model ID can be used |
messages.text |
Supported | Supports basic text messages | Standard chat clients work |
messages.image |
Supported | Supports image blocks and translates them into the internal image format | Suitable for common multimodal clients |
messages.tool_use |
Partial | Accepts Anthropic-style tool_use history blocks, then translates/folds them internally |
Not native upstream tool-call semantics |
messages.tool_result |
Partial | Accepts tool_result and converts it into bridge-layer tool-result text |
Details such as is_error are not guaranteed to be preserved |
system |
Partial | Merged into the prompt prefix | Not preserved as a native upstream system layer |
messages (multi-turn) |
Partial | Multi-turn history is compacted / translated | Structured conversation semantics are approximate |
tools[] |
Partial | Supports the basic {name,input_schema,description} shape |
Implemented through prompt/XML simulation, not native upstream tool execution |
tool_choice |
Partial | Supports the basic auto / any / tool / none modes |
Relies on prompt steering and retry hints, not an upstream hard guarantee |
thinking |
Partial | Currently accepts legacy thinking: {type:"enabled", budget_tokens:N} and maps it approximately |
Not equivalent to Anthropic's newer adaptive thinking / effort semantics |
stream |
Supported | Returns an Anthropic-style SSE event sequence | Suitable for Claude Code and other streaming clients |
max_tokens |
Ignored with warning | Currently does not enforce an upstream output limit | Exposed through warning headers / logs |
stop_sequences |
Ignored with warning | Not currently mapped to upstream stop behavior | Exposed through warning headers / logs |
metadata |
Ignored with warning | Not used in the upstream request | Exposed through warning headers / logs |
temperature / top_p / top_k |
Ignored with warning | Not currently mapped to upstream sampling controls | Exposed through warning headers / logs |
service_tier |
Ignored with warning | Not supported | Exposed through warning headers / logs |
container |
Ignored with warning | Not supported | Exposed through warning headers / logs |
output_config |
Ignored with warning | Does not currently support official structured outputs / effort semantics | Exposed through warning headers / logs |
mcp_servers |
Not supported yet | Anthropic MCP runtime semantics are not supported | Currently surfaced as a risk warning; a later version may convert this into an explicit error |
context_management |
Not supported yet | Official compaction / context-editing semantics are not supported | Currently surfaced as a risk warning; a later version may convert this into an explicit error |
When a request includes approximate or unsupported fields, the response may include these headers:
X-Qwen2API-Anthropic-CompatibilityX-Qwen2API-Anthropic-Warnings
These headers indicate which Anthropic capabilities are Partial and which fields were Ignored with warning. They do not change the basic successful response body shape.
Request Example (with tool calling):
{
"model": "qwen3-coder-plus",
"max_tokens": 1024,
"messages": [
{"role": "user", "content": "Check Guangzhou weather"}
],
"tools": [
{
"name": "get_weather",
"input_schema": {
"type": "object",
"properties": { "city": { "type": "string" } },
"required": ["city"]
}
}
],
"tool_choice": { "type": "any" }
}Note:
max_tokensis accepted in the example above, but it is currently Ignored with warning and does not enforce an upstream output limit the way the official Anthropic API does.
Non-streaming Response:
{
"id": "msg_xxx",
"type": "message",
"role": "assistant",
"model": "qwen3-coder-plus",
"content": [
{
"type": "tool_use",
"id": "call_xxx",
"name": "get_weather",
"input": { "city": "Guangzhou" }
}
],
"stop_reason": "tool_use",
"stop_sequence": null,
"usage": { "input_tokens": 233, "output_tokens": 25 }
}Streaming SSE Event Sequence:
event: message_start
data: {"type":"message_start","message":{...}}
event: content_block_start
data: {"type":"content_block_start","index":0,"content_block":{"type":"tool_use","id":"call_xxx","name":"get_weather","input":{}}}
event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"input_json_delta","partial_json":"{\"city\":\"Guangzhou\"}"}}
event: content_block_stop
data: {"type":"content_block_stop","index":0}
event: message_delta
data: {"type":"message_delta","delta":{"stop_reason":"tool_use","stop_sequence":null},"usage":{"input_tokens":234,"output_tokens":25}}
event: message_stop
data: {"type":"message_stop"}
Currently supports two calling methods:
- Using
/v1/chat/completions+ model suffix:-image,-image-edit,-video - Using OpenAI-style resource endpoints:
/v1/images/generations,/v1/images/edits,/v1/videos
In the following examples, please refer to the id field returned by /v1/models for model names.
Text-to-image:
{
"model": "Qwen3-Omni-Flash-image",
"messages": [
{
"role": "user",
"content": "Draw a kitten playing in the garden, cartoon style"
}
],
"size": "1:1",
"stream": false
}Image editing:
{
"model": "Qwen3-Omni-Flash-image-edit",
"messages": [
{
"role": "user",
"content": [
{
"type": "text",
"text": "Change this image to light blue tech style poster"
},
{
"type": "image_url",
"image_url": {
"url": "data:image/png;base64,..."
}
}
]
}
],
"stream": false
}Video generation:
{
"model": "Qwen3-Omni-Flash-video",
"messages": [
{
"role": "user",
"content": "Generate a 3-second night scene time-lapse video, city street neon lights flickering"
}
],
"size": "9:16",
"stream": false
}Supported Size Parameters:
- Image/video generation under
/v1/chat/completionssupports1:1,4:3,3:4,16:9,9:16 /v1/images/generations,/v1/images/edits,/v1/videoscompatible with1024x1024,1536x1024,1024x1536,1792x1024,1024x1792
Image generation:
POST /v1/images/generations
Content-Type: application/json
Authorization: Bearer sk-your-api-key{
"model": "Qwen3-Omni-Flash",
"prompt": "An orange cat sitting on a wooden table looking at the camera, realistic style",
"size": "1024x1024",
"response_format": "url"
}Image editing:
POST /v1/images/edits
Content-Type: multipart/form-data
Authorization: Bearer sk-your-api-keyForm fields:
- model: Optional, automatically selects default model supporting image editing when not passed
- prompt: Optional, defaults to "Please complete the edit based on the uploaded image"
- image: Required, supports multipart file upload, also supports JSON string form of image URL/data URI
- size: Optional, supports OpenAI-style size notation
response_format: Optional, supports url,b64_json
Video generation:
POST /v1/videos
Content-Type: application/json
Authorization: Bearer sk-your-api-key{
"model": "Qwen3-Omni-Flash",
"prompt": "A brief 3-second night scene time-lapse video, city street neon lights flickering",
"size": "1024x1792"
}Image generation response example:
{
"created": 1776126402,
"data": [
{
"url": "https://cdn.qwenlm.ai/output/example/generated-image.png"
}
]
}Video generation response example:
{
"id": "video_1776126509490",
"object": "video",
"created": 1776126509,
"model": "qwen3-omni-flash-2025-12-01",
"status": "completed",
"data": [
{
"url": "https://cdn.qwenlm.ai/output/example/generated-video.mp4"
}
]
}Append -search suffix to model name to enable search functionality:
{
"model": "Qwen3.6-Plus-search",
"messages": [...]
}Append -thinking suffix to model name to enable thinking process output:
{
"model": "Qwen3.6-Plus-thinking",
"messages": [...]
}Enable both search and reasoning functionality simultaneously:
{
"model": "Qwen3.6-Plus-thinking-search",
"messages": [...]
}API automatically handles image and video uploads, supports sending image/video URLs or Base64 data URIs in conversations.
Image understanding example:
{
"model": "Qwen3.5-Omni-Plus",
"messages": [
{
"role": "user",
"content": [
{
"type": "text",
"text": "What's in this picture?"
},
{
"type": "image_url",
"image_url": {
"url": "data:image/jpeg;base64,..."
}
}
]
}
]
}Video understanding example:
{
"model": "Qwen3.5-Omni-Plus",
"messages": [
{
"role": "user",
"content": [
{
"type": "text",
"text": "Describe this video in one sentence"
},
{
"type": "input_video",
"input_video": {
"url": "data:video/mp4;base64,..."
}
}
]
}
]
}Supported video fields:
input_videovideo_urlvideo
CLI endpoint accesses using Qwen Code / Qwen Cli's OAuth token, supports 256K context and tool calling (Function Calling).
Supported Models:
| Model ID | Description |
|---|---|
qwen3-coder-plus |
Qwen3 Coder Plus |
qwen3-coder-flash |
Qwen3 Coder Flash (faster speed) |
coder-model |
Qwen 3.5 Plus (with reasoning chain, 256K context) |
qwen3.5-plus |
Alias for coder-model, automatic redirect |
Send chat requests through CLI endpoint, supports streaming and non-streaming responses.
POST /cli/v1/chat/completions
Content-Type: application/json
Authorization: Bearer API_KEYRequest Body:
{
"model": "qwen3-coder-plus",
"messages": [
{
"role": "user",
"content": "Hello, please introduce yourself."
}
],
"stream": false,
"temperature": 0.7,
"max_tokens": 2000
}Using coder-model (i.e., Qwen 3.5 Plus) or its alias qwen3.5-plus:
{
"model": "coder-model",
"messages": [
{
"role": "user",
"content": "Write a quick sort algorithm."
}
],
"stream": false
}Streaming Request:
{
"model": "qwen3-coder-flash",
"messages": [
{
"role": "user",
"content": "Write a poem about spring."
}
],
"stream": true
}Response Format:
Non-streaming response is the same as standard OpenAI API format:
{
"id": "chatcmpl-123",
"object": "chat.completion",
"created": 1677652288,
"model": "qwen3-coder-plus",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Hello! I am an AI assistant..."
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 20,
"completion_tokens": 50,
"total_tokens": 70
}
}Streaming response uses Server-Sent Events (SSE) format:
data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1677652288,"model":"qwen3-coder-flash","choices":[{"index":0,"delta":{"content":"Hello"},"finish_reason":null}]}
data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1677652288,"model":"qwen3-coder-flash","choices":[{"index":0,"delta":{"content":"!"},"finish_reason":null}]}
data: [DONE]
- This project is for learning and communication purposes only, strictly prohibited for any commercial use.
- All consequences arising from the use of this project are borne by the user, and the project developer assumes no responsibility.
- This project does not guarantee the service availability or stability of
chat.qwen.aiandQwen Code / Qwen Cli. - Users should comply with laws and regulations in their respective regions, as well as Tongyi Qianwen's terms of service and usage policies.
- For infringement, please contact the author for deletion.
This project uses a personal learning and research only license:
- Prohibited from using this project or its derivatives for any commercial purposes, including but not limited to: selling, renting, providing paid services, embedding in commercial products, etc.
- Prohibited from using this project to perform any actions that violate Tongyi Qianwen's terms of service.
- Prohibited from using this project for large-scale automated calls, malicious attacks, or abuse of upstream services.

