Qwen API tutorial
How to authenticate, which endpoint to call, and the mistakes that cost people an afternoon.
The Qwen API is OpenAI-compatible: the official documentation shows requests built with the standard OpenAI SDKs. In practice that means three values decide whether your first call works — the API key, the base URL and the model ID.
1. Create an account and an API key
- Sign in to the developer platform: www.aliyun.com/product/bailian.
- Create an API key in the console and store it in an environment variable rather than in source code.
- Full details: API documentation
2. Point your client at the right endpoint
https://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/compatible-mode/v1 (legacy: https://dashscope.aliyuncs.com)
Overseas endpoints documented: Singapore https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1 (legacy dashscope-intl.aliyuncs.com), US https://dashscope-us.aliyuncs.com/compatible-mode/v1, Tokyo https://{WorkspaceId}.ap-northeast-1.maas.aliyuncs.com/compatible-mode/v1
If you are migrating existing code, keep your current SDK and change only base_url and the model name. If you see 401 responses, the key is usually wrong; if you see 404 on the model, the model ID is wrong or deprecated.
3. A minimal working request
from openai import OpenAI
client = OpenAI(
api_key=os.environ["DASHSCOPE_API_KEY"],
base_url="https://dashscope.aliyuncs.com/compatible-mode/v1", # legacy domain; workspace domains are also documented
)
resp = client.chat.completions.create(
model="qwen-plus",
messages=[{"role": "user", "content": "Explain vector databases in three sentences."}],
)
print(resp.choices[0].message.content)
4. Watch out for these
- The base URL is migrating to workspace-specific domains; older documentation still shows dashscope.aliyuncs.com style URLs.
- Prices differ by region and by input-length tier, so the numbers above are examples, not a quote.
- Saving-plan and Token Plan prices are not published.
- Old model IDs get retired: check the model list before hard-coding a name in production.
- Context limits are documented per model — sending more input than the window allows is the most common cause of rejection in long-document workloads.
5. Controlling cost
- Cache-hit input tokens are cheaper than cache-miss input tokens, so keep prompts stable where possible.
- Batch or off-peak options are documented for some models and are billed below the standard rate.
- Current rates: official pricing page.
6. Next steps
- Qwen overview: what the product is, which models exist and what each is documented to do.
- Published Qwen pricing: free allowances, token rates and subscription tiers, with sources.
- Qwen vs Western AI models: how the interfaces, context limits and availability compare.
- Qwen in the AI-Mind directory: ratings, categories and alternatives.
Facts on this page were checked against the official pages linked above on 2026-09-19. Prices and model IDs change frequently: confirm them on the vendor’s own pricing page before you rely on them.