GLM-5.3-Flash API
Call GLM-5.3-Flash through Wokey and review published catalog pricing, upstream-native API forms, context, and model capabilities.
- Model ID
glm-5.3-flash- Vendors
- zhipu, volcengine, opencode
- Supply status
- Currently callable
Catalog pricing snapshot
Prices are per 1M tokens; request-time pricing controls settlement.
| Meter | Wokey | Official reference | Savings |
|---|---|---|---|
| Input | $0.075 | $0.15 | 50% |
| Output | $0.25 | $0.5 | 50% |
| Cache read | $0.015 | $0.03 | 50% |
| Cache write | — | — | — |
Official price source: https://docs.z.ai/guides/overview/pricing
Model capabilities
- Context
- 1,000,000 tokens
- Max output
- 131,072 tokens
- Streaming
- Supported
- Tools
- Supported
- Vision
- Supported
- Supported API forms
/v1/chat/completions·/v1/messages
Send a request
/v1/chat/completions
curl https://api.wokey.ai/v1/chat/completions \
-H "Authorization: Bearer $WOKEY_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"glm-5.3-flash","messages":[{"role":"user","content":"Hello"}]}'/v1/messages
curl https://api.wokey.ai/v1/messages \
-H "x-api-key: $WOKEY_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{"model":"glm-5.3-flash","max_tokens":1024,"messages":[{"role":"user","content":"Hello"}]}'Use GLM-5.3-Flash in your tools
The gateway is https://api.wokey.ai (OpenAI-compatible clients usually take https://api.wokey.ai/v1), and the model ID is glm-5.3-flash. Open the guide for the client you use.
- Use GLM-5.3-Flash in Claude Code: Anthropic Messages · base URL and API key setup
- Use GLM-5.3-Flash in Claude Desktop: Anthropic Messages · base URL and API key setup
- Use GLM-5.3-Flash in opencode: OpenAI Chat Completions · base URL and API key setup
- Use GLM-5.3-Flash in OpenClaw: OpenAI Chat Completions · base URL and API key setup
- Use GLM-5.3-Flash in Hermes Agent: OpenAI Chat Completions · base URL and API key setup
- OpenRouter alternative: Wokey vs OpenRouter: Compare token prices, payment fees, API forms, and upstream sources.
- Verifiable AI API: which models and API forms carry response proofs: Response proofs currently cover only raw passthrough Claude Messages and GPT Responses on official routes.
Frequently asked questions
How much does the GLM-5.3-Flash API cost, and how much does it save versus the official API?
Wokey lists input at $0.075 per 1M tokens and output at $0.25 per 1M tokens; the official references are $0.15 and $0.5. That is 50% lower for input and 50% lower for output. Request-time pricing controls settlement, and the vendor pricing source is linked on this page.
What is the GLM-5.3-Flash API model ID?
The model ID for GLM-5.3-Flash on Wokey is glm-5.3-flash. Use it exactly as written in the request body's model field or in your client's model setting (for example ANTHROPIC_MODEL in Claude Code, or model in the Codex CLI config.toml); GET https://api.wokey.ai/v1/models also lists it.
Which API forms does GLM-5.3-Flash support, and how do I call it through Wokey?
GLM-5.3-Flash supports /v1/chat/completions (Chat Completions requests) and /v1/messages (Messages requests). Set the API base URL to https://api.wokey.ai, authenticate with a Wokey API key, and use the canonical model ID glm-5.3-flash; you can start with /v1/chat/completions. The Send a request section includes a curl example for every supported API form.
How do I use GLM-5.3-Flash in Claude Code?
Claude Code calls GLM-5.3-Flash over /v1/messages. Set ANTHROPIC_BASE_URL="https://api.wokey.ai", ANTHROPIC_AUTH_TOKEN to your Wokey API key, and ANTHROPIC_MODEL="glm-5.3-flash", then start claude. Claude Desktop also uses /v1/messages; the setup guides cover both step by step.
What are the context window and maximum output for GLM-5.3-Flash?
The catalog context window is 1,000,000 tokens and the maximum output is 131,072 tokens. These are separate limits, and a request must also satisfy the upstream constraints in effect when it is processed.
How are cache reads and cache writes priced for GLM-5.3-Flash?
Cache read: Wokey lists $0.015 per 1M tokens and the official reference is $0.03 per 1M tokens. Cache write: the vendor does not publish a separate meter, so Wokey also shows — as the public price.
How does calling GLM-5.3-Flash through Wokey differ from OpenRouter?
Wokey lists GLM-5.3-Flash at $0.075 input / $0.25 output per 1M tokens, against an official reference of $0.15 / $0.5. OpenRouter states it adds no markup on inference and charges its fee when you buy credits (5.5% by card). Wokey carries a curated set of models and has no per-provider routing parameters, so OpenRouter fits better if you need a wider catalog. The OpenRouter alternative page has the full comparison.
How do GLM-5.3-Flash and GLM-5.3 differ in price and context?
In the published price comparison, GLM-5.3-Flash lists input/output at $0.075 / $0.25 with a 1,000,000-token context window; GLM-5.3 lists $0.28 / $0.88 with a 1,000,000-token context window. This is a factual price-and-capacity comparison, not a quality ranking; also compare their supported API forms and capabilities.
Related models and documentation
Get API key · glm-5.3-flash