GLM-5.3-Flash API

Call GLM-5.3-Flash through Wokey and review dated catalog pricing, upstream-native API forms, context, and model capabilities.

Model ID
glm-5.3-flash
Vendors
zhipu
Supply status
Currently callable

Catalog pricing snapshot

Prices are per 1M tokens; request-time pricing controls settlement.

MeterWokeyOfficial referenceSavings
Input$0.075$0.0750%
Output$0.25$0.250%
Cache read$0.015$0.0150%
Cache write

Price checked:

Official price source: https://docs.z.ai/guides/overview/pricing

Model capabilities

Context
1,000,000 tokens
Max output
131,072 tokens
Streaming
Supported
Tools
Supported
Vision
Supported
Supported API forms
/v1/chat/completions · /v1/messages

Send a request

/v1/chat/completions

curl https://api.wokey.ai/v1/chat/completions \
  -H "Authorization: Bearer $WOKEY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"glm-5.3-flash","messages":[{"role":"user","content":"Hello"}]}'

/v1/messages

curl https://api.wokey.ai/v1/messages \
+  -H "x-api-key: $WOKEY_API_KEY" \
+  -H "anthropic-version: 2023-06-01" \
+  -H "Content-Type: application/json" \
+  -d '{"model":"glm-5.3-flash","max_tokens":1024,"messages":[{"role":"user","content":"Hello"}]}'

Frequently asked questions

How much does the GLM-5.3-Flash API cost, and how much does it save versus the official API?

As of 2026-08-26, Wokey lists input at $0.075 per 1M tokens and output at $0.25 per 1M tokens; the official references are $0.075 and $0.25. That is 0% lower for input and 0% lower for output. Request-time pricing controls settlement, and the vendor pricing source is linked on this page.

Which API forms does GLM-5.3-Flash support, and how do I call it through Wokey?

GLM-5.3-Flash supports /v1/chat/completions (Chat Completions requests) and /v1/messages (Messages requests). Set the API base URL to https://api.wokey.ai, authenticate with a Wokey API key, and use the canonical model ID glm-5.3-flash; you can start with /v1/chat/completions. The Send a request section includes a curl example for every supported API form.

What are the context window and maximum output for GLM-5.3-Flash?

The catalog context window is 1,000,000 tokens and the maximum output is 131,072 tokens. These are separate limits, and a request must also satisfy the upstream constraints in effect when it is processed.

How are cache reads and cache writes priced for GLM-5.3-Flash?

Cache read: Wokey lists $0.015 per 1M tokens and the official reference is $0.015 per 1M tokens. Cache write: the vendor does not publish a separate meter, so Wokey also shows — as the public price.

How do GLM-5.3-Flash and GLM-5.3 differ in price and context?

In this price snapshot, GLM-5.3-Flash lists input/output at $0.075 / $0.25 with a 1,000,000-token context window; GLM-5.3 lists $0.56 / $1.76 with a 1,000,000-token context window. This is a factual price-and-capacity comparison, not a quality ranking; also compare their supported API forms and capabilities.

Related models and documentation

Get API key · glm-5.3-flash