Kimi K3 API

Kimi K3 is an open-weights model Moonshot AI released on 2026-07-17, with 2.8T total and 104B active parameters, native image and video input and a 1M context window. It lists at $3 input / $15 output per million tokens. This page covers the official and third-party benchmarks, how it compares with Claude Opus 5.5 and GLM-5.3, how to set reasoning_effort, the hardware needed to run it locally, the request rules that trip up integrations, and Wokey API key pricing and examples.

Price check: official, OpenRouter and Wokey

USD per million tokens. Official and OpenRouter rates were checked by hand on 2026-09-23, excluding tax; the Wokey rate is read from the live price list.

ChannelInputCache readOutput
Official API$3$0.3$15
OpenRouter$3$0.3$15
Wokey$0.9$0.09$4.5

The cheapest OpenRouter host is Sail Research's fp4-quantized deployment ($1.4989 input / $10.758 output). Quantized output can differ from the official weights, so it is not in the table.

· Official pricing page · OpenRouter endpoint list

Use it in Claude Code

This model supports /v1/messages on Wokey. Set these variables, then start claude.

export ANTHROPIC_BASE_URL="https://api.wokey.ai"
export ANTHROPIC_AUTH_TOKEN="$WOKEY_API_KEY"
export ANTHROPIC_MODEL="kimi-k3"

Kimi K3 API quickstart

Get a Wokey API key and follow these three steps to send your first request.

  1. Get an API key

    Register or sign in to Wokey, open the API console, and create and copy a key in API key management. Check your available balance before calling.

  2. Configure your client

    Use the base URL below in your OpenAI-compatible client, enter your Wokey API key, and select kimi-k3 as the model.

  3. Send your first request

    Replace YOUR_WOKEY_API_KEY in the example with your own key and run it in a terminal. Read the answer in choices, then review usage and charges in the API console.

Get a Wokey API key · Read API docs

Client base URL
https://api.wokey.ai/v1
Model ID
kimi-k3
Full request URL
https://api.wokey.ai/v1/chat/completions

The client base URL includes /v1. The curl example uses the full request URL.

Minimal request

export WOKEY_API_KEY="YOUR_WOKEY_API_KEY"

curl https://api.wokey.ai/v1/chat/completions \
  -H "Authorization: Bearer $WOKEY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "kimi-k3",
    "messages": [{"role": "user", "content": "Hello, Kimi!"}],
    "max_tokens": 1024
  }'

Kimi K3 at a glance

Release date
2026-07-17, on the Kimi web app, the Kimi Work desktop app, Kimi Code and the Kimi API platform on day one; the full weights went public on Hugging Face on 2026-07-27
Positioning
Moonshot says K3 still trails Claude Fable 5 and GPT-5.6 Sol overall but outperformed every other model in its evaluations, and recommends Kimi Code CLI as its agent framework
Model and weights
A 2.8T-parameter MoE with 104B active, routing each token to 16 of 896 experts, built on Kimi Delta Attention and Attention Residuals; weights at moonshotai/Kimi-K3, quantized to MXFP4 during training, about 1.56 TB in total, under the Kimi K3 License
Context and output
1,048,576-token context; max_completion_tokens defaults to 131,072 and goes up to 1,048,576; native text, image and video input, text output
Thinking and effort
Thinking is always on and responses carry reasoning_content; reasoning_effort accepts low / high / max with max as the default (only max existed at launch); temperature, top_p and the other sampling parameters are fixed and should be omitted
Official price
$3 input / $15 output per million tokens and $0.30 for cache hits; no context-length tiers
API and tools
The Kimi platform offers OpenAI- and Anthropic-compatible APIs with tool calling (tool_choice can be required), json_schema structured output and context caching; Wokey serves /v1/chat/completions and /v1/messages

· Sources: Moonshot · Hugging Face · K3 quickstart · Pricing · Artificial Analysis

Official benchmarks: Kimi K3 vs Fable 5 vs GPT-5.6 Sol vs Opus 4.8 vs GLM-5.2

From Moonshot’s model card on Hugging Face; these are vendor-reported results, and “—” means no score was published. The comparison models are those current at K3’s launch (2026-07), so Claude Opus 5.5 and GLM-5.3, both released later, are not in the table; see “Which to use” below for those comparisons.

BenchmarkKimi K3Fable 5GPT-5.6 SolOpus 4.8GLM-5.2
GPQA Diamond93.5%92.6%94.1%91.0%91.2%
HLE-Full (tools)56.0%63.0%58.0%57.9%—
DeepSWE v1.167.5%70.0%73.0%59.0%46.2%
Terminal-Bench 2.188.3%88.0%88.8%84.6%82.7%
ProgramBench77.8%76.8%77.6%71.9%63.7%
SWE-Marathon (pre-v1.1)42.0%35.0%39.0%40.0%13.0%
BrowseComp91.2%88.0%90.4%84.3%—
Toolathlon-Verified76.5%77.9%74.9%76.2%59.9%
OSWorld-Verified84.8%85.0%83.0%83.4%—
GDPval-AA v216861747173615931510
Video-MME (w/ subs)90.0%—89.5%86.0%—

Every K3 score uses reasoning_effort=max and temperature=1.0. The Fable 5 column ran with fallback, the card’s GPT-5.5 column is left out, and the HLE-Full row is the with-tools score. On DeepSWE, Terminal-Bench 2.1 and ProgramBench K3 ran in the Kimi Code harness, while most other scores come from public leaderboards or each model’s best harness, so the conditions differ: on the official DeepSWE leaderboard’s mini-SWE-agent harness K3 scores 67.3%. The SWE-Marathon row uses an H20-recalibrated task set from before the v1.1 release; Z.ai’s table gives K3 48.1% on v1.1.

Which to use: Kimi K3, Claude Opus 5.5 or GLM-5.3

vs Claude Opus 5.5

Opus 5.5 shipped on 2026-09-22, two months after K3. On Artificial Analysis’ Intelligence Index v4.3.2 Opus 5.5 scores 58 at max and K3 44, a clear gap. K3 lists at $3 / $15, only 25% below Opus 5.5’s $4 / $20, but Opus 5.5 used 260M output tokens across the index against K3’s 160M, so a task costs $5.98 versus $2.00 at official prices, about a third for K3. Opus 5.5 outputs about 92 tokens/s and K3 about 34. Pick Opus 5.5 for the hardest coding, long-horizon agents and work that has to be right the first time; pick K3 for high volume, tight budgets, video input or self-hosting.

vs GLM-5.3

The two are level on Artificial Analysis’ Intelligence Index: GLM-5.3 scores 45 at max and K3 44, at almost the same cost per task ($2.01 vs $2.00). GLM-5.3 lists at $1.4 / $4.4, less than half of K3, but writes longer outputs (210M tokens across the index versus 160M) and runs about twice as fast (about 66 vs 34 tokens/s). Z.ai’s GLM-5.3 launch page lists both: they are close on DeepSWE v1.1 (66.9% vs 67.5%), GLM-5.3 leads on Terminal Bench 3.0 (28.3% vs 17.4%), and K3 leads on SWE-Marathon v1.1 (48.1% vs 42.5%) and Toolathlon (76.5% vs 73.0%). Either works for text-only coding; pick K3 for image or video input, since GLM-5.3 is text-only; for self-hosting, GLM-5.3 (753B total parameters) is far easier than K3 (2.8T).

Choosing effort

K3 has only low, high and max, with max as the default, and thinking cannot be turned off. Artificial Analysis measured max at 44 for $2.00 per task and low at 30 for $1.15: low is about 40% cheaper per task but drops 14 points. Use the default max for coding and agents; there is no third-party data for high yet, so compare it with max on your own tasks; keep low for simple Q&A and format conversion. Reasoning tokens bill at the output rate, so higher effort means more output.

Kimi effort docs

Third-party and community reviews

  • Artificial Analysis: On Intelligence Index v4.3.2 it scores 44 at max and 30 at low; it used 160M output tokens across the index for $2.00 per task, outputs about 34 tokens/s and takes about 4.5 seconds to first token. On launch day, on the index version of the time, it scored 57 and ranked third, level with Claude Opus 4.8 and GPT-5.5 and behind Fable 5 and GPT-5.6 Sol, while its hallucination rate rose from K2.6’s 39% to 51%. Every Wokey meter is 30% of the official rate, so the same task comes to about $0.60 at max and $0.34 at low.
  • Fireworks: Fireworks, which also serves K3, ran it against Claude Fable 5 in the same agent loop on about 1,030 agentic tasks. On SWE tasks K3 passed 92.4% and Fable 5 92.6%, essentially level, but K3 took about 55 turns and 1.3M tokens per task against Fable 5’s 21 turns and 130K tokens.
  • Hacker News (2,107-point thread): Price drew the most discussion: many found $3 / $15 steep for an open-weights model, while others said it is fair if K3 really trails only Fable and Sol. Next came thinking length: some complained it thinks too long and repeats itself; Simon Willison’s pelican test produced 16,658 output tokens, 13,241 of them reasoning, for about $0.25; another commenter measured about 60% of K2.6’s reasoning tokens.

Using K3 in Claude Code, and request rules to check

These rules come from Kimi’s K3 quickstart and from issues Wokey has debugged in production. Check them before switching from K2.x to kimi-k3 or wiring it up for the first time.

  • Thinking cannot be turned off, and responses always carry reasoning_content. When migrating from K2.x, replace the old thinking parameter with the top-level reasoning_effort (low / high / max).
  • In multi-turn chats and tool calls, pass back the complete assistant message the API returned, including reasoning_content and tool_calls; sending back only content drops the earlier reasoning.
  • temperature, top_p, n, presence_penalty and frequency_penalty are fixed at 1.0, 0.95, 1, 0 and 0, and Kimi asks you to omit them from requests.
  • Images cannot be public URLs: put a base64 data URL such as data:image/png;base64,... in image_url, because an https image link returns 400 unsupported image url; video also goes in as base64. Kimi’s upload-then-reference flow with ms://<file_id> is not available through Wokey yet.
  • max_completion_tokens defaults to 131,072 and goes up to 1,048,576. Reasoning tokens bill as output and the default max effort writes long outputs, so leave room in output budgets.
  • With a stop sequence set, K3 can be cut off while still reasoning and return finish_reason stop with empty content. That comes from upstream and Wokey does not remove text, so test stop-dependent flows on your own tasks first.
  • Using it in Claude Code: kimi-k3 on Wokey supports /v1/messages, so set ANTHROPIC_BASE_URL to https://api.wokey.ai, ANTHROPIC_AUTH_TOKEN to your Wokey API key and ANTHROPIC_MODEL to kimi-k3; it costs $0.90 input / $4.50 output per million tokens. Moonshot recommends Kimi Code CLI as the agent framework, and OpenCode and other OpenAI-compatible clients use /v1/chat/completions.

Request body with an image (/v1/chat/completions)

{
  "model": "kimi-k3",
  "reasoning_effort": "high",
  "messages": [{
    "role": "user",
    "content": [
      { "type": "image_url", "image_url": { "url": "data:image/png;base64,iVBORw0KGgo..." } },
      { "type": "text", "text": "Find the layout bug in this screenshot" }
    ]
  }]
}
Model ID
kimi-k3
Vendors
moonshot, volcengine
Supply status
Currently callable

Catalog pricing snapshot

Prices are per 1M tokens; request-time pricing controls settlement.

MeterWokeyOfficial referenceSavings
Input$0.9$370%
Output$4.5$1570%
Cache read$0.09$0.370%
Cache write———

Official price source: https://platform.kimi.ai/docs/pricing/chat-k3

Model capabilities

Context
1,048,576 tokens
Max output
1,048,576 tokens
Streaming
Supported
Tools
Supported
Vision
Supported
Supported API forms
/v1/chat/completions · /v1/messages

Send a request

/v1/chat/completions

curl https://api.wokey.ai/v1/chat/completions \
  -H "Authorization: Bearer $WOKEY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"kimi-k3","messages":[{"role":"user","content":"Hello"}]}'

/v1/messages

curl https://api.wokey.ai/v1/messages \
  -H "x-api-key: $WOKEY_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{"model":"kimi-k3","max_tokens":1024,"messages":[{"role":"user","content":"Hello"}]}'

Use Kimi K3 in your tools

The gateway is https://api.wokey.ai (OpenAI-compatible clients usually take https://api.wokey.ai/v1), and the model ID is kimi-k3. Open the guide for the client you use.

Example production call

These usage values come from a completed Kimi K3 call that succeeded on its first attempt and was reconciled with its bill.

Total input
76023
Output
574
Cached share of input
99.3%
Call duration
25.60 s

Chat Completions usage

Compiled from the billing record into API field examples. prompt_tokens is total input, including cached tokens.

{
  "model": "kimi-k3",
  "usage": {
    "prompt_tokens": 76023,
    "prompt_tokens_details": {
      "cached_tokens": 75520
    },
    "completion_tokens": 574,
    "total_tokens": 76597
  }
}

Token accounting

These are the final per-million-token rates recorded for this call.

MeterTokensRate for this call / 1M
Uncached input503$0.9
Cache read75520$0.09
Output574$4.5
Billed amount
$0.009833

Cost = sum of tokens × their rate ÷ 1,000,000, rounded to six decimals after summing.

Source of this call

This call used the Volcengine channel at ark.cn-beijing.volces.com with API Key authorization. Usage was metered upstream and the request returned HTTP 200.

About Provider Node · Integration docs

Verify Prompt Cache

  1. Keep the shared prefix of long prompts identical and reuse context according to the selected API’s cache rules.
  2. Check prompt_tokens_details.cached_tokens; zero means no cache reads were recorded for this call.
  3. Uncached input = prompt_tokens minus cached tokens. Price each bucket separately and add the output charge.

Frequently asked questions

How do I get an API key for Kimi K3?

To call Kimi K3 through Wokey, use a Wokey API key. Register or sign in, open the /api console, then create and copy a key in API key management. Make sure your account has available balance before calling.

What base URL and model ID should I use for Kimi K3?

In an OpenAI-compatible client, use https://api.wokey.ai/v1 as the base URL and kimi-k3 as the model ID. The full HTTP endpoint is https://api.wokey.ai/v1/chat/completions. Pass your Wokey API key using Authorization: Bearer.

How do a Kimi web subscription and a Wokey API key work?

This example uses an API key created in your Wokey account and the Wokey gateway. Usage is billed against your Wokey balance. Kimi web subscriptions and keys issued by other platforms are managed by their respective platforms.

Kimi K3 vs Claude Opus 5.5: which is better?

On Artificial Analysis’ Intelligence Index v4.3.2 Opus 5.5 scores 58 at max and K3 44, so Opus 5.5 is clearly stronger, and it outputs about 2.7 times as fast. But a K3 index task costs $2.00, about a third of Opus 5.5’s $5.98, and its $3 / $15 list price is below Opus 5.5’s $4 / $20. Choose Opus 5.5 for the highest quality and first-time-right work; choose K3 for high-volume calls, tight budgets, video input or self-hosting.

Kimi K3 vs GLM-5.3: which is better?

On Artificial Analysis’ Intelligence Index GLM-5.3 scores 45 at max and K3 44, at almost the same cost per task ($2.01 vs $2.00). GLM-5.3 lists at $1.4 / $4.4, less than half of K3’s $3 / $15, and outputs about twice as fast. On Z.ai’s GLM-5.3 launch page GLM-5.3 leads on Terminal Bench 3.0, K3 leads on SWE-Marathon and Toolathlon, and they are close on DeepSWE. Pick K3 when you need image or video input; either works for text-only coding.

Can I run Kimi K3 locally, and what hardware does it need?

Yes. The weights are public on Hugging Face (moonshotai/Kimi-K3) in native MXFP4, about 1.56 TB in total, under the Kimi K3 License, so read the license before commercial use. vLLM’s official recipe needs at least 8 GB300 or 8 MI355X/MI350X GPUs in one node, and Unsloth’s smallest dynamic 1-bit quant (UD-IQ1_S) is still 594 GB and needs about 610 GB of combined RAM and VRAM. One community project streams the weights from four SSDs on a MacBook Pro at about 1 token/s, which is an experiment rather than a workable setup. Without that hardware, the official API or a per-token channel such as Wokey is more practical.

How should I set reasoning_effort on Kimi K3, and can I turn thinking off?

Thinking cannot be turned off. K3’s reasoning_effort accepts only low, high and max, with max as the default; when migrating from K2.x, replace the old thinking parameter with reasoning_effort. Artificial Analysis measured max at 44 for $2.00 per task and low at 30 for $1.15. Use max for coding and agents, low for simple tasks, and compare high with max on your own tasks before switching.

Why does Kimi K3 return unsupported image url?

Kimi’s vision input does not accept public image URLs, so upstream returns 400 unsupported image url even when the image opens fine in a browser. Download the image, convert it to a base64 data URL such as data:image/png;base64,..., and put that in image_url; video also goes in as base64. Wokey does not fetch remote images and re-encode them for you, so retrying the same URL fails again.

How much does the Kimi K3 API cost, and how much does it save versus the official API?

Wokey lists input at $0.9 per 1M tokens and output at $4.5 per 1M tokens; the official references are $3 and $15. That is 70% lower for input and 70% lower for output. Request-time pricing controls settlement, and the vendor pricing source is linked on this page.

Which API forms does Kimi K3 support, and how do I call it through Wokey?

Kimi K3 supports /v1/chat/completions (Chat Completions requests) and /v1/messages (Messages requests). Set the API base URL to https://api.wokey.ai, authenticate with a Wokey API key, and use the canonical model ID kimi-k3; you can start with /v1/chat/completions. The Send a request section includes a curl example for every supported API form.

How do I use Kimi K3 in Claude Code?

Claude Code calls Kimi K3 over /v1/messages. Set ANTHROPIC_BASE_URL="https://api.wokey.ai", ANTHROPIC_AUTH_TOKEN to your Wokey API key, and ANTHROPIC_MODEL="kimi-k3", then start claude. Claude Desktop also uses /v1/messages; the setup guides cover both step by step.

What are the context window and maximum output for Kimi K3?

The catalog context window is 1,048,576 tokens and the maximum output is 1,048,576 tokens. These are separate limits, and a request must also satisfy the upstream constraints in effect when it is processed.

How are cache reads and cache writes priced for Kimi K3?

Cache read: Wokey lists $0.09 per 1M tokens and the official reference is $0.3 per 1M tokens. Cache write: the vendor does not publish a separate meter, so Wokey also shows — as the public price.

How does calling Kimi K3 through Wokey differ from OpenRouter?

Wokey lists Kimi K3 at $0.9 input / $4.5 output per 1M tokens, against an official reference of $3 / $15. OpenRouter states it adds no markup on inference and charges its fee when you buy credits (5.5% by card). Wokey carries a curated set of models and has no per-provider routing parameters, so OpenRouter fits better if you need a wider catalog. The OpenRouter alternative page has the full comparison.

How do Kimi K3 and Kimi K2.8 Preview differ in price and context?

In the published price comparison, Kimi K3 lists input/output at $0.9 / $4.5 with a 1,048,576-token context window; Kimi K2.8 Preview lists $0.3 / $1.2 with a 1,048,576-token context window. This is a factual price-and-capacity comparison, not a quality ranking; also compare their supported API forms and capabilities.

Related models and documentation

Get API key · kimi-k3