MODEL COMPARISON & BENCHMARKS

Kimi, DeepSeek, GLM, or Qwen?

Run the same prompts and track the same metrics across Chinese AI models for writing, coding, long context, and API use.

The first test cases are public; results will be updated with model versions.

Model names and marketing claims are not enough for a production decision. The page publishes prompts and the recording method first, then adds model IDs, latency, token usage, and CNY cost so you can reproduce the comparison for your own workload.

01 / FOCUS

What we test for each model

K

Kimi

Test Chinese long-context work, document organization, and multi-turn context retention.

Chinese long contextContext retentionMulti-turn chat
D

DeepSeek

Test complex reasoning, coding tasks, and request cost.

ReasoningCodingToken cost
G

GLM

Test Chinese workflows, structured output, and tool calling.

Chinese tasksJSON outputTool calling
Q

Qwen

Test multilingual, coding, and tool-use performance together.

MultilingualCode and toolsContext handling
02 / MATRIX

One test matrix for every model

Every model runs through the same public dimensions. We do not publish a ranking before the test is run.

DimensionWhat we recordWhy it matters
Chinese and long contextFact retention, context consistency, summary structureWhether real business documents are usable
Coding and reasoningPass rate, explanation quality, edge casesWhether fixes reduce manual rework
Structured outputJSON parsing, field completeness, tool callsWhether automation can depend on it
Cost and latencyTime to first token, total time, token costWhether the production budget and UX are controllable
03 / REPRODUCIBLE TESTS

Reproducible test cases

Each case is a real API request with a concrete input, acceptance checks, and an output summary. Copy the prompt and change only the model parameter to rerun it.

01

Case 01 · Support-ticket routing

A charged-but-missing API balance ticket tests whether a model can emit reliable JSON for a support queue.

Prompt你是客服工单路由器。只输出 JSON,不要 Markdown,不要解释。字段必须为 category、priority、needs_human、reply_language、reason。category 只能是 billing、account、api;priority 只能是 low、medium、high;reply_language 使用 BCP-47。 工单:我昨晚给 API 账户充值 200 元,银行卡已经扣款,但余额仍然是 0。我们今天要上线,能否尽快处理?
Acceptance checkscategory=billingpriority=highneeds_human=truereply_language=zh-CNOutput parses with JSON.parse
02

Case 02 · TypeScript production bug

A tiny but common field-name bug tests whether the model can return the smallest runnable fix.

Prompt只输出修复后的 TypeScript 函数,不要解释,不要 Markdown。保留函数名和类型定义。 type LineItem = { price: number; qty: number }; export function subtotal(items: LineItem[]): number { return items.reduce((sum, item) => sum + item.price * item.quantity, 0); }
Acceptance checksUse item.qtyRemove item.quantityKeep subtotal as the function nameOutput passes type checking
03

Case 03 · Order analytics SQL

A common paid-order report tests filtering, aggregation, grouping, and sorting in one small query.

Prompt只输出一条 SQLite SQL,不要解释,不要 Markdown。统计过去 30 天内 status 为 paid 的订单,按 user_id 汇总 amount,结果列必须叫 user_id 和 total_amount,并按 total_amount 从高到低排序。 表结构:orders(id INTEGER, user_id INTEGER, status TEXT, amount REAL, created_at TEXT)
Acceptance checksSUM(amount) aggregates amountFilter status='paid'Group by user_idSort by total_amount descendingKeep only the last 30 days

Case 01 · Support-ticket routing

Results are single streaming requests recorded on 2026-10-01 UTC with max_tokens=512. TTFT, total latency, and tokens come from the request record; cost is calculated from the active model catalog and usage.

ModelChecksOutput summaryTTFTTotalInput / output tokensUsage cost
DeepSeek V4 ProPassed4/4 checks passedValid JSON; billing / high / true / zh-CN11.79s12.92s118 / 134¥0.004680
GLM-5.2Passed4/4 checks passedValid JSON; billing / high / true / zh-CN8.24s8.75s99 / 64¥0.001938
Qwen3.7 MaxPassed4/4 checks passedValid JSON; billing / high / true / zh-CN9.37s10.22s106 / 627¥0.019075
Kimi K3Passed4/4 checks passedValid JSON; billing / high / true / zh-CN17.12s17.79s174 / 209¥0.014628

Case 02 · TypeScript production bug

Results are single streaming requests recorded on 2026-10-01 UTC with max_tokens=512. TTFT, total latency, and tokens come from the request record; cost is calculated from the active model catalog and usage.

ModelChecksOutput summaryTTFTTotalInput / output tokensUsage cost
DeepSeek V4 ProPassed3/3 checks passedReplaced item.quantity with item.qty5.14s6.40s93 / 98¥0.003483
GLM-5.2Passed3/3 checks passedReplaced item.quantity with item.qty5.64s6.03s69 / 54¥0.001548
Qwen3.7 MaxPassed3/3 checks passedReplaced item.quantity with item.qty3.01s3.73s80 / 126¥0.004397
Kimi K3Passed3/3 checks passedReplaced item.quantity with item.qty16.01s16.01s153 / 179¥0.012576

Case 03 · Order analytics SQL

Results are single streaming requests recorded on 2026-10-01 UTC with max_tokens=512. TTFT, total latency, and tokens come from the request record; cost is calculated from the active model catalog and usage.

ModelChecksOutput summaryTTFTTotalInput / output tokensUsage cost
DeepSeek V4 ProPassed5/5 checks passedComplete SQLite query with a 30-day filter9.91s10.92s98 / 42¥0.002016
GLM-5.2Passed5/5 checks passedComplete SQLite query with a 30-day filter19.73s25.99s77 / 46¥0.001428
Qwen3.7 MaxPassed5/5 checks passedComplete SQLite query with a 30-day filter9.84s10.46s85 / 752¥0.022474
Kimi K3Rerun neededIncompleteEmpty output; the 256-token cap was exhausted, so rerun with a higher limit—52.78s159 / 256¥0.017268

How to read the results

Results are single streaming requests recorded on 2026-10-01 UTC with max_tokens=512. TTFT, total latency, and tokens come from the request record; cost is calculated from the active model catalog and usage.

FieldMeaningDecision ruleDisplay
PassedAll acceptance checks passAutomated checks plus a manual spot checkGreen Passed
IncompleteOutput is capped or emptyDo not call it a capability failureRerun needed
CostInput plus output tokensCatalog price calculationCNY
LatencyFirst token and full responseSingle-request observationTTFT / Total

Results vary by model version, parameters, and test set. We publish conclusions only after a run can be reproduced.

04 / METHOD

How we run the tests

Keep the prompt, temperature, maximum output tokens, and protocol fixed; change only the model ID.

Record time to first token, total latency, input/output tokens, and request ID; calculate CNY cost from the active catalog rates.

Store qualitative scores separately from raw outputs so one score does not hide the full result.

Show the test date and model version; rerun after upstream changes instead of treating old results as permanent.

05 / FAQ

FAQ

Which model is best for coding?

Use the same code-fix cases and compare pass rate, latency, and cost instead of relying on the model name alone.

Can these models use one API?

KimiSeek supports OpenAI, Anthropic Messages, Responses, and Gemini-compatible interfaces. Change the model parameter to switch enabled models.

How is comparison pricing calculated?

Requests are settled from input, cached-input, and output tokens at the active CNY selling price.

KIMISEEK

Use one codebase and verify the difference

Copy the test cases, then choose a model based on your actual prompts, latency target, and budget.

Start
testing
Chinese AI Models Compared: Kimi, DeepSeek, GLM, and Qwen | KimiSeek