Model names and marketing claims are not enough for a production decision. The page publishes prompts and the recording method first, then adds model IDs, latency, token usage, and CNY cost so you can reproduce the comparison for your own workload.
What we test for each model
Kimi
Test Chinese long-context work, document organization, and multi-turn context retention.
DeepSeek
Test complex reasoning, coding tasks, and request cost.
GLM
Test Chinese workflows, structured output, and tool calling.
Qwen
Test multilingual, coding, and tool-use performance together.
One test matrix for every model
Every model runs through the same public dimensions. We do not publish a ranking before the test is run.
| Dimension | What we record | Why it matters |
|---|---|---|
| Chinese and long context | Fact retention, context consistency, summary structure | Whether real business documents are usable |
| Coding and reasoning | Pass rate, explanation quality, edge cases | Whether fixes reduce manual rework |
| Structured output | JSON parsing, field completeness, tool calls | Whether automation can depend on it |
| Cost and latency | Time to first token, total time, token cost | Whether the production budget and UX are controllable |
Reproducible test cases
Each case is a real API request with a concrete input, acceptance checks, and an output summary. Copy the prompt and change only the model parameter to rerun it.
Case 01 · Support-ticket routing
A charged-but-missing API balance ticket tests whether a model can emit reliable JSON for a support queue.
你是客服工单路由器。只输出 JSON,不要 Markdown,不要解释。字段必须为 category、priority、needs_human、reply_language、reason。category 只能是 billing、account、api;priority 只能是 low、medium、high;reply_language 使用 BCP-47。
工单:我昨晚给 API 账户充值 200 元,银行卡已经扣款,但余额仍然是 0。我们今天要上线,能否尽快处理?Case 02 · TypeScript production bug
A tiny but common field-name bug tests whether the model can return the smallest runnable fix.
只输出修复后的 TypeScript 函数,不要解释,不要 Markdown。保留函数名和类型定义。
type LineItem = { price: number; qty: number };
export function subtotal(items: LineItem[]): number {
return items.reduce((sum, item) => sum + item.price * item.quantity, 0);
}Case 03 · Order analytics SQL
A common paid-order report tests filtering, aggregation, grouping, and sorting in one small query.
只输出一条 SQLite SQL,不要解释,不要 Markdown。统计过去 30 天内 status 为 paid 的订单,按 user_id 汇总 amount,结果列必须叫 user_id 和 total_amount,并按 total_amount 从高到低排序。
表结构:orders(id INTEGER, user_id INTEGER, status TEXT, amount REAL, created_at TEXT)Case 01 · Support-ticket routing
Results are single streaming requests recorded on 2026-10-01 UTC with max_tokens=512. TTFT, total latency, and tokens come from the request record; cost is calculated from the active model catalog and usage.
| Model | Checks | Output summary | TTFT | Total | Input / output tokens | Usage cost |
|---|---|---|---|---|---|---|
| DeepSeek V4 Pro | Passed4/4 checks passed | Valid JSON; billing / high / true / zh-CN | 11.79s | 12.92s | 118 / 134 | ¥0.004680 |
| GLM-5.2 | Passed4/4 checks passed | Valid JSON; billing / high / true / zh-CN | 8.24s | 8.75s | 99 / 64 | ¥0.001938 |
| Qwen3.7 Max | Passed4/4 checks passed | Valid JSON; billing / high / true / zh-CN | 9.37s | 10.22s | 106 / 627 | ¥0.019075 |
| Kimi K3 | Passed4/4 checks passed | Valid JSON; billing / high / true / zh-CN | 17.12s | 17.79s | 174 / 209 | ¥0.014628 |
Case 02 · TypeScript production bug
Results are single streaming requests recorded on 2026-10-01 UTC with max_tokens=512. TTFT, total latency, and tokens come from the request record; cost is calculated from the active model catalog and usage.
| Model | Checks | Output summary | TTFT | Total | Input / output tokens | Usage cost |
|---|---|---|---|---|---|---|
| DeepSeek V4 Pro | Passed3/3 checks passed | Replaced item.quantity with item.qty | 5.14s | 6.40s | 93 / 98 | ¥0.003483 |
| GLM-5.2 | Passed3/3 checks passed | Replaced item.quantity with item.qty | 5.64s | 6.03s | 69 / 54 | ¥0.001548 |
| Qwen3.7 Max | Passed3/3 checks passed | Replaced item.quantity with item.qty | 3.01s | 3.73s | 80 / 126 | ¥0.004397 |
| Kimi K3 | Passed3/3 checks passed | Replaced item.quantity with item.qty | 16.01s | 16.01s | 153 / 179 | ¥0.012576 |
Case 03 · Order analytics SQL
Results are single streaming requests recorded on 2026-10-01 UTC with max_tokens=512. TTFT, total latency, and tokens come from the request record; cost is calculated from the active model catalog and usage.
| Model | Checks | Output summary | TTFT | Total | Input / output tokens | Usage cost |
|---|---|---|---|---|---|---|
| DeepSeek V4 Pro | Passed5/5 checks passed | Complete SQLite query with a 30-day filter | 9.91s | 10.92s | 98 / 42 | ¥0.002016 |
| GLM-5.2 | Passed5/5 checks passed | Complete SQLite query with a 30-day filter | 19.73s | 25.99s | 77 / 46 | ¥0.001428 |
| Qwen3.7 Max | Passed5/5 checks passed | Complete SQLite query with a 30-day filter | 9.84s | 10.46s | 85 / 752 | ¥0.022474 |
| Kimi K3 | Rerun neededIncomplete | Empty output; the 256-token cap was exhausted, so rerun with a higher limit | — | 52.78s | 159 / 256 | ¥0.017268 |
How to read the results
Results are single streaming requests recorded on 2026-10-01 UTC with max_tokens=512. TTFT, total latency, and tokens come from the request record; cost is calculated from the active model catalog and usage.
| Field | Meaning | Decision rule | Display |
|---|---|---|---|
| Passed | All acceptance checks pass | Automated checks plus a manual spot check | Green Passed |
| Incomplete | Output is capped or empty | Do not call it a capability failure | Rerun needed |
| Cost | Input plus output tokens | Catalog price calculation | CNY |
| Latency | First token and full response | Single-request observation | TTFT / Total |
Results vary by model version, parameters, and test set. We publish conclusions only after a run can be reproduced.
How we run the tests
Keep the prompt, temperature, maximum output tokens, and protocol fixed; change only the model ID.
Record time to first token, total latency, input/output tokens, and request ID; calculate CNY cost from the active catalog rates.
Store qualitative scores separately from raw outputs so one score does not hide the full result.
Show the test date and model version; rerun after upstream changes instead of treating old results as permanent.
FAQ
Which model is best for coding?
Use the same code-fix cases and compare pass rate, latency, and cost instead of relying on the model name alone.
Can these models use one API?
KimiSeek supports OpenAI, Anthropic Messages, Responses, and Gemini-compatible interfaces. Change the model parameter to switch enabled models.
How is comparison pricing calculated?
Requests are settled from input, cached-input, and output tokens at the active CNY selling price.
Use one codebase and verify the difference
Copy the test cases, then choose a model based on your actual prompts, latency target, and budget.
testing