看完总榜(Agent 场景算账裸价排行)再看各家内部:同一个厂里也有便宜和贵的,差距有时比厂商之间还大。这篇把 11 家厂商 109 款模型的内部价格全排列出来(按 Agent 场景模拟花费),想用谁家的先看它家内部怎么选。

一张总表:11 家的最贵与最便宜

厂商款数最贵(模拟花费)最便宜厂内价差
阿里10qwen3.8-max-prime ¥60qwen-turbo ¥0.6987 倍
百川智能12Baichuan4(2024 定价)¥1,020Baichuan4-Air ¥10102 倍
百度5ernie-5.0 ¥64.8ernie-4.5-turbo ¥3.2420 倍
字节跳动15doubao-seed-evolving ¥22.8doubao-seed-1.6-flash ¥0.7232 倍
DeepSeek6deepseek-v4-pro-peak ¥17.1v4-flash-offpeak ¥2.856 倍
MiniMax6m3-priority-gt512k ¥22.7minimax-m2.7 ¥7.563 倍
Moonshot4kimi-k3 ¥58kimi-k2.6 ¥21.82.7 倍
阶跃星辰9step-1o-audio(语音)¥82step-3.5-flash ¥2.3834 倍
腾讯18tencent-yuanqi(平台价)¥1,020hy-mt2-lite(翻译)¥3.24315 倍
小米2mimo-v2.5-pro ¥4.43mimo-v2.5 ¥1.582.8 倍
智谱 AI22GLM-4-AirX ¥102GLM-4.6V-FlashX ¥0.72142 倍

每家内部怎么选(快评)

阿里(10 款):档位清晰,turbo 是白送价

qwen3.8-max-prime(¥60)→ qwen3.8-max(¥30)→ qwen3.7-max 5 折(¥15)→ qwen3-max(¥6.75)→ qwen-max(¥6.48)→ qwq-plus(¥3.84)→ qwen-plus(¥1.92)→ qwen-turbo(¥0.69)。每档价格差一半,能力档位分明,是最容易按预算选型的家族。qwen3.7-max 现在挂 5 折,性价比窗口期。

字节跳动(15 款):全家桶策略,flash 系列极其便宜

seed-2.1-pro(¥22.8)领衔,但真正的重点是 2.0-mini(¥0.96)、1.6-lite(¥0.96)、1.6-flash(¥0.72)——高频调用场景直接无脑 flash。还有专门的 translation/character/voice 分支,按需取用。

智谱(22 款):库存最庞大,老模型清库存价

GLM-5.3/5.2(¥31.6)是新旗舰,GLM-4 系列整体还在卖但属于上一代;GLM-4.7-FlashX 一类 flash 档 ¥0.72~2,22 款里一半是体验价。选型建议直接锁定 5.3/5.1 和 flash 档,中间档意义不大。

DeepSeek(6 款):简单粗暴,峰谷价是特色

v4-pro-peak(¥17.1)/ v4-pro-offpeak(¥8.55)/ v4-flash-peak(¥5.7)/ v4-flash-offpeak(¥2.85),pro 和 flash 各配峰谷,缓存命中价只要原价 3%,批量任务挪谷时直接半价。

腾讯(18 款):混元分支多,认准 hunyuan 主线

hunyuan-turbos-vision(¥31.8)是视觉旗舰,hunyuan-a13b(¥5.4)是轻量主力;translation 系列单独计价。元器平台统一价(¥1,020)不是 API 价,忽略即可。

Moonshot(4 款):k3 定高价,k2.6 是性价比位

kimi-k3(¥58)对标旗舰,k2.6(¥21.8)和 k2.7-code(¥23.6)才是主力走量款。长输出型选手(输出 27~100),算成本时输出占比要放大

阶跃星辰(9 款):语音模型是价格天花板

step-1o-audio(¥82)等语音模型霸占高价段,文本主力 step-3.7-flash(¥5.4)/ step-3.5-flash(¥2.38)其实很便宜。别被语音榜吓到,文本线是性价比路线

MiniMax / 小米 / 百度 / 百川

MiniMax 全系挂 5 折(m3 系列 ¥7.56~22.7),Context 长度分档(>512K 加价);小米 mimo-v2.5(¥1.58)配 2% 缓存比,重上下文场景黑马;百度 ernie 5.0(¥64.8)偏贵,4.5-turbo(¥3.24)是实用位;百川老刊例价偏historic,4-Air(¥10)以下再考虑。

11 家厂商内部榜长图

阿里内部价格分榜
阿里(10 款)
DeepSeek 内部价格分榜
DeepSeek(6 款)
字节跳动内部价格分榜
字节跳动(15 款)
智谱内部价格分榜
智谱 AI(22 款)
腾讯内部价格分榜
腾讯(18 款)
百度内部价格分榜
百度(5 款)
Moonshot 内部价格分榜
Moonshot(4 款)
MiniMax 内部价格分榜
MiniMax(6 款)
阶跃星辰内部价格分榜
阶跃星辰(9 款)
百川内部价格分榜
百川智能(12 款)
小米内部价格分榜
小米(2 款)

选型心法

  • 先定能力档,再看厂内价:绝大多数任务用不到旗舰,各厂的 flash/turbo/lite 档就够,成本差 10~30 倍
  • 重上下文看缓存比:DeepSeek、小米的缓存价是原价 1~3%,Agent 场景首选验证
  • 折扣窗口要抓住:qwen3.7-max 5 折、MiniMax 全系 5 折都是限时的,快照价未必一直在

同系列:第 1 篇 Agent 场景算账总榜 · 第 2 篇裸价三榜

免责声明:模拟花费口径为输入 1000 万 tokens(缓存命中 90%)+ 输出 20 万 tokens;价格为 2026 年 8 月官网刊例,手工整理可能不全或有错漏,波动频繁仅供参考。

Beyond the global boards (agent workloads, sticker prices), the inside of each vendor matters: in-vendor price gaps can exceed gaps between vendors. All 109 models from 11 Chinese vendors, ranked internally by simulated agent cost (10M input tokens @90% cache hit + 200K output).

One-table summary

VendorModelsPriciestCheapestGap
Alibaba10qwen3.8-max-prime ¥60qwen-turbo ¥0.6987×
Zhipu22GLM-4-AirX ¥102GLM-4.6V-FlashX ¥0.72142×
ByteDance15doubao-seed-evolving ¥22.8doubao-seed-1.6-flash ¥0.7232×
DeepSeek6v4-pro-peak ¥17.1v4-flash-offpeak ¥2.85
Moonshot4kimi-k3 ¥58kimi-k2.6 ¥21.82.7×
Tencent18yuanqi platform ¥1,020hy-mt2-lite ¥3.24315×

Per-vendor quick takes

  • Alibaba: cleanest ladder — each tier costs ~half the one above (60 → 30 → 15 → 6.75 → 3.84 → 1.92 → 0.69). qwen3.7-max currently at 50% off.
  • ByteDance: the flash family (¥0.72–0.96) makes high-frequency use essentially free; pro tier at ¥22.8.
  • Zhipu: 22 models, half are near-experience pricing; lock onto GLM-5.3/5.1 or the flash tier.
  • DeepSeek: pro/flash × peak/off-peak grid; cache at 3% of input rate; off-peak halves batch costs.
  • Moonshot: k3 prices like a flagship (¥58); k2.6/k2.7-code (¥21.8/23.6) are the volume picks. Long outputs — weight output cost heavily.
  • Step: voice models top the board (¥82) but the text line (step-3.5-flash ¥2.38) is a bargain.
  • MiniMax: everything 50% off, context-length tiers above 512K.
  • Xiaomi MiMo: tiny prices plus ~1% cache ratio — a dark horse for long-context agents.

Selection heuristics

  • Pick capability tier first; flash/turbo/lite tiers usually suffice and cost 10–30× less
  • For long-context agents, verify the cache ratio before committing
  • Discount windows (Qwen 50%, MiniMax 50%) are time-limited — snapshot prices won't last

Series: part 1 · part 2.

Disclaimer: simulated cost = 10M input tokens (90% cache hit) + 200K output; official list prices August 2026, hand-collected, volatile.
返回文章列表