Model Pricing and Billing

Review multipliers, price conversion rules, standardized billing units, and model costs.

1. Billing Rules

Core billing unit: The platform bills per 100 million Tokens. A 1x multiplier = 40 yuan per 100 million Tokens, equivalent to 0.4 yuan per 1 million Tokens. The "multiplier" and "price per 1 million Tokens" columns in the table below use the same conversion basis, so you can read the prices directly from the table.

Billing ItemStandardDescription
1x multiplier40 yuan / 100 million TokensEquivalent to 0.4 yuan per 1 million Tokens and used as the base conversion standard for the price table.
2x multiplier80 yuan / 100 million TokensPrices scale linearly from the base multiplier, so the price per 1 million Tokens also doubles.
5x multiplier200 yuan / 100 million TokensUsed for models with higher costs or resource requirements.
10x multiplier400 yuan / 100 million TokensHigher-cost models use higher multipliers. See the model price table for details.

One Key works with every model. Multipliers only represent differences in model costs. You do not need a separate Key for each model.

Actual charges are based on system records. If a model receives a temporary subsidy, promotional discount, or multiplier adjustment, the price table on this page will be updated accordingly.

2. Cache Reads and Cache Writes

What is a cache hit? When consecutive requests contain a large amount of identical context, the system can reuse previously processed content. The cached portion does not need to be recomputed at the full input rate. Common examples include long conversations, consecutive follow-up questions, code completion, agent workflows, and multi-turn tool calls.

Input tokens are billed in three separate parts. Reading from the cache (a hit) is heavily discounted. Placing new content into the cache — a cache write, also called a cache miss — is billed at its own rate, which on some models is higher than standard input because the upstream provider charges a premium for it. Ordinary uncached input sits between the two. Each part appears as its own column in your usage records, so the three always add up to the input you actually sent.

ItemBilling MethodDescription
Cache read (cache hit)Usually billed at a 0.1x multiplierThe cached portion costs about 10% of standard input, reducing costs for long-context use cases.
Cache write (cache miss)Billed at its own multiplier — often above standard inputWriting context into the cache so later turns can hit it. On the Claude models this is about 1.25x standard input (1.5x on Claude Haiku 4.5); on some other models it is cheaper than standard input. The exact per-model rate is in the price table on the pricing page.
Uncached inputBilled at the model's corresponding multiplierNew content that is neither read from nor written to the cache is billed at the normal input rate.
Special models / periods of instabilityMay be billed at a 0.2x multiplier or temporarily bypass cachingFor example, cache billing may be adjusted based on actual capabilities for high-cost models or when an upstream provider is unstable.

Example: In a long conversation, the first turn that establishes the context pays the cache write rate on that context. Every later turn that reuses it pays the much lower cache read rate. New questions and newly generated output are still billed at the model's corresponding multiplier. This is why a conversation's first request can cost more than the ones that follow it.

Whether a request results in a cache hit depends on the request content, model capabilities, upstream caching policy, and current system availability. Final charges are based on billing records.

3. Model List and Multiplier Price Table (All Models Available)

How to read the table: Start with the model name, then check the multiplier, and finally review the "price per 1 million Tokens." A lower multiplier means a lower cost per Token. Multipliers and prices use the same conversion basis.

FieldDescription
Model nameThe model name entered in a client, workflow, or API request. We recommend copying capitalization exactly as shown in the table.
MultiplierRepresents differences in model costs. Lower multipliers are better for frequent everyday use, while higher multipliers are better suited to high-value tasks.
Price per 1 million TokensAlready converted using the multiplier, so you can use this column directly to estimate usage costs.
Temporary subsidy / promotional priceWhen a temporary subsidy or limited-time promotion applies, the table will show the latest multiplier. Pricing may return to its previous rate or change again when the promotion ends.
ProviderDisplay NameModel IDMultiplierPriceContextImage
AI API ProxyClaude Sonnet 4.6 (Optimized)claude-sonnet-4-60.8x¥0.32 / M Tokens1MSupported
OpenAIGPT-5.4gpt-5.43x¥1.2 / M Tokens1MSupported
OpenAIGPT-5.5gpt-5.56x¥2.4 / M Tokens258KSupported
OpenAI / CodexGPT-5.3 Codex Sparkgpt-5.3-codex-spark1x¥0.4 / M Tokens128KNot supported
OpenAIGPT Image 2gpt-image-2Per image¥0.05–0.10 / image1–2K outputGeneration
OpenAIGPT Image 2 4Kgpt-image-2-4kPer image¥0.5–1 / image4K outputGeneration
OpenAIGPT-5.6 Solgpt-5.6-sol6x¥2.4 / M Tokens258KSupported
OpenAIGPT-5.6 Terragpt-5.6-terra6x¥2.4 / M Tokens258KSupported
OpenAIGPT-5.6 Lunagpt-5.6-luna6x¥2.4 / M Tokens258KSupported
xAI / GrokGrok 4.5grok-4.53x¥1.2 / M Tokens500KSupported
AnthropicClaude Haiku 4.5claude-haiku-4-5-202510011x¥0.4 / M Tokens256KSupported
AnthropicClaude Sonnet 5claude-sonnet-512x¥4.8 / M Tokens1MSupported
AnthropicClaude Fable 5claude-fable-530x¥12 / M Tokens1MSupported
AnthropicClaude Opus 4.6claude-opus-4-620x¥8 / M Tokens1MSupported
AnthropicClaude Opus 4.7claude-opus-4-720x¥8 / M Tokens1MSupported
AnthropicClaude Opus 4.8claude-opus-4-820x¥8 / M Tokens1MSupported
Alibaba Cloud / QwenQwen 3.6 Plusqwen3.6-plus3x¥1.2 / M Tokens1MSupported
Alibaba Cloud / QwenQwen 3.7 Plusqwen3.7-plus4x¥1.6 / M Tokens1MSupported
Alibaba Cloud / QwenQwen 3.7 Maxqwen3.7-max8x¥3.2 / M Tokens1MNot supported
Alibaba Cloud / QwenQwen 3.8 Maxqwen3.8-max10x¥4 / M Tokens1MSupported
Meituan / LongCatLongCat 2.0LongCat-2.01x¥0.4 / M Tokens1MNot supported
Tencent HunyuanHunyuan 3hy31x¥0.4 / M Tokens256KNot supported
MiniMaxMiniMax M3MiniMax-M31x¥0.4 / M Tokens1MSupported
MiniMaxImage 01image-01Per image¥0.2 / imageGeneration
MiniMaxImage 01 Liveimage-01-livePer second¥2 / secondGeneration
StepFunStep 3.7 Flashstep-3.7-flash1x¥0.4 / M Tokens256KSupported
ByteDance / DoubaoDoubao Seed 2.0 Codedoubao-seed-2.0-code3x¥1.2 / M Tokens200KSupported
ByteDance / DoubaoDoubao Seed 2.0 Prodoubao-seed-2.0-pro3x¥1.2 / M Tokens128KSupported
Xiaomi / MiMoMiMo V2.5 Promimo-v2.5-pro3x¥1.2 / M Tokens1MNot supported
Xiaomi / MiMoMiMo V2.5mimo-v2.52x¥0.8 / M Tokens1MSupported
DeepSeekDeepSeek V4 Prodeepseek-v4-pro7.5x¥3 / M Tokens1MNot supported
DeepSeekDeepSeek V4 Flashdeepseek-v4-flash2.5x¥1 / M Tokens1MNot supported
Moonshot AI / KimiKimi K3kimi-k325x¥10 / M Tokens1MSupported
Zhipu / Z.aiGLM 5.1glm-5.15x¥2 / M Tokens256KSupported
Zhipu / Z.aiGLM 5.2glm-5.28x¥3.2 / M Tokens1MSupported

API endpoint: Use https://jzstoken.com/api/v1 whenever possible. If a client is incompatible with an address that includes /v1, try the base URL without /v1 instead. Different API gateway platforms, clients, and plugins may have different Base URL requirements. After switching models, we recommend testing with a small request first.

Price Table Change Monitoring Skill Installation Guide

The price table change monitoring Skill has been updated to 1.0.4. The new version fixes issues reading large tables in public Feishu documents and prioritizes the actual model price table, preventing explanatory tables from being misidentified as price tables.

What It DoesDescription
Automatically detects changesMonitors model, multiplier, and price tables in public Feishu documents to identify added or removed models, multiplier changes, and price changes per 1 million Tokens.
Automatically formats notificationsOutputs a Markdown notification when changes occur, ready to send to WeChat groups, Telegram, internal company groups, or bot messages.
Designed for customer monitoringResellers, team administrators, and customer-community managers can use it to track price changes without manually opening the document and comparing tables every day.
Stays quiet when nothing changesOutputs NO_REPLY when no changes are detected, making it suitable for OpenClaw, cron, or other automation tasks.

One-command installation

clawhub install feishu-public-table-monitor

Update an existing installation to the latest version

clawhub update feishu-public-table-monitor --force

Recommended update command for an OpenClaw / Lobster workspace

clawhub --workdir ~/.openclaw/workspace update feishu-public-table-monitor --force

Example command for monitoring this price table

python3 ~/.openclaw/workspace/skills/feishu-public-table-monitor/scripts/monitor_feishu_price_table.py \
  "https://ycnfsn8b2u4z.feishu.cn/wiki/XFBBweJGoiboM7kKdjKca4Fvnvb" \
  --section-title "三、模型列表与倍率价格表(所有模型可用)"

How to use it: The first run creates a baseline. Run it on a schedule after that. An output of NO_REPLY means the price table has not changed. When changes are detected, it automatically generates a notification covering added models, removed models, multiplier changes, and price changes.

Scope: This Skill is designed for publicly accessible price tables in Feishu documents. For private documents or pages that require authentication, verify access permissions and data security requirements before connecting them.