https://new.weeanno.shop/v1Use this first for most clients. If your client rejects it, switch to the fallback.
https://new.weeanno.shopDifferent API routers validate Base URLs differently, so test the address when you switch models or clients.
Model multipliers, cache hits, and first API-key purchase
The RMB site uses a CNY 100 per 100M token benchmark, with CNY 10, 50, 100, 298, and 488 packages.
Existing usersAlready have an API key? Top it up directlyTop up, look up orders, and confirm delivery here; 500M tokens cost CNY 488.Go to top-uphttps://new.weeanno.shop/v1Use this first for most clients. If your client rejects it, switch to the fallback.
https://new.weeanno.shopDifferent API routers validate Base URLs differently, so test the address when you switch models or clients.
When part of a request matches previously processed content, the system can reuse that computation instead of charging the full input cost again.
Cache reuse can improve response speed and reduce usage cost, especially in multi-turn chats, repeated workflow calls, and code-completion loops.
If you repeatedly include a large shared context across a long conversation, the matched repeated portion can be charged at a lower multiplier.
Matched cached input is billed at only 10% of the regular input cost.
Follow-up chats, workflow calls, and code-completion loops are more likely to reuse repeated context.
When upstream costs move sharply, cache multipliers may change and can be paused in extreme cases.
Note: cache hits depend on whether the request content qualifies for system cache reuse.
Compare multipliers, converted per-token cost, context size, and vision support before you buy.
A practical reference for understanding model cost before purchase.
| Provider | Model | MultiplierExample: 2x means 100M tokens deducts 200M quota | Recommended context windowAssumes max_tokens=8192 | Vision support |
|---|---|---|---|---|
| China aggregate route(High concurrency) | claude-sonnet-4-6Fast response | 0.5x | 1M | ✅ |
| OpenAI | gpt-5.6-sol | 6x | 258K | ✅ |
| OpenAI | gpt-5.6-terra | 1x | 258K | ✅ |
| OpenAI | gpt-5.4 | 2x | 1M | ✅ |
| OpenAI | gpt-5.5 | 4x | 258K | ✅ |
| OpenAI | gpt-5.3-codex-spark | 1x | 128K | |
| OpenAI | gpt-image-2 | per image | 1–2K 出图 | |
| OpenAI | gpt-image-2-4k | per image | 4K 出图 | |
| 通义千问 | qwen3.6-plus | 3x | 1M | ✅ |
| 通义千问 | qwen3.7-plus | 4x | 1M | ✅ |
| 通义千问 | qwen3.7-max | 10x | 1M | |
| 通义千问 | qwen3.8-max | 20x | 1M | ✅ |
| 美团 LongCat | LongCat-2.0Agentic Coding | 1x | 1M | |
| 腾讯混元 | hy3 | 1x | 256K | |
| MiniMax | MiniMax-M3 | 1x | 1M | ✅ |
| MiniMax | image-01 | per image | ||
| MiniMax | image-01-live | per second | ||
| StepFun | step-3.7-flashFast response | 1x | 256K | ✅ |
| ByteDance | doubao-seed-2.1-turbo | 3x | 128K | ✅ |
| Xiaomi | mimo-v2.5-pro | 4x | 1M | |
| Xiaomi | mimo-v2.5 | 2x | 1M | ✅ |
| DeepSeek | deepseek-v4-pro | 7.5x | 1M | |
| DeepSeek | deepseek-v4-flash | 2.5x | 1M | |
| Kimi | kimi-k3 | 25x | 1M | ✅ |
| Zhipu AI | glm-5.1 | 5x | 256K | ✅ |
| Zhipu AI | glm-5.2 | 8x | 1M | ✅ |
| xAI / Grok | grok-4.5 | 5x | 500K | ✅ |
| Anthropic | claude-haiku-4-5-20251001 | 1x | 256K | ✅ |
| Anthropic | claude-sonnet-5高能力主力 | 10x | 1M | ✅ |
| Anthropic | claude-fable-5顶级任务 | 40x | 1M | ✅ |
| Anthropic | claude-opus-4-6 | 15x | 1M | ✅ |
| Anthropic | claude-opus-4-7 | 15x | 1M | ✅ |
| Anthropic | claude-opus-4-8 | 15x | 1M | ✅ |
| Anthropic | claude-opus-5 | 15x | 1M | ✅ |
gemini-3.1-pro | 10x | 1M | ✅ | |
gemini-3.5-flash | 6x | 1M | ✅ |
Prices above are reference conversions based on current multipliers. Image models are billed per image or second. Actual usage follows the platform multiplier and request consumption; ✅ means vision is supported.
This page explains multipliers, cache hits, and model cost. New users can buy an API key; existing users can top up an existing key.