Flex mode, batch and other discounts: what actually applies to video
Updated 2026-10-02
"Is there a discount if I can wait?" is a reasonable question for any API bill, and a frequent one for video, where jobs are slow and spend adds up. The honest answer on VideoRouter is specific: its documented discount mechanisms apply to chat completions, and the video route is priced per second at the host's rate plus the platform fee. This page lays out what the documentation says, what it does not say, and how to ask the same question of any provider without being misled.
What flex mode is, per the documentation
VideoRouter's Flex Mode page describes one extra field on a normal POST /v1/chat/completions call: service_tier: "flex". You trade a slower response for a discount, on the same endpoint with the same request and response shape. The call blocks, so you give it a long client timeout, and the response reports which tier actually served it. The documentation says that when the underlying model has a native flex tier, the request is forwarded with a real flex tier and billed at the discounted price the upstream provider itself charges. If the model has no native flex tier, the request still succeeds at the standard price.
Note what that means. Flex mode is a pass-through of a provider-side feature, not a markdown invented by the gateway, and the documentation names the providers whose flex tiers are involved. It is documented for chat completions. It is not described as applying to video or image generation, so do not plan a video budget around it.
What the batch and async options cover
The same documentation set describes a file-based Batch API for bulk work with a longer completion window, and an Async Jobs endpoint for single requests that can wait longer, delivered by polling or a webhook. Their listed endpoints are chat completions. The documentation does not list video generation among them, so the correct assumption for video is that no batch discount exists. If that changes, it will appear in the Batch API and Flex Mode pages, and those pages, not this article, are the authority.
The trade-off, stated plainly
| Option | You give up | You get |
|---|---|---|
| Flex tier (chat) | Latency; you hold a connection open for minutes | The provider's own discounted rate, if the model has one |
| Batch (chat) | Immediacy; results arrive within a window | Bulk submission and a lower rate where the upstream offers it |
| Async job (chat) | Synchronous response | Poll or webhook delivery for long waits |
| Standard (video, per second) | Nothing | Predictable per-second billing at creation, free polling |
The pattern is the usual one: you can buy a lower price with flexibility on time, but only where the upstream provider sells it. A gateway cannot invent a discount it does not receive without losing money, so be wary of any provider that advertises a batch or flex discount on a modality where its own supplier has no such tier.
Where video savings really come from
If flex does not apply, the levers are the ones that every per-second API shares, and they are bigger than most discounts anyway:
- Host choice. Prices for the same model differ between hosts. Unpinned requests go to the cheapest healthy host; the live table below shows the spread.
- Tier choice. Draft on the cheaper variant or resolution and promote only approved shots.
- Requested duration. Billing uses requested seconds; do not ask for 10 when you keep 6.
- Retries. One fewer rejected generation per accepted clip is a larger saving than most percentage discounts.
| Model | Cheapest host | Priciest host | Cheapest is | Hosts |
|---|---|---|---|---|
| bytedance/seedance-2.5 (480p) | OpenSand $0.0525 / second | Fal-US $0.2646 / second | 80% lower | 9 |
| bytedance/seedance-2.0 (2160p) | MachGen $0.59 / second | Fal $1.5552 / second | 62% lower | 9 |
| alibaba/wan-3.0 (480p) | Replicate $0.025 / second | Alibaba $0.05 / second | 50% lower | 10 |
| google/gemini-omni-flash | Google $0.1 / second | Fal $0.13 / second | 23% lower | 4 |
| minimax/h3 (768p) | MachGen $0.04 / second | WaveSpeedAI-resell $0.1 / second | 60% lower | 14 |
| alibaba/happyhorse-1.1 (720p) | Pika $0.098 / second | Alibaba $0.14 / second | 30% lower | 4 |
| bytedance/seedance-2.0-fast (480p) | Atlas Cloud $0.027 / second | Fal $0.2419 / second | 89% lower | 9 |
| bytedance/seedance-2.0-mini (480p) | OpenSand $0.0104 / second | Fal $0.0721 / second | 86% lower | 8 |
| seedance-2-mini-unrestricted (480p) | OpenSand $0.0114 / second | SandBase $0.0721 / second | 84% lower | 3 |
| kling-o3 (720p) | SandBase $0.0588 / second | Tencent TokenHub $0.084 / second | 30% lower | 4 |
| minimax/h3-max (480p) | SandBase $0.01 / second | MiniMax $0.05 / second | 80% lower | 5 |
| kling-v3 (2160p) | SandBase $0.294 / second | Tencent TokenHub $0.42 / second | 30% lower | 8 |
Per second, before VideoRouter's 2% platform fee. For tiered models each row compares the resolution tier with the widest host-to-host gap. Built 2026-10-02 from the live catalog.
See how to find the cheapest video API for your workload and hidden costs in video APIs for the mechanics, and the calculator to model the effect of each lever.
Do not budget on a discount you have not verified
A sensible default for planning is to price video at the standard per-second rate and treat any discount as upside you confirm later. If you later find a documented discount that covers your modality, your budget improves; if you had assumed one that does not exist, you have an overrun. The asymmetry favours conservative planning, and it is cheap to verify by running one short job and comparing the charge with the rate you assumed.
How to ask any provider about discounts
Whether you are comparing VideoRouter with a direct vendor or another gateway, these questions get to a number you can use:
- Does the discount apply to the exact modality, model and resolution I will use, or only to text models?
- Is it the supplier's real discounted rate passed through, or a gateway markdown? Where does the saving come from?
- What is the completion window, and what happens if the job is late or fails: refund, retry, or silent charge?
- Does the discount stack with the platform fee, and is the fee quoted before or after the discount?
- Are there volume tiers or commitments, and what do they require up front?
- Are failed or timed-out jobs billed?
Write the answers down next to the unit-cost formula from the budget guide. If a discount cannot be tied to a specific model and window, treat it as marketing until the provider's documentation says otherwise. VideoRouter's own pages are the place to check what it currently offers; to test the video route with a capped key, create an account.
Frequently asked questions
Does flex mode apply to video generation?
VideoRouter's Flex Mode documentation describes it on chat completions via service_tier. It does not list video generation, so do not budget for a flex discount on video.
Is there a batch discount for video?
The Batch API documentation lists chat completions only, so assume no batch discount for video unless the documentation changes.
What is the trade-off of flex mode?
Slower responses in exchange for the upstream provider's discounted rate where the model has a native flex tier; otherwise the request is served at the standard price.
How can I lower video cost without a discount?
Compare hosts, draft on a cheaper tier, request only the seconds you keep, and reduce retries. These usually matter more than a percentage discount.
Keep reading
- How AI Video API Pricing Works — Per-Second, Per-Tier, Per-Host
- Cheapest Video Generation API: How to Find It for Your Workload
- Video API Pricing by Resolution: 480p to 4K Per-Second Costs
- How to Estimate Your AI Video API Budget, Step by Step
VideoRouter puts it next to dozens of other video and image models behind one API key, so you can compare providers, prices and fail over automatically. Compare providers on VideoRouter →