Home › Guides › Flex mode and discounts

Flex mode, batch and other discounts: what actually applies to video

Updated 2026-10-02

"Is there a discount if I can wait?" is a reasonable question for any API bill, and a frequent one for video, where jobs are slow and spend adds up. The honest answer on VideoRouter is specific: its documented discount mechanisms apply to chat completions, and the video route is priced per second at the host's rate plus the platform fee. This page lays out what the documentation says, what it does not say, and how to ask the same question of any provider without being misled.

What flex mode is, per the documentation

VideoRouter's Flex Mode page describes one extra field on a normal POST /v1/chat/completions call: service_tier: "flex". You trade a slower response for a discount, on the same endpoint with the same request and response shape. The call blocks, so you give it a long client timeout, and the response reports which tier actually served it. The documentation says that when the underlying model has a native flex tier, the request is forwarded with a real flex tier and billed at the discounted price the upstream provider itself charges. If the model has no native flex tier, the request still succeeds at the standard price.

Note what that means. Flex mode is a pass-through of a provider-side feature, not a markdown invented by the gateway, and the documentation names the providers whose flex tiers are involved. It is documented for chat completions. It is not described as applying to video or image generation, so do not plan a video budget around it.

What the batch and async options cover

The same documentation set describes a file-based Batch API for bulk work with a longer completion window, and an Async Jobs endpoint for single requests that can wait longer, delivered by polling or a webhook. Their listed endpoints are chat completions. The documentation does not list video generation among them, so the correct assumption for video is that no batch discount exists. If that changes, it will appear in the Batch API and Flex Mode pages, and those pages, not this article, are the authority.

The trade-off, stated plainly

OptionYou give upYou get
Flex tier (chat)Latency; you hold a connection open for minutesThe provider's own discounted rate, if the model has one
Batch (chat)Immediacy; results arrive within a windowBulk submission and a lower rate where the upstream offers it
Async job (chat)Synchronous responsePoll or webhook delivery for long waits
Standard (video, per second)NothingPredictable per-second billing at creation, free polling

The pattern is the usual one: you can buy a lower price with flexibility on time, but only where the upstream provider sells it. A gateway cannot invent a discount it does not receive without losing money, so be wary of any provider that advertises a batch or flex discount on a modality where its own supplier has no such tier.

Where video savings really come from

If flex does not apply, the levers are the ones that every per-second API shares, and they are bigger than most discounts anyway:

ModelCheapest hostPriciest hostCheapest isHosts
bytedance/seedance-2.5 (480p)OpenSand
$0.0525 / second
Fal-US
$0.2646 / second
80% lower9
bytedance/seedance-2.0 (2160p)MachGen
$0.59 / second
Fal
$1.5552 / second
62% lower9
alibaba/wan-3.0 (480p)Replicate
$0.025 / second
Alibaba
$0.05 / second
50% lower10
google/gemini-omni-flashGoogle
$0.1 / second
Fal
$0.13 / second
23% lower4
minimax/h3 (768p)MachGen
$0.04 / second
WaveSpeedAI-resell
$0.1 / second
60% lower14
alibaba/happyhorse-1.1 (720p)Pika
$0.098 / second
Alibaba
$0.14 / second
30% lower4
bytedance/seedance-2.0-fast (480p)Atlas Cloud
$0.027 / second
Fal
$0.2419 / second
89% lower9
bytedance/seedance-2.0-mini (480p)OpenSand
$0.0104 / second
Fal
$0.0721 / second
86% lower8
seedance-2-mini-unrestricted (480p)OpenSand
$0.0114 / second
SandBase
$0.0721 / second
84% lower3
kling-o3 (720p)SandBase
$0.0588 / second
Tencent TokenHub
$0.084 / second
30% lower4
minimax/h3-max (480p)SandBase
$0.01 / second
MiniMax
$0.05 / second
80% lower5
kling-v3 (2160p)SandBase
$0.294 / second
Tencent TokenHub
$0.42 / second
30% lower8

Per second, before VideoRouter's 2% platform fee. For tiered models each row compares the resolution tier with the widest host-to-host gap. Built 2026-10-02 from the live catalog.

See how to find the cheapest video API for your workload and hidden costs in video APIs for the mechanics, and the calculator to model the effect of each lever.

Do not budget on a discount you have not verified

A sensible default for planning is to price video at the standard per-second rate and treat any discount as upside you confirm later. If you later find a documented discount that covers your modality, your budget improves; if you had assumed one that does not exist, you have an overrun. The asymmetry favours conservative planning, and it is cheap to verify by running one short job and comparing the charge with the rate you assumed.

How to ask any provider about discounts

Whether you are comparing VideoRouter with a direct vendor or another gateway, these questions get to a number you can use:

  1. Does the discount apply to the exact modality, model and resolution I will use, or only to text models?
  2. Is it the supplier's real discounted rate passed through, or a gateway markdown? Where does the saving come from?
  3. What is the completion window, and what happens if the job is late or fails: refund, retry, or silent charge?
  4. Does the discount stack with the platform fee, and is the fee quoted before or after the discount?
  5. Are there volume tiers or commitments, and what do they require up front?
  6. Are failed or timed-out jobs billed?

Write the answers down next to the unit-cost formula from the budget guide. If a discount cannot be tied to a specific model and window, treat it as marketing until the provider's documentation says otherwise. VideoRouter's own pages are the place to check what it currently offers; to test the video route with a capped key, create an account.

Frequently asked questions

Does flex mode apply to video generation?

VideoRouter's Flex Mode documentation describes it on chat completions via service_tier. It does not list video generation, so do not budget for a flex discount on video.

Is there a batch discount for video?

The Batch API documentation lists chat completions only, so assume no batch discount for video unless the documentation changes.

What is the trade-off of flex mode?

Slower responses in exchange for the upstream provider's discounted rate where the model has a native flex tier; otherwise the request is served at the standard price.

How can I lower video cost without a discount?

Compare hosts, draft on a cheaper tier, request only the seconds you keep, and reduce retries. These usually matter more than a percentage discount.

Keep reading

Using AI video is one part of the job.

VideoRouter puts it next to dozens of other video and image models behind one API key, so you can compare providers, prices and fail over automatically. Compare providers on VideoRouter →