Sep 17, 2026
Wan 3.0 vs Wan 2.7: What’s Different, Pricing, and Which to Use (2026)
Cost Optimization
Distributed Inference
Wan 3.0 doubles the clip length to 30 seconds and reads documents; Wan 2.7 is 40% cheaper at 1080p and still has the edit endpoint. Specs, Gateway prices, and API calls for both.

Alibaba’s Wan 3.0 replaced four models with one and doubled the clip length. It did not replace Wan 2.7 on price. Here’s the split, with Gateway rates and working API calls.
Wan 2.7 shipped in early April 2026 as a suite of four models: text-to-video, image-to-video, reference-to-video, and video edit, each capped at 15 seconds. Wan 3.0 followed on August 6 as a single model that handles all of those modes in one architecture, generates up to 30 seconds in a single pass, and takes documents, spreadsheets, and web pages as references. Both are on Yotta AI Gateway, and the choice between them is less obvious than the version numbers suggest, because 2.7 is the cheaper model at the resolutions most people use.
The two models side by side
| Wan 3.0 | Wan 2.7 | |
| Released | August 6, 2026 (public beta) | Early April 2026 |
| Architecture | One unified model for all modes | Four separate models |
| Max length per request | 30 seconds (2 to 30, or smart duration) | 15 seconds (2 to 15) |
| Resolutions | 480p, 720p, 1080p | 720p, 1080p |
| Audio | Native, on by default (generate_audio) | Native, or driven by an uploaded audio file (audio_url) |
| Reference inputs | Up to 10 images, 5 videos, and 5 audio clips (text, image, audio, video, documents) | Up to 5 mixed images and videos, with a voice-reference audio per video |
| Document and web page input | Yes (doc, xls, ppt, pdf, txt, key, pages, URLs) | No |
| Prompt length | 20,000 characters | 5,000 characters |
| Gateway modes | Text-to-video, image-to-video, reference-to-video | Text-to-video, image-to-video, reference-to-video, video edit |
| Gateway price, 720p | $0.07 per second | $0.055 per second |
| Gateway price, 1080p | $0.14 per second | $0.0825 per second |
| Gateway price, 480p | $0.035 per second | Not offered |
| Open weights | API-only, as far as we can verify | API-only, as far as we can verify |
Three rows decide most of it. Length: 3.0 does 30 seconds in one generation, 2.7 stops at 15. Inputs: 3.0 reads documents and takes twice the reference images. Price: 2.7 is cheaper at every resolution both models share, by 21% at 720p and 41% at 1080p. Everything else is detail.
What Wan 3.0 changed
The big change is structural. Wan 2.7 was four models with four endpoints and four sets of limits. Wan 3.0 is one model, which Alibaba’s Tongyi Lab describes as unified support for reference, editing, replication, and driving, with omni-modal reference: text, images, audio, video, and structured files. On the Gateway that shows up as three endpoints (text, image, and reference-to-video) that share one parameter set and one price list.
The 30-second limit is the number people will notice. Wan 2.7 and every other model on the Gateway except Seedance 2.5 stop at 15 seconds; Wan 3.0 generates 30 in a single pass with no stitching, and its “smart duration” mode (duration: -1) lets the model pick the length from the prompt. That’s the difference between a product clip and a product ad.
Document input is the feature nobody else has. Wan 3.0 accepts doc, xls, ppt, pdf, txt, key, and pages files as references and can parse web pages, which means a pitch deck or a spec sheet can be the source material for a video. Whether that produces something usable depends on the document, but it’s a mode that didn’t exist before August.
Reference-to-video got bigger too. Wan 2.7 takes up to 5 references total, images and videos mixed. Wan 3.0 takes up to 10 images, 5 videos, and 5 audio clips, each category with its own cap, with total video and audio reference length capped at 15 seconds each. For character consistency across a series of clips, that’s the difference between a face and a costume and a whole cast.
What Wan 2.7 still does better
Price. At 1080p, the resolution both models default to, Wan 2.7 is $0.0825 per second on the Gateway and Wan 3.0 is $0.14. A 15-second 1080p clip is $1.24 on 2.7 and $2.10 on 3.0. At 720p it’s $0.055 against $0.07. Wan 3.0’s only price advantage is the 480p tier at $0.035 per second, which 2.7 doesn’t offer; if 480p is acceptable, 3.0 is the cheaper model, and it’s the cheapest way on the Gateway to get 30 seconds of anything.
Audio-driven generation. Wan 2.7’s text-to-video and image-to-video endpoints take an audio_url, and the model uses it as the driving track: lip sync, action timing, cuts on the beat. Wan 3.0’s text and image endpoints generate audio but don’t take one as input; audio goes in through reference-to-video instead. If you’re making a video to a specific voiceover or song, 2.7’s path is more direct.
Voice cloning. Wan 2.7’s reference-to-video takes a reference_audio per reference video, 1 to 10 seconds, and generates the character’s voice in that register. Wan 3.0 takes audio references too, but 2.7’s per-character pairing is the more controllable version.
Video edit. Wan 2.7 has a fourth endpoint on the Gateway, wan2.7-video-edit, for prompt-driven edits to an existing clip: localized or global changes, element replacement from image references, and replicating motion, effects, and camera moves from a reference. Wan 3.0’s unified model covers editing in Alibaba’s description, but there’s no dedicated edit endpoint for it on the Gateway as of this writing.
Price, worked out
| Clip | Wan 3.0 | Wan 2.7 |
| 5 seconds, 720p | $0.35 | $0.28 |
| 10 seconds, 1080p | $1.40 | $0.83 |
| 15 seconds, 1080p | $2.10 | $1.24 |
| 30 seconds, 1080p | $4.20 | Not possible in one request |
| 30 seconds, 480p | $1.05 | Not possible in one request |
Audio doesn’t change the price on either model; the Gateway lists the same rate with and without it. For scale, the same 15-second 1080p clip is $6.45 on Seedance 2.5 and about $1.17 on PixVerse V6 with audio, so Wan 2.7 sits next to PixVerse as one of the two cheapest current 1080p options on the Gateway, and Wan 3.0 sits in the middle of the catalog.
Which one to use
Choose Wan 3.0 if you need more than 15 seconds, if a document or web page is the source material, if you need more than 5 references in one shot, or if 480p output is fine and you want the lowest per-second rate on the Gateway for a long clip.
Choose Wan 2.7 if the clip is 15 seconds or under and 1080p matters, since it’s 41% cheaper there; if you’re generating to an existing audio track; if you need voice cloning per character; or if you need the edit endpoint.
For most volume work, that means 2.7 by default and 3.0 when the brief needs its length or its inputs. Since both sit behind the same key and the same endpoints, switching is a model string.
How to call Wan on Yotta AI Gateway
All video generation on the Gateway is asynchronous: submit, get a request ID, poll until the status is completed, read the output URL. Base URL https://gateway.yottalabs.ai/api/maas, key in the X-API-KEY header. Model strings: wan3.0-t2v, wan3.0-i2v, wan3.0-r2v, wan2.7-t2v, wan2.7-i2v, wan2.7-r2v, wan2.7-video-edit.
Text-to-video, Wan 3.0, 30 seconds
curl -X POST "https://gateway.yottalabs.ai/api/maas/text-to-video/generations" \
-H "X-API-KEY: $YOTTA_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "wan3.0-t2v",
"parameters": {
"prompt": "A 30-second product spot for a matte black espresso machine: open on steam rising in morning light, cut to the portafilter locking in, a slow orbit around the machine, and a final shot of the cup on a marble counter. Warm, quiet soundtrack.",
"aspect_ratio": "16:9",
"resolution": "1080p",
"duration": 30,
"generate_audio": true,
"enhance_prompt": true
}
}'
aspect_ratio accepts adaptive, 16:9, 9:16, 1:1, 4:3, and 3:4 and defaults to adaptive. resolution defaults to 1080p. duration is 2 to 30, or -1 to let the model choose. enhance_prompt rewrites short prompts and is on by default. On Wan 2.7 the same call uses wan2.7-t2v, duration caps at 15, resolution is 720p or 1080p, and you can add "audio_url" to drive the video from a track.
Reference-to-video, Wan 3.0, product plus voice
curl -X POST "https://gateway.yottalabs.ai/api/maas/reference-to-video/generations" \
-H "X-API-KEY: $YOTTA_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "wan3.0-r2v",
"parameters": {
"prompt": "The presenter from @Image1 holds the bottle from @Image2 and talks to camera in the voice from @Audio1, vertical, studio lighting, 15 seconds.",
"reference_items": [
{"url": "https://your-cdn.example.com/presenter.png", "type": "image"},
{"url": "https://your-cdn.example.com/bottle.png", "type": "image"},
{"url": "https://your-cdn.example.com/voice.mp3", "type": "audio"}
],
"aspect_ratio": "9:16",
"resolution": "1080p",
"duration": 15,
"generate_audio": true
}
}'
Wan 3.0 references are numbered by type in order (@Image1, @Image2, @Audio1, @Video1); there are no custom names. Images are jpeg, jpg, png, bmp, or webp up to 20 MB; videos mp4 or mov up to 100 MB; audio wav or mp3 up to 15 MB. Including a reference video shortens the maximum output by the reference’s length. On Wan 2.7 the equivalent call uses wan2.7-r2v, up to 5 items of type image or video, an optional reference_audio on each video item for voice, and an optional first_frame.
Poll for the result
curl "https://gateway.yottalabs.ai/api/maas/reference-to-video/generations/$REQUEST_ID" \
-H "X-API-KEY: $YOTTA_API_KEY"
The submit call returns data.request_id; the poll endpoint matches the mode and returns status, then output_url and duration when complete. The Seedance API guide has a Python client with the polling loop that works for every video model on the Gateway.
Frequently asked questions
What’s the difference between Wan 3.0 and Wan 2.7? Wan 3.0 is one model that generates up to 30 seconds, takes documents and web pages as input, and accepts up to 10 image references. Wan 2.7 is four models capped at 15 seconds, with audio-driven generation, per-character voice cloning, and a video edit endpoint. On Yotta AI Gateway, 2.7 is cheaper at 720p and 1080p; 3.0 adds a 480p tier.
Which is cheaper, Wan 3.0 or Wan 2.7? Wan 2.7 at 720p ($0.055 vs $0.07 per second) and 1080p ($0.0825 vs $0.14). Wan 3.0 is the only one with 480p, at $0.035 per second.
How long can a Wan 3.0 video be? Up to 30 seconds in a single request, or the model chooses the length in smart duration mode. Wan 2.7 caps at 15 seconds, and 10 when a reference video is included.
Does Wan 3.0 generate audio? Yes, on by default, at no extra cost on the Gateway. Wan 2.7 generates audio too, and can also take an uploaded audio file as the driving track on text-to-video and image-to-video.
Can Wan 3.0 make a video from a PDF or a web page? That’s its headline feature: doc, xls, ppt, pdf, txt, key, and pages files and web pages are accepted as references. Wan 2.7 can’t.
Is Wan 2.7 still worth using? Yes. It’s the cheaper model at 1080p by 41%, it has the edit endpoint, and it has the more direct audio controls. It’s the default for volume work under 15 seconds.
Are Wan 3.0 and Wan 2.7 on Yotta AI Gateway? Both, in text-to-video, image-to-video, and reference-to-video, plus video edit for 2.7, all behind one key alongside Seedance 2.5 and 2.0, PixVerse V6 and C1, HappyHorse-1.0, and Kling v3.
Bottom line
Wan 3.0 is the model for length and inputs: 30 seconds, documents, ten references. Wan 2.7 is the model for price and control: 41% cheaper at 1080p, audio-driven, voice-cloned, editable. Alibaba shipped an upgrade that costs more, which is unusual, and the practical result is that both models have a job. Run both through Yotta AI Gateway with one key and pick per brief. The rest of the video lineup, with per-second pricing, is in the Gateway video announcement.



