What is Wan 3.0?

Wan 3.0 is Alibaba Tongyi Lab's video model, the newest release in the Wan (Tongyi Wanxiang) family that started with Wan 2.1 and moved through Wan 2.7. It is a text-to-video model at heart: type a prompt and it renders a moving clip, but it also takes an image or a reference clip as the starting point.

The headline change from earlier Wan releases is length and sound. Wan 3.0 renders up to 15 seconds of video in a single generation pass, at resolutions up to 1080p, instead of stitching short clips together. Audio is generated in the same pass as the picture, so dialogue, ambience, and motion land in sync without a separate voiceover step.

How it works inside MagicShot

On MagicShot, Wan 3.0 sits alongside other video models as one more engine you can pick for a job. You give it a prompt, a photo, or a short reference clip, choose an aspect ratio, and it renders the finished video.

  • Text-to-video: describe the shot and Wan 3.0 builds it from scratch.

  • Image-to-video: upload a product photo or portrait and animate it.

  • Reference-to-video: feed it a clip and it keeps characters, props, and scenes consistent across the new shot.

Where it fits

Wan 3.0 is built for continuous camera movement and one-take shot language, the kind of long, unbroken pan that used to mean joining several short clips. That makes it a fit for product walkthroughs, short-form social edits, and ad concepts where a single flowing shot reads better than a cut-up sequence.

Every aspect ratio from 16:9 down to 9:16 is supported, so the same generation can target YouTube, Reels, or TikTok without a separate crop pass.