No video yet. Submit the form to generate video.
Explore different use cases and parameter configurations
Transparent pricing with no hidden fees. Pay as you go.
| Rule & Modality | Channel | Credits | Price (USD) | Official / Reference Price | Daily Savings |
|---|---|---|---|---|---|
test videowan | Regular | 999per run | $4.541 | - | - |
Complete guide to using wan-3-0
Use Wan 3.0 on ApiPass to build text-to-video, image-to-video, and reference-material-to-video workflows. Create longer, more controllable videos from prompts, keyframes, images, video clips, audio, documents, and public webpages through a unified asynchronous API.

Wan 3.0 is an all-in-one video generation model designed for multimodal creation, longer-form storytelling, richer reference understanding, stronger visual consistency, and more precise creative control. It can use text prompts, images, videos, audio, keyframes, documents, and webpages to guide scenes, motion, sound, visual identity, and narrative direction. On ApiPass, Wan 3.0 uses a unified task workflow: create a task, receive an ApiPass task ID, then poll or receive a callback when the generated video is ready.
Turn detailed prompts into complete video concepts. Written instructions can guide characters, environments, camera movement, action logic, lighting, pacing, atmosphere, and emotional rhythm.
Use a first-frame image, and optionally a last-frame image, to anchor visual identity and scene progression while Wan 3.0 creates natural motion between key moments.
Guide character appearance, clothing, props, locations, materials, composition, and visual style with image references so details remain recognizable during motion.
Use existing footage as a motion and scene-flow reference. Reference clips can help define performance cues, timing, camera behavior, and dynamic movement.
Use speech, music, rhythm, vocals, or sound references to help shape pacing, mood, performance energy, and audiovisual coordination.
Bring briefs, presentations, PDFs, text documents, and public webpages into the creative workflow. ApiPass supports at most one reference file or one reference link per request.
Wan 3.0 expands video creation beyond prompts by incorporating richer creative context from images, videos, audio, files, and webpages. This makes it easier to guide scene development, visual consistency, and creative direction.
Wan 3.0 supports normal output durations from 2 to 30 seconds, plus intelligent duration mode with duration=-1. Longer generation gives scenes more room to establish context, develop actions, transition between story beats, and reach a satisfying conclusion.
Characters, clothing, props, environments, audio cues, and spatial relationships can remain closely aligned with their references throughout changing actions and scenes, making reference-driven creation more dependable.
Existing videos can be reshaped through detailed creative instructions without unnecessarily changing everything in the original footage. Use reference-guided editing to refine subjects, backgrounds, lighting, visual style, and scene elements.
Wan 3.0 pairs visual storytelling with expressive audio direction, helping motion, performance, atmosphere, music, speech, and environmental sound develop with stronger cohesion.
Control lighting, color, composition, character performance, camera language, scene dynamics, action choreography, physical interactions, materials, and nuanced movement for more cinematic results.
Create short films that develop across multiple actions, locations, camera changes, and story beats. Wan 3.0 is suitable for urban chases, miniature adventures, surreal comedy, cinematic action scenes, and other connected narratives.
Use images, videos, and other materials as persistent guidance for characters, costumes, environments, styling, spatial relationships, and movement continuity in fashion films, choreography sequences, character showcases, and historical scenes.
Start from existing video and apply controlled creative changes such as replacing a subject, changing an environment, removing an accessory, shifting visual style, adjusting selected scene elements, or extending a sequence.
Produce videos built around dialogue, music, environmental sound, rhythmic motion, or expressive atmosphere, including performance videos, cinematic commercials, immersive story scenes, and rhythm-driven motion pieces.
Wan 3.0 is designed for projects that need stronger multimodal references, longer generation, and more controllable audiovisual storytelling.
Many video workflows begin and end with a prompt. Wan 3.0 can still generate from text alone, but it also supports richer creative context from images, video clips, audio, keyframes, files, and public links, making it more practical for production workflows that need continuity and control.
Short video models are useful for isolated moments, but longer ideas need more room. Wan 3.0 supports durations from 2 to 30 seconds and intelligent duration mode, giving scenes space to develop action, pacing, transitions, and narrative structure.
Wan 3.0 is especially useful when character identity, clothing, props, scene layout, motion language, and audiovisual direction need to remain consistent. ApiPass automatically selects reference-material-to-video mode when supported reference inputs are present.
Current Wan 3.0 baseline pricing uses ApiPass model credit rules. Normal durations from 2 to 30 seconds cost 30 credits per second.
Requests with duration=-1, or requests where duration is omitted or empty, use the current per-video baseline of 150 credits per video.
Credits are pre-deducted before task creation. If provider generation fails and no successful result is produced, pre-deducted credits are refunded according to the existing ApiPass process.
Create or log in to your ApiPass account, generate an API key, and prepare your request environment with Authorization: Bearer ${APIPASS_API_KEY} and Content-Type: application/json.
Build a request with model set to wan/wan-3-0 and provide a prompt. Add first_frame_url for image-to-video, or add reference_image_urls, reference_video_urls, reference_audio_urls, reference_file_urls, or reference_link_urls for reference-material generation.
Send a POST request to /api/v1/jobs/createTask. Save data.taskId from the response, then poll /api/v1/jobs/recordInfo every 3–5 seconds until data.state becomes success or fail.
When the task succeeds, read data.resultJson.resultUrls for the generated video URLs processed by the ApiPass CDN result flow. If you provided callBackUrl, ApiPass can also POST the terminal task record to your endpoint.
All APIs require authentication via Bearer Token.
Authorization: Bearer
Submit a Wan 3.0 video generation task. ApiPass automatically routes the request as text-to-video, image-to-video, or reference-material-to-video based on the input parameters.
The API accepts a JSON payload with the following structure:
1{
2 "model": "wan/wan-3-0",
3 "callBackUrl": "string (optional)",
4 "channel": "auto",
5 "input": {
6 "prompt": "string",
7 "resolution": "1080p",
8 "duration": 5,
9 "aspect_ratio": "adaptive",
10 "audio": true
11 }
12}modelRequiredstringThe model name to use. For this endpoint it must be set to wan/wan-3-0.
"wan/wan-3-0"
callBackUrlOptionalstringTop-level task completion callback URL. It is not part of the input object. If omitted, no callback notification will be sent.
"https://your-domain.com/api/callback"
channelOptionalstringYou may specify the corresponding provider within APIPASS via the channel parameter; these providers handle the actual image and video generation tasks. APIPASS currently offers three provider options:
The default channel is auto. APIPASS automatically allocates tasks across available providers based on real-time pricing and stability metrics. starter is ultra-low-cost with limited quotas and weaker stability; regular is standard and much cheaper than official APIs with moderate stability; official uses the model native API with high stability, fast task execution, and official pricing.
Available options:
auto
The input object contains the following parameters:
input.promptRequiredstringDescription of the video content, motion, and style. It cannot be empty and supports up to 20000 characters.
Describe the subject, action, camera movement, style, lighting, and atmosphere for best results.
"A cinematic ocean wave rolls toward a quiet beach at sunrise, realistic water and natural camera movement."
input.resolutionOptionalstringOutput video resolution. Default: 1080p. If both resolution and quality are provided, resolution takes priority.
Available options:
"1080p"
input.qualityOptionalstringCompatibility alias for input.resolution. If both fields are provided, input.resolution takes priority.
Available options:
"1080p"
input.durationOptionalinteger or numeric stringVideo duration. Supports integers from 2 to 30, or -1 for intelligent duration. Numeric strings such as "10" are converted to integers. Default: 5.
Billing is based on duration. Normal duration from 2 to 30 seconds costs credits per second, while duration=-1 or omitted duration uses the per-video baseline.
5
input.aspect_ratioOptionalstringOutput video aspect ratio. Default: adaptive.
Available options:
"adaptive"
input.aspectRatioOptionalstringCamelCase compatibility alias for input.aspect_ratio.
Available options:
"16:9"
input.ratioOptionalstringUpstream compatibility alias for input.aspect_ratio.
Available options:
"16:9"
input.audioOptionalbooleanWhether to generate audio for the video. Default: true.
true
input.generate_audioOptionalbooleanCompatibility alias for input.audio.
true
input.generateAudioOptionalbooleanCamelCase compatibility alias for input.audio.
true
input.seedOptionalintegerGeneration seed. Use -1 for a random seed, or a non-negative value from 0 to 2147483647.
apipass does not support seed=-1. If AtlasCloud is unavailable and the request uses seed=-1, apipass returns a provider capability error.
12345
input.first_frame_urlOptionalstringPublic URL of the first-frame image for image-to-video generation.
If last_frame_url is provided, first_frame_url is required. Frame inputs cannot be mixed with reference material inputs.
"https://cdn.example.com/first-frame.png"
input.last_frame_urlOptionalstringPublic URL of the last-frame image for image-to-video generation. It must be provided together with first_frame_url.
"https://cdn.example.com/last-frame.png"
input.reference_image_urlsOptionalstring[]Public image reference URLs for reference-material-to-video generation. The platform public validator allows image, video, and audio references to total up to 20 items.
apipass currently supports up to 10 reference images. Provider-specific limits may be stricter than the public ApiPass validation.
["https://cdn.example.com/character.png"]
input.reference_video_urlsOptionalstring[]Public video reference URLs for reference-material-to-video generation.
apipass currently supports up to 5 reference videos. Provider-specific limits may be stricter than the public ApiPass validation.
["https://cdn.example.com/motion.mp4"]
input.reference_audio_urlsOptionalstring[]Public audio reference URLs for reference-material-to-video generation.
ApiPass currently supports up to 5 reference audio files. Provider-specific limits may be stricter than the public ApiPass validation.
["https://cdn.example.com/dialogue.mp3"]
input.reference_file_urlsOptionalstring or string[]Public reference file URL for reference-material-to-video generation. At most one non-empty file URL is supported.
Providing a file automatically sets enable_thinking=true. file and link are mutually exclusive.
"https://cdn.example.com/storyboard.pdf"
input.reference_link_urlsOptionalstring or string[]Public reference webpage or link URL for reference-material-to-video generation. At most one non-empty link URL is supported.
Providing a link automatically sets enable_thinking=true. file and link are mutually exclusive.
"https://example.com/storyboard"
input.enable_thinkingOptionalbooleanEnables reference-material thinking mode. Reference files or links automatically set this to true, so users usually do not need to provide it. Default: false.
Providing enable_thinking=true by itself also selects reference-material mode. AtlasCloud forwards this field in reference mode; apipass currently forwards only actual reference material, files, or links and does not send enable_thinking to wan/3-0-video.
false
input.enableThinkingOptionalbooleanCamelCase compatibility alias for input.enable_thinking.
false
1curl --request POST \
2 --url "https://api.apipass.dev/api/v1/jobs/createTask" \
3 --header "Authorization: Bearer ${APIPASS_API_KEY}" \
4 --header "Content-Type: application/json" \
5 --data '{
6 "model": "wan/wan-3-0",
7 "callBackUrl": "https://your-domain.com/api/callback",
8 "channel": "auto",
9 "input": {
10 "prompt": "A cinematic ocean wave rolls toward a quiet beach at sunrise, realistic water and natural camera movement.",
11 "resolution": "1080p",
12 "duration": 5,
13 "aspect_ratio": "adaptive",
14 "audio": true
15 }
16 }'1{
2 "code": 200,
3 "message": "success",
4 "data": {
5 "taskId": "task_xxxxxxxxxxxxxxxx"
6 }
7}codeStatus code. 200 indicates that the task was created successfully; other values indicate failure.
messageResponse message. When the request fails, this field contains the error description.
data.taskIdThe ApiPass task ID. Save this ID and use it with the Query Task API. Do not use provider task IDs from AtlasCloud or apipass logs for polling.

ByteDance
ByteDance
Starting from
0 credits
ByteDance
ByteDance
Starting from
0 credits
ByteDance
ByteDance
Turn text, images, video clips, and audio into cinematic, audio-synced videos with a single API call.
Starting from
0 credits
MiniMax
MiniMax
Starting from
0 credits
ByteDance
ByteDance
Generate cinematic 4K AI videos up to 30 seconds with Seedance 2.5's multimodal power.
Starting from
0 credits
Google Veo 3.1 is the latest state-of-the-art generative video model designed to transform text and image prompts into high-fidelity cinematic visuals. Building upon its predecessors, Veo 3.1 features significant upgrades in prompt adherence, visual realism, and native audio generation, creating 8-second clips with synchronized sound effects and dialogue.
Starting from
0 credits
Runway
Runway
Starting from
0 credits
Kling
Kling
Transform text prompts and images into high-quality videos with advanced AI. Generate cinematic content with customizable duration, aspect ratio, and audio capabilities.
Starting from
0 credits
MiniMax
MiniMax
Generate high-quality AI videos from text prompts or images with MiniMax's Hailuo 2.3 model. Support for both 6s and 10s durations, 768p and 1080p resolutions, with intelligent prompt optimization.
Starting from
0 credits
grok
grok
Grok Imagine is xAI’s AI video generation model for creating high-quality videos with realistic motion, native audio, and synchronized speech. Grok Imagine 1.5 introduces faster generation speeds, improved motion physics, and stronger visual consistency for modern AI video applications.
Starting from
0 credits
Luma
Luma
Starting from
0 credits
Kling
Kling
Kling v3 video generation model, supports text-to-video and image-to-video, up to 15 seconds, 1080p pro mode
Starting from
0 credits
wan
wan
Generate videos with cinematic consistency by reimagining existing footage through high-fidelity visual transformation and structural control.
Starting from
0 credits
Starting from
0 credits
grok
grok
Starting from
0 credits
kling
kling
Starting from
0 credits