Filtrix model field guide / 014Released June 2026
Alibaba Cloud · Image to video

Happy Horse1.1

Turn one still frame into a finished audiovisual shot—with cinematic motion, native sound, multilingual lip-sync, and up to 1080p output.

1080p

Maximum resolution

3–15s

Flexible clip length

24 fps

MP4 video output

Audio

Generated in the same pass

Sample outputPicture + sound
Official model demo

Input

First frame

Output

Video + audio

Range

3–15 seconds

First-frame I2V
720p / 1080p
3–15 seconds
Native audio
Multilingual lip-sync

The 1.1 upgrade

Built to hold the shot together.

Alibaba Cloud positions version 1.1 as a production-focused upgrade across motion, consistency, instruction following, texture, and sound.

01

Stronger dynamic expression

Happy Horse 1.1 is tuned for energetic, readable action—from subtle gestures to fast product and character movement.

02

More stable character identity

Improved consistency helps a subject retain its defining facial, wardrobe, and visual details while the shot develops.

03

Better prompt direction

Describe the action, camera, lighting, dialogue, and sound as one production brief with stronger instruction following.

04

Richer visual texture

The 1.1 release improves material detail, lighting texture, and overall finish for more production-ready footage.

05

Native audiovisual generation

Create motion and sound together, including ambience, effects, dialogue, and multilingual lip-sync.

Prompt as a shot list

Direct what moves—and what the audience hears.

Treat the prompt like a compact production brief. Concrete actions, a clear camera move, and explicit sound direction give the model more to work with than mood words alone.

Shot direction / 00:00–00:101080p
01 / Subject

A ceramic perfume bottle on a wet black-stone pedestal

02 / Action

Water beads slide down the glass as the cap turns slowly

03 / Camera

Macro opening, then a smooth half-orbit at product level

04 / Light

A warm rim light catches the glass; cool reflections move behind it

05 / Sound

Soft rain, glass resonance, one precise metallic click; no music

Be specific about the visible action and the audible result.Try your prompt

Production use cases

More than motion for motion's sake.

The release focuses on the high-frequency content workflows where identity, movement, texture, and sound all need to survive the same take.

Narrative

Short drama & story beats

Animate a character-led keyframe into a complete 3–15 second moment with action, camera movement, dialogue, and atmosphere.

Commerce

Product advertising

Turn a clean product still into motion-led ecommerce creative with tactile detail and a soundtrack generated in the same pass.

Campaigns

Brand marketing

Develop visual concepts for launches and social campaigns while preserving the composition and identity of the source image.

Worldbuilding

Game CG concepts

Bring character art, environments, and cinematic keyframes to life for pitches, mood films, and previsualization.

Image-to-video workflow

From first frame to final take.

The fal endpoint keeps the controls focused: one image, one optional prompt, resolution, duration, seed, and safety checking.

Step 01

Choose a strong first frame

Upload a JPEG, PNG, BMP, or WebP image at least 300px wide and high. The file can be up to 20 MB.

Step 02

Direct motion and sound

Write up to 2,500 characters covering the subject, action, camera, lighting, dialogue, ambience, and sound effects.

Step 03

Set the final take

Choose 720p or 1080p, select any duration from 3 to 15 seconds, then generate and refine the prompt if needed.

Input

JPEG · PNG · BMP · WebP

Image size

300px minimum · 20 MB max

Aspect ratio

1:2.5 to 2.5:1

Safety

Input + output checking

FAQ

Before you roll.

Confirmed against the Alibaba Cloud release notes and the current image-to-video API schema.

01What is Happy Horse 1.1?+

Happy Horse 1.1 is Alibaba's image-to-video model for creating cinematic clips from a first-frame image, with native audio, multilingual lip-sync, and output up to 1080p.

02Does the model generate audio?+

Yes. It can generate video and synchronized audio in the same pass, including ambience, effects, dialogue, and multilingual lip-sync.

03What clip lengths are supported?+

The API accepts any whole-number duration from 3 to 15 seconds.

04What resolutions are available?+

You can generate at 720p or 1080p. The returned result is an MP4 video at 24 frames per second.

05What image formats can I upload?+

The API accepts JPEG, PNG, BMP, and WebP images. Images must be at least 300 × 300 pixels, up to 20 MB, and within a 1:2.5 to 2.5:1 aspect ratio.

06How should I write a prompt?+

Describe the visible subject, intended action, camera movement, lighting, dialogue, ambience, and sound effects. The prompt can contain up to 2,500 characters.

Ready for your next frame?

Give a still image its voice.

Open the Filtrix image-to-video workspace and turn your keyframe into a polished, shareable video.

Technical sources: Alibaba Cloud release notes and fal image-to-video API schema.

Specifications checked July 2026