Real Photo, New Moves
A portrait in a red dress takes on a yoga balance sequence while the meadow, light, and fabric stay true to the photo.
Upload one character image and one motion reference video. Kling motion control reads the dance, gesture, or action in the clip and performs it with your character, keeping the face, outfit, and style consistent.
No distortion, the character's movements are consistent with the video. The woman in the red dress does the yoga sequence in the same sunny meadow.
Reference example. It is not a result generated from your upload.
Kling motion control is a motion transfer feature of Kling AI, the video model family from Kuaishou. Instead of guessing movement from a text prompt, it takes two inputs, a character image and a motion reference video, and generates a new clip in which your character performs the exact moves, timing, and gestures of the person in the video. That makes it a true motion control AI: the performance comes from real footage, so dances, sports moves, hand gestures, and acting beats come out far more predictable than ordinary image-to-video. Think of it as Kling AI motion capture without a mocap suit: an AI motion control video generator that only needs a phone clip.
Kling 2.6 motion control introduced the feature with full-body motion transfer, precise hand performance, and one-shot clips up to 30 seconds. Kling 3.0 motion control (often shortened to Kling 3 motion control) builds on it with steadier facial identity from any angle, better handling of complex expressions and briefly hidden faces, and an option to keep the background from your image instead of the video. On EzImgMaker you can run both Kling motion versions in one simple editor, no node graph or API setup needed.
What each version accepts and produces on EzImgMaker.
| Spec | Kling 3.0 Motion Control | Kling 2.6 Motion Control |
|---|---|---|
| Inputs | 1 character image + 1 motion video, optional prompt | 1 character image + 1 motion video, optional prompt |
| Character image | JPG or PNG, up to 10 MB, over 340 px, ratio 2:5 to 5:2 | JPG or PNG, up to 10 MB, over 300 px, ratio 2:5 to 5:2 |
| Motion reference video | MP4 or MOV, 3–30 s, up to 100 MB, over 340 px | MP4 or MOV, 3–30 s, up to 100 MB |
| Output length | Follows the reference: up to 30 s (match video) or 10 s (match image) | Follows the reference: up to 30 s (match video) or 10 s (match image) |
| Resolution | 720p Standard or 1080p Pro | 720p Standard or 1080p Pro |
| Background source | From the video or from the image | From the video |
| Prompt length | Up to 2,500 characters | Up to 2,500 characters |
| Best for | Close-ups, expressive faces, multi-angle and long takes | Full-body dances, sports moves, hand gestures, lower cost |
The generated clip normally matches the length of the usable motion in your reference video.
Why AI motion control gives you more say over the movement than a prompt alone.
| Motion control | Image-to-video | |
|---|---|---|
| Where the motion comes from | A real reference video | Your text prompt |
| Timing and choreography | Copied beat for beat | Interpreted by the model |
| Repeatability | Same video, same moves | Varies between generations |
| Inputs | Character image + motion video | Image + prompt |
| Good for | Dance trends, choreography, gestures, acting | Ambient motion, camera moves, short scenes |
Each demo shows the character image (top left), the motion reference video (bottom left), and the generated motion control video (right).
A portrait in a red dress takes on a yoga balance sequence while the meadow, light, and fabric stay true to the photo.
A simple cartoon character dances real ballet steps. The flat art style survives every spin and arm movement.
Slow, flowing tai chi from a live-action clip becomes smooth anime motion with the same outfit and starry background.
Three steps from a still character to a moving one.
Add a clear image of your character, a photo, illustration, or 3D render, plus a 3–30 second reference clip of one person doing the moves you want. Head, shoulders, and torso should be visible in both.
Choose Kling 3.0 or Kling 2.6 motion control, then decide whether the character faces the way it does in your image or follows the video. Add an optional prompt for the scene, lighting, or style.
Choose 720p or 1080p and generate. The motion control AI video follows the length of your reference clip, ready to download and share.
Posture, rhythm, and weight shifts carry over from the reference video, even through balances, turns, and big movements. Kling motion control keeps the whole body coordinated instead of animating parts in isolation.
Kling 3.0 motion control holds facial identity steady as the head turns, the camera moves, or a hand briefly covers the face, and it reproduces subtle expressions like a smile or a look of surprise.
Match video copies the orientation and camera framing of the reference for clips up to 30 seconds. Match image keeps the pose direction from your picture while it performs the moves, for clips up to 10 seconds.
The motion comes from the video, so the prompt is free for everything else: move the dancer to a beach, add stage lights, or keep your image background with Kling 3.0. Motion control AI handles the choreography; you direct the scene.
Put trending dances and gestures on an original character or avatar for Shorts and Reels without filming the character.
Drive illustrated characters, mascots, and anime avatars with real human performance instead of keyframes.
Turn a static mascot or virtual presenter into a motion control video that waves, points, and presents a product.
Test choreography, blocking, and acting beats with a consistent character before a shoot.
Clean inputs matter more than long prompts.
Use a clip with a single performer. Several people or heavy overlap confuse the motion extraction.
Both the image and the video should clearly show the upper body. Full-body shots work best for dances.
A full-body image pairs best with a full-body video; a half-body image with a half-body clip.
Avoid fast camera whips and dark footage. Very fast or chaotic motion can make the output shorter than your clip.
A clear, well-lit face in the character image helps Kling 3.0 keep identity stable from every angle.
The video already defines the motion. Use the prompt for background, lighting, and style, and add "no distortion".
Common questions about Kling AI motion control.
It is a Kling AI feature that transfers movement from a reference video onto a character image. You provide the character and the performance, and the model generates a new video of your character doing those moves, with optional prompt control over the scene.
Both take the same inputs and produce 720p or 1080p clips up to 30 seconds. Kling 3.0 motion control keeps faces more consistent across angles and long takes, reproduces complex expressions better, and can keep the background from your image. Kling 2.6 motion control is strong for full-body dances and hand gestures and costs fewer credits.
One character image (JPG or PNG, up to 10 MB, larger than 340 px for Kling 3.0 or 300 px for Kling 2.6, aspect ratio between 2:5 and 5:2) and one motion reference video (MP4 or MOV, 3–30 seconds, up to 100 MB). Both should clearly show the head, shoulders, and torso. A prompt is optional.
With orientation set to match the video, the output can be up to 30 seconds; with match image, up to 10 seconds. The result normally follows the length of your reference clip.
If parts of the reference are very fast, blurry, or chaotic, the model keeps only the clean, continuous motion it can follow. As long as at least 3 seconds of usable motion is found, a video is generated. Steadier footage gives a full-length result.
It is Kling, from Kling AI by Kuaishou. Searches like "King AI motion control", "king motion control", "kling motion contro", or "motion control Kling AI" all point to the same feature. The model names are Kling AI 2.6 motion control and Kling motion control 3.0, and on this page both are simply called Kling motion control.
Yes. The character image can be a photo, an illustration, an anime character, a mascot, or a 3D render, as long as the head, shoulders, and torso are visible. The motion reference should be a real or clearly human performance.
Generation uses credits. New users get free credits to try it, and the credit cost for your chosen model, resolution, and clip length is shown before you start.
No. EzImgMaker is an independent service that offers access to Kling motion control models. It is not affiliated with, endorsed by, or sponsored by Kuaishou or Kling AI.
Only upload footage and images you own or have permission to use, and get consent from anyone who appears in them. Do not use motion control to impersonate real people, create misleading videos, or animate photos of minors.
Upload a character image and a motion video, pick Kling 3.0 or 2.6, and watch your character take on every move.
Kling and Kling AI are trademarks of Kuaishou Technology. EzImgMaker is an independent service and is not affiliated with, endorsed by, or sponsored by Kuaishou. Model names are used only to describe compatibility. Demo clips are placeholders from public model showcases and will be replaced with our own generations.