Kling Lip Sync on AI Compare Hub
Kling Lip Sync is KwaiVGI's lipsync AI video model — adding lip-sync to any short video with either an audio file or a text prompt voiced through one of 45 built-in voices. Available on AI Compare Hub, Kling Lip Sync is the most affordable lipsync tier in this lineup and is uniquely designed to plug directly into the Kling video pipeline, supporting both English and Mandarin voice synthesis with adjustable speech rate.
What you can create
-
Lipsync for Kling-generated videos
Add dialogue to videos previously created with Kling by passing the original video_id directly to Kling Lip Sync. The pipeline keeps content authoring inside the Kling ecosystem from generation through to voiced output.
-
Text-driven voiceover with built-in voices
Skip a separate text-to-speech step entirely. Provide the script as text and let Kling Lip Sync synthesise the voiceover with one of 45 built-in voices, then re-animate the source video's mouth in a single generation.
-
Bilingual social-format content
Produce English and Mandarin variants of the same source video using the bilingual voice catalogue. Kling Lip Sync's built-in voices cover narrator, character, age-range, and regional-accent options in both languages.
-
Quick voiced product clips and demos
Voice short product demos, explainers, and announcements directly from a script. Kling Lip Sync's text-mode workflow is ideal when iteration speed and cost matter more than custom voice talent.
Why creators choose Kling Lip Sync
-
Audio file or text — your choice
Kling Lip Sync uniquely accepts either a recorded audio file or a text prompt as the driving input. Text mode generates the voiceover internally with a chosen voice_id and voice_speed, removing the need for a separate text-to-speech step.
-
45 built-in voices, English and Mandarin
The voice_id catalogue includes 45 voices spanning narrator, character, age, and regional-accent options across English and Mandarin. Kling Lip Sync makes bilingual lipsync content straightforward without integrating a third-party TTS provider.
-
Voice speed control
The voice_speed parameter (0.8–2.0) lets creators tune speech rate for energy, pacing, or accessibility. Kling Lip Sync handles slow narration, conversational pace, and high-energy delivery from a single text input.
-
Lowest per-second cost in this lineup
Kling Lip Sync is priced at the most affordable per-second rate among the lipsync models on AI Compare Hub, making it a strong default choice for short-form, social, and bulk lipsync work where cost discipline matters.
How to generate your first Kling Lip Sync video
- Provide your source video. Upload an MP4 or MOV file (under 100 MB, 2–10 seconds long, 720p–1080p) or pass the video_id of a previously generated Kling video.
- Choose audio mode or text mode. Either upload an audio file (MP3, WAV, M4A, or AAC, ≤ 5 MB) or enter a text script and pick a voice_id from the catalogue plus a voice_speed.
- Generate. Run Kling Lip Sync on AI Compare Hub and review the output. Compare side-by-side with other lipsync models on the same source video.
Common questions
What is Kling Lip Sync?
Kling Lip Sync is KwaiVGI's AI video lipsync model, available via the Replicate API. It re-animates a source video's mouth to match either an uploaded audio file or a text prompt voiced through one of 45 built-in voices in English or Mandarin. Kling Lip Sync is the most affordable lipsync tier on AI Compare Hub and integrates directly with the Kling video pipeline.
Does Kling Lip Sync support text-driven voice synthesis?
Yes — Kling Lip Sync uniquely supports two driving input modes. In audio mode, it lipsyncs to an uploaded audio file (MP3, WAV, M4A, or AAC up to 5 MB). In text mode, it synthesises the voiceover internally from a text script using a chosen voice_id (45 voices, English and Mandarin) and an adjustable voice_speed (0.8–2.0), then re-animates the source video's mouth in the same generation.
What are the source video requirements for Kling Lip Sync?
The source video must be an MP4 or MOV file, less than 100 MB, between 2 and 10 seconds in duration, and at a resolution between 720p and 1080p (720–1920 px on the longer side). You can either upload the video directly via video_url or pass the video_id of a previously generated Kling video — but not both at once.
How can you use Kling Lip Sync on AI Compare Hub?
To generate lipsync videos with Kling Lip Sync on AI Compare Hub, click the "Kling Lip Sync" button at the top of this page. Provide your source video, choose audio mode (upload an audio file) or text mode (enter a script and select a voice), and generate in seconds. You can also compare Kling Lip Sync side-by-side with other leading AI lipsync models — all in one place, for free.
Key Parameters
- Category: Video
- Video-to-Video generation supported
- Processing speed: fast
For the Use of This Model
Kling Lip Sync by Kwaivgi (Kuaishou) is a video lipsync model that re-animates the mouth in a source video to match either an uploaded audio file or text-to-speech audio. Available via the Replicate API on AI Compare Hub.
- Use responsibly. Do not create or share content that is harmful, misleading, or that violates others' rights.
- Outputs & responsibility. You control how you use the videos generated here. Your inputs and outputs may be temporarily retained by the service provider for abuse monitoring and reliability.
- Model focus. Kling Lip Sync is a video-to-video lipsync model: it requires a source video plus either a replacement audio file OR text input (with optional voice ID + speed) and returns a re-animated video.
- Terms of use. Your use of this model is governed by the provider's page on Replicate: replicate.com/kwaivgi/kling-lip-sync.
Try Kling Lip Sync on AI Compare Hub
Generate with Kling Lip Sync directly on AI Compare Hub. Compare results side by side with other leading AI models using the same prompt.