Kling Avatar V2 on AI Compare Hub

Kling Avatar V2 by KwaiVGI turns a single portrait image and an audio file into an expressive, lip-synced talking-avatar video. It works on realistic humans, cartoons, illustrations, and animals — and outputs up to 1080p at 48 fps in Pro mode. Video duration automatically matches the length of the audio you upload, so there's no separate duration parameter to set.

What you can create

Why creators choose Kling Avatar V2

How to generate your first avatar video

  1. Pick your portrait. Use a clear, front-facing portrait with a single subject. Realistic photos, illustrations, cartoons, and animal portraits all work. Avoid heavy occlusion (sunglasses, masks) for the best lip sync results.
  2. Provide your audio. Upload an audio file (MP3, WAV, M4A, or AAC, max ~5 MB). You can also generate audio first using AI Compare Hub's Text-to-Speech feature and select it from your gallery. The output video duration matches the audio length.
  3. Configure mode and prompt. Pick Standard (720p) or Pro (1080p, 48 fps). Optionally add a prompt to direct emotion, energy, or camera framing — e.g., "warm, smiling, looking directly at the camera." Generate and review.

Common questions

What is Kling Avatar V2?

Kling Avatar V2 is KwaiVGI's portrait-driven avatar animation model, available via the Replicate API. It turns a single image plus an audio file into an expressive, lip-synced talking-avatar video — supporting realistic humans, cartoons, illustrations, and animals, with output up to 1080p at 48 fps in Pro mode.

How is the video duration determined?

Output duration matches the length of the audio file you upload. There is no separate duration parameter — a 10-second audio file produces a 10-second video; a 60-second audio file produces a 60-second video.

What inputs are required?

A portrait image and an audio file are both required. An optional text prompt can refine the animation style, emotion, or camera framing — e.g., "calm and authoritative, slight head movement, warm lighting."

How much does Kling Avatar V2 cost in credits?

Standard mode (720p) costs 28 credits per second. Pro mode (1080p, 48 fps) costs 50 credits per second. A 15-second voiceover renders at ~420 credits in Standard or ~750 credits in Pro.

What audio formats are supported?

MP3, WAV, M4A, and AAC are supported. The audio file must be ≤ 5 MB. For longer voiceovers, compress the file or split into multiple shorter generations.

Can I generate the audio on AI Compare Hub first?

Yes — generate the voiceover with AI Compare Hub's Text-to-Speech tool first, then select the resulting audio from your gallery when configuring Kling Avatar V2. The two-step workflow lets you craft the script, pick a voice, and animate the portrait without leaving the platform.

How can I use Kling Avatar V2 on AI Compare Hub?

Open the AI Video Generator's Image-to-Video panel and pick "Kling Avatar V2" from the model picker. Upload (or pick from gallery) a portrait image and an audio file, choose Standard or Pro mode, optionally add a prompt for emotion/framing, and generate.

Key Parameters

For the Use of This Model

Kling Avatar V2

Animate a portrait by pairing it with a voice or song clip. Kling Avatar V2 produces lip-synced facial animation up to 1080p at 48 fps and works with realistic photos, illustrations, cartoons, and animals. The output duration matches the length of the supplied audio file (no manual duration parameter needed).

Inputs

  • image - Portrait reference (.jpg, .jpeg, .png; max 10MB; aspect 1:2.5 to 2.5:1).
  • audio - Driving audio (.mp3, .wav, .m4a, .aac; max 5MB).
  • prompt - Optional text refining style, emotion, or camera framing.
  • mode - std (720p, lower cost) or pro (1080p, higher quality).

Try Kling Avatar V2 on AI Compare Hub

Generate with Kling Avatar V2 directly on AI Compare Hub. Compare results side by side with other leading AI models using the same prompt.