Kling O1 on AI Compare Hub

Kling O1 by KwaiVGI performs context-aware video editing using natural language — swap subjects, change backgrounds, restyle scenes, or insert specific elements from reference images while preserving the original clip's timing and motion. Output duration matches the input video (3–10 seconds). Resolution spans 720p through 4K, with optional element-reference inputs for precise visual control.

What you can create

Why creators choose Kling O1

How to generate your first edit

  1. Upload your reference video. The clip you want to edit (3–10 seconds). The original timing, motion, and camera move will be preserved in the output.
  2. Describe the edit in plain language. Be specific about what to change and what to keep: "Change the man's blue shirt to a red leather jacket. Keep the lighting, background, and motion exactly the same." Optionally add reference images of specific elements you want inserted (props, characters, logos).
  3. Configure mode and elements. Pick Standard (720p) or Pro (1080p). If you've added element references, you'll be quoted at the higher elements tier (57 or 76 cr/sec). Generate and review.

Common questions

What is Kling O1?

Kling O1 is KwaiVGI's context-aware video editing model, available via the Replicate API. It performs natural-language edits — subject/outfit swaps, background changes, style transfers, and element insertions — while preserving the original clip's timing and motion. Output duration matches the input video (3–10 seconds).

How much does Kling O1 cost in credits?

Standard (720p): 42 cr/sec for prompt-only edits, 57 cr/sec when element references are added. Pro (1080p): 50 cr/sec for prompt-only, 76 cr/sec with element references. The credit quote uses your reference video's actual duration.

When are higher pricing tiers used?

When you provide element references (the elements array — reference images of specific props, characters, or assets you want inserted), the model uses the higher "video-input + elements" pricing tier. Standard mode without elements is the cheapest path; Pro mode with elements is the most expensive.

How is the output duration determined?

Output duration matches the input video's duration (3–10 seconds). There is no separate duration parameter to set — the credit quote on the platform pulls the reference video's true length once it's uploaded.

What kinds of edits work best?

Subject swaps, outfit changes, background replacement, prop insertion, and style/look transfers are all strong use cases. Edits that preserve the original scene's structure (motion, framing, lighting) work most reliably. Highly destructive edits that change everything at once may be better handled by re-generating from scratch.

Can I keep the audio from my source video?

Yes — toggle "keep audio" to preserve the audio track from the input video in the edited output. This keeps dialogue, ambience, and music in sync with the original timing.

How can I use Kling O1 on AI Compare Hub?

Open the AI Video Generator's Video-to-Video panel and pick "Kling O1" from the model picker. Upload your source video, write a natural-language edit description, optionally add element-reference images for specific assets, choose Standard or Pro mode, then generate.

Key Parameters

For the Use of This Model

Kling O1

Context-aware video editing model: takes a source video plus a text instruction and outputs an edited video that preserves the original motion and timing. Supports up to 4 combined reference inputs (element and style images) which can be referenced in the prompt as @Element1, @Element2, @Image1, @Image2.

Inputs

  • reference_video - The source video to edit (.mp4/.mov/.webm/.m4v/.gif; 3-10s; max 200MB).
  • prompt - Required natural-language editing instruction.
  • elements - Optional structured array of element references (up to 4 combined elements + style images).
  • mode - std or pro.
  • keep_audio - Preserve original audio from the reference video.

Try Kling O1 on AI Compare Hub

Generate with Kling O1 directly on AI Compare Hub. Compare results side by side with other leading AI models using the same prompt.