VEED has two lipsync APIs, and which one you need depends on what you're starting with. Fabric 1.0 turns a still image into a talking video. Lipsync 2.0 re-syncs footage you already have to a new audio track. This article covers both, and how to get the best results from either.
Which lipsync API should I use?
Start from what you have. A photo or illustration: use Fabric 1.0. Existing video footage: use Lipsync 2.0.
| Fabric 1.0 | Lipsync 2.0 |
You send | An image + an audio track | A video + an audio track |
You get | A talking video (MP4) | The same video, re-synced to the new audio (MP4) |
Use it for | Talking avatars, mascots, personalized videos | Dubbing, translation, replacing a voice track |
Resolution | 480p or 720p | Same as source |
Price | $0.08/s (480p), $0.15/s (720p) | $0.07/s |
Fabric 1.0 API (image to video)
Fabric 1.0 turns a still image and an audio track into a talking video, with the mouth moving in sync with the speech. One call does the whole job — no separate face detection or audio steps.
Endpoint
POST /v1/fabric-1.0, then poll GET /v1/fabric-1.0/{job_id}
Inputs
Input | Required | Notes |
| Yes | Public URL of the image. Must contain a face. |
| Yes | Public URL of the audio to speak. |
| Yes |
|
Output, pricing and processing time
Output: one MP4 video.
Pricing: per second of generated video — $0.08/s at 480p, $0.15/s at 720p. A 10-second 720p video costs $1.50.
Processing time: usually a minute or two.
Fabric 1.0 FAQs
Q: Can I use any image?
Yes, as long as it has a clear face and you own it or have permission. Photos, illustrations, mascots and stylized characters all work.
Q: How long can the video be?
Longer audio fails with
audio_too_long.
Q: Can I send text instead of audio?
No. Generate the audio first with a text-to-speech tool, then send its URL.
Open Fabric 1.0 docs and playground at api.veed.io/models/fabric-1.0
Lipsync 2.0 API (video to video)
Lipsync 2.0 takes an existing video and a new audio track, and re-renders the video so the speaker's mouth matches the new audio.
Endpoint
POST /v1/lipsync-2.0, then poll GET /v1/lipsync-2.0/{job_id}
Inputs
Input | Required | Notes |
| Yes | Public URL of the source video. |
| Yes | Public URL of the new audio track. |
Output, pricing and processing time
Output: one MP4 video.
Pricing: $0.07 per second of generated video. A 10-second video costs $0.70.
Processing time: usually a minute or two.
Lipsync 2.0 FAQs
Q: Can I use it to translate a video?
Yes. Create the translated audio first, then send it with the original video.
Q: Does it work with more than one speaker?
No.
Open Lipsync 2.0 docs and playground at api.veed.io/models/lipsync-2.0
Tips for the best lipsync results
These apply to both models
Use clean audio. One voice, no music or background noise.
Show the face clearly. Front-facing, well lit, and not covered by hands, microphones or hair.
Use a sharp, high-resolution source.
For Lipsync 2.0, avoid fast movement. Footage where the speaker turns away or moves quickly gives worse results.
Test a short clip first before running long jobs.
