Skip to main content

Lipsync APIs: Fabric 1.0 and Lipsync 2.0

Generate talking videos from an image with Fabric 1.0, or re-sync existing footage to new audio with Lipsync 2.0. Endpoints, pricing, and inputs.

VEED has two lipsync APIs, and which one you need depends on what you're starting with. Fabric 1.0 turns a still image into a talking video. Lipsync 2.0 re-syncs footage you already have to a new audio track. This article covers both, and how to get the best results from either.


Which lipsync API should I use?

Start from what you have. A photo or illustration: use Fabric 1.0. Existing video footage: use Lipsync 2.0.

Fabric 1.0

Lipsync 2.0

You send

An image + an audio track

A video + an audio track

You get

A talking video (MP4)

The same video, re-synced to the new audio (MP4)

Use it for

Talking avatars, mascots, personalized videos

Dubbing, translation, replacing a voice track

Resolution

480p or 720p

Same as source

Price

$0.08/s (480p), $0.15/s (720p)

$0.07/s


Fabric 1.0 API (image to video)

Fabric 1.0 turns a still image and an audio track into a talking video, with the mouth moving in sync with the speech. One call does the whole job — no separate face detection or audio steps.

Endpoint

POST /v1/fabric-1.0, then poll GET /v1/fabric-1.0/{job_id}

Inputs

Input

Required

Notes

image_url

Yes

Public URL of the image. Must contain a face.

audio_url

Yes

Public URL of the audio to speak.

resolution

Yes

480p or 720p

Output, pricing and processing time

  • Output: one MP4 video.

  • Pricing: per second of generated video — $0.08/s at 480p, $0.15/s at 720p. A 10-second 720p video costs $1.50.

  • Processing time: usually a minute or two.

Fabric 1.0 FAQs

  • Q: Can I use any image?

    • Yes, as long as it has a clear face and you own it or have permission. Photos, illustrations, mascots and stylized characters all work.

  • Q: How long can the video be?

    • Longer audio fails with audio_too_long.

  • Q: Can I send text instead of audio?

    • No. Generate the audio first with a text-to-speech tool, then send its URL.

Open Fabric 1.0 docs and playground at api.veed.io/models/fabric-1.0


Lipsync 2.0 API (video to video)

Lipsync 2.0 takes an existing video and a new audio track, and re-renders the video so the speaker's mouth matches the new audio.

Endpoint

POST /v1/lipsync-2.0, then poll GET /v1/lipsync-2.0/{job_id}

Inputs

Input

Required

Notes

video_url

Yes

Public URL of the source video.

audio_url

Yes

Public URL of the new audio track.

Output, pricing and processing time

  • Output: one MP4 video.

  • Pricing: $0.07 per second of generated video. A 10-second video costs $0.70.

  • Processing time: usually a minute or two.

Lipsync 2.0 FAQs

  • Q: Can I use it to translate a video?

    • Yes. Create the translated audio first, then send it with the original video.

  • Q: Does it work with more than one speaker?

    • No.

Open Lipsync 2.0 docs and playground at api.veed.io/models/lipsync-2.0


Tips for the best lipsync results

These apply to both models

  • Use clean audio. One voice, no music or background noise.

  • Show the face clearly. Front-facing, well lit, and not covered by hands, microphones or hair.

  • Use a sharp, high-resolution source.

  • For Lipsync 2.0, avoid fast movement. Footage where the speaker turns away or moves quickly gives worse results.

  • Test a short clip first before running long jobs.

Did this answer your question?