Skip to main content

Clean Audio API

Clean up noisy speech recordings with VEED's Clean Audio API. Send a file, get clear speech audio back in FLAC or WAV.

The Clean Audio API removes background noise from speech. Send an audio or video file, and get clean speech audio back. It works with any language.


Endpoint

POST /v1/clean-audio

The job is accepted immediately: you get 202 Accepted with a job_id and a status of PROCESSING. Poll GET /v1/clean-audio/{job_id} until the status is COMPLETED or FAILED.

You can also try it without writing any code in the playground.


Inputs

Input

Required

Default

Notes

audio_url

Yes

—

Public http(s) URL of the recording. Any audio or video file, up to 30 minutes and 512 MB. A video's audio track is used, and multi-channel audio is mixed down to mono.

strength

No

0.874

Between 0 and 1. How much of the original is allowed to remain underneath speech — the suppression floor is 1 − strength. Lower values keep more room tone behind the voice. Silence between words is always fully cleaned.

target_lufs

No

-19

Between −40 and −8. Integrated loudness of the output in LUFS (ITU-R BS.1770); true peak is capped at −1.1 dBTP. Ignored when normalize_loudness is false.

normalize_loudness

No

true

Set to false to skip loudness normalization and keep the input level.

output_format

No

flac

flac or wav. Both are 48 kHz mono 16-bit. FLAC is lossless at about half the size of WAV.


Supported files

Anything ffmpeg can read:

  • Audio: MP3, WAV, M4A/AAC, OGG/Opus, FLAC

  • Video: MP4, MOV, WebM

Up to 30 minutes and 512 MB per file.


Output

One denoised audio file at 48 kHz mono 16-bit, the same duration as the input. FLAC by default, or WAV on request.


Pricing and processing time

  • Pricing: $0.0125 per minute of input, rounded up to the next whole minute. There's a minimum of 1 minute ($0.0125) per call. A 4-minute 10-second file is billed as 5 minutes, so $0.0625.

  • Processing time: roughly 0.2x to 0.5x the length of the audio. A 10-minute file takes 2 to 5 minutes.

The response that creates the job includes credits_estimated, a quote given before the job ran. Once the job is COMPLETED, credits_charged gives the actual charge. The estimate isn't the price — the charged amount can be higher or lower.


FAQs

  • Q: Does it work on music?

    • No. It's built for speech. Music and sound effects are treated as background noise and removed.

  • Q: Which languages are supported?

    • All of them.

  • Q: Can I send a video file?

    • Yes. The audio track is used, and you get the cleaned audio back as FLAC or WAV — not a video.

  • Q: My file is longer than 30 minutes or over 512 MB. What do I do?

    • Split it into smaller files and send each one separately.

  • Q: What happens to stereo or multi-channel audio?

    • It's mixed down to mono before processing, and the output is always mono.

  • Q: Is this the same as Clean Audio in the editor?

    • Yes, the same model. The API gives you programmatic access and finer control over strength and loudness. See how to use the Clean Audio tool for the in-editor version.


Tips for the best results

  • Only use it on speech. If you want to keep music or sound effects, don't run the track through Clean Audio.

  • If the voice sounds processed or thin, lower strength. That leaves more of the original room tone underneath the speech, which usually sounds more natural.

  • Use target_lufs to get consistent loudness across files — useful when you're cleaning a batch that will play back-to-back.

  • Set normalize_loudness to false if you're handling levels yourself further down your pipeline.

  • Use wav if your pipeline can't read FLAC, otherwise stick with the default — it's the same audio at about half the size.

Did this answer your question?