The Clean Audio API removes background noise from speech. Send an audio or video file, and get clean speech audio back. It works with any language.
Endpoint
POST /v1/clean-audio
The job is accepted immediately: you get 202 Accepted with a job_id and a status of PROCESSING. Poll GET /v1/clean-audio/{job_id} until the status is COMPLETED or FAILED.
You can also try it without writing any code in the playground.
Inputs
Input | Required | Default | Notes |
| Yes | — | Public http(s) URL of the recording. Any audio or video file, up to 30 minutes and 512 MB. A video's audio track is used, and multi-channel audio is mixed down to mono. |
| No |
| Between 0 and 1. How much of the original is allowed to remain underneath speech — the suppression floor is 1 − strength. Lower values keep more room tone behind the voice. Silence between words is always fully cleaned. |
| No |
| Between −40 and −8. Integrated loudness of the output in LUFS (ITU-R BS.1770); true peak is capped at −1.1 dBTP. Ignored when |
| No |
| Set to |
| No |
|
|
Supported files
Anything ffmpeg can read:
Audio: MP3, WAV, M4A/AAC, OGG/Opus, FLAC
Video: MP4, MOV, WebM
Up to 30 minutes and 512 MB per file.
Output
One denoised audio file at 48 kHz mono 16-bit, the same duration as the input. FLAC by default, or WAV on request.
Pricing and processing time
Pricing: $0.0125 per minute of input, rounded up to the next whole minute. There's a minimum of 1 minute ($0.0125) per call. A 4-minute 10-second file is billed as 5 minutes, so $0.0625.
Processing time: roughly 0.2x to 0.5x the length of the audio. A 10-minute file takes 2 to 5 minutes.
The response that creates the job includes credits_estimated, a quote given before the job ran. Once the job is COMPLETED, credits_charged gives the actual charge. The estimate isn't the price — the charged amount can be higher or lower.
FAQs
Q: Does it work on music?
No. It's built for speech. Music and sound effects are treated as background noise and removed.
Q: Which languages are supported?
All of them.
Q: Can I send a video file?
Yes. The audio track is used, and you get the cleaned audio back as FLAC or WAV — not a video.
Q: My file is longer than 30 minutes or over 512 MB. What do I do?
Split it into smaller files and send each one separately.
Q: What happens to stereo or multi-channel audio?
It's mixed down to mono before processing, and the output is always mono.
Q: Is this the same as Clean Audio in the editor?
Yes, the same model. The API gives you programmatic access and finer control over strength and loudness. See how to use the Clean Audio tool for the in-editor version.
Tips for the best results
Only use it on speech. If you want to keep music or sound effects, don't run the track through Clean Audio.
If the voice sounds processed or thin, lower
strength. That leaves more of the original room tone underneath the speech, which usually sounds more natural.Use
target_lufsto get consistent loudness across files — useful when you're cleaning a batch that will play back-to-back.Set
normalize_loudnesstofalseif you're handling levels yourself further down your pipeline.Use
wavif your pipeline can't read FLAC, otherwise stick with the default — it's the same audio at about half the size.
