Skip to content
Catalog

cn8/studio-speech-clearance

Send an audio file (wav, mp3, m4a, flac, ogg) or a video file (mp4, mov, mkv, avi, webm)

Learn more about Speech Clearance
Audio & Voiceasync (job)Instant previewaudioprocessingstudio

Try it

Source type
Audio or video with noisy speech

Result

The result shows up here

Upload your audio or video on the left and hit Run.

Overview

Send an audio file (wav, mp3, m4a, flac, ogg) or a video file (mp4, mov, mkv, avi, webm) and get back a clean, enhanced version of the speech. A neural speech-restoration model removes hiss, hum, and background noise and brings the voice forward, with smooth, click-free, high-quality output. For audio input you get an enhanced audio file; for video input you get the same video back with the cleaned speech track in place — the picture is untouched. This is an async job: submit it, poll until it's ready, then download the result.

Key capabilities

Audio and video input

Accepts audio files (wav, mp3, m4a, flac, ogg) and video files (mp4, mov, mkv, avi, webm). For video, the audio track is enhanced and put back into the same video.

Noise removal and speech enhancement

Removes hiss, hum, and background noise and brings the voice forward for clear, intelligible speech.

High-quality, click-free output

The enhanced audio is delivered at high quality with smooth, click-free results from start to finish.

Video soundtrack enhancement

For video input, only the audio is replaced — the video itself is left exactly as it was. Output is a single MP4.

When to use it

Podcast and interview cleanup

Remove hiss, hum, and ambient noise from voice recordings before publishing.

Voiceover and narration

Enhance dialogue or narration tracks for video production.

Video soundtrack speech enhancement

Clean the speech in a video without touching the picture.

Input & output

input

Audio URL (wav, mp3, m4a, flac, ogg) or video URL (mp4, mov, mkv, avi, webm)

JSON body

output

Enhanced speech audio file URL, or the original video with enhanced audio (async job result)

JSONmedia URL

Guides & tips

How it works

  • Submit an audio or video URL. For video, the audio track is enhanced and put back into the same video — the picture is left untouched.
  • A neural speech-restoration model removes hiss, hum, and background noise and brings the voice forward, with smooth, click-free, high-quality output.
  • Because it processes the whole file, this is an async job — submit it, poll until it's ready, then download the result.

Supported formats

  • Audio: wav, mp3, m4a, flac, ogg
  • Video: mp4, mov, mkv, avi, webm
  • For video, only the audio is enhanced; the picture is left as-is.

Specs

Latency
Async; depends on audio duration
Async
true
Rate Limit
Per API key
Max Input
Duration-dependent; longer files take longer

Schema

Request body

mediaUrlrequiredstring

URL of source audio (wav, mp3, m4a, flac, ogg) or video (mp4, mov, mkv, avi, webm)

Pricing

Billed per second of source audio. Two-part tariff: 5 credits base fee per job + 2 credits/second, minimum 5 billed seconds. Failed jobs are not billed.

ServiceUnitPrice
Speech Clearancesecond2 credits/second

For video input, billing is based on the audio track duration.

FAQ

Related models