Legal · Sample Media
Sample Media Attributions
Every audio, video, and image file bundled with SonicVox® for tool demos is listed here along with its license and upstream source. All samples are commercial-use safe (CC0, Pixabay Content License, Pexels License, LibriVox public domain, CC-BY with attribution, or first-party SonicVox assets).
For open-source library credits, see the credits page.
Speech-to-Text models
10 modelsOur transcription & diarization stack is built on these open models, all licensed for commercial use. CC-BY-4.0 and NVIDIA Open Model License models are credited here per their attribution terms.
NVIDIA Canary-1B v2
https://huggingface.co/nvidia/canary-1b-v2NVIDIA Parakeet TDT 0.6B v3
https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3Cohere Transcribe (03-2026)
https://huggingface.co/CohereLabs/cohere-transcribe-03-2026IBM Granite Speech 4.1-2B
https://huggingface.co/ibm-granite/granite-speech-4.1-2bFunAudioLLM SenseVoice Small
https://huggingface.co/FunAudioLLM/SenseVoiceSmallUseful Sensors Moonshine (streaming-medium)
https://huggingface.co/UsefulSensors/moonshine-streaming-mediumpyannote speaker-diarization-3.1
https://huggingface.co/pyannote/speaker-diarization-3.1NVIDIA Streaming Sortformer Diarizer 4spk v2.1
https://huggingface.co/nvidia/diar_streaming_sortformer_4spk-v2NVIDIA TitaNet-Large (speaker re-ID)
https://huggingface.co/nvidia/speakerverification_en_titanet_largeOur commercial posture
Public domain / CC0 / SonicVox-generated only. No third-party stock licenses and no attribution obligations.
Audiogram Generator
Audio track (spoken narration)
Expressive storytelling narration — drives the animated waveform. Swap in your own podcast clip for production.
/samples/audiogram/speech.mp3audio/mpeg19sCover image
1080x1080 clean indigo-to-magenta cover background for audiogram overlays.
/samples/audiogram/cover-clean.jpgimage/jpegFormat Converter
Uncompressed WAV source
44.1 kHz stereo WAV of clean spoken narration. Try converting to MP3, Opus, FLAC, M4A, or AIFF and compare file sizes.
/samples/convert/source.wavaudio/wav16sMusic Ducker
Trigger track (spoken narration)
A clear spoken narration used as the trigger track — the music bed ducks under it whenever this speaks.
/samples/duck/voice.mp3audio/mpeg16sBackground music bed
Instrumental march that gets ducked under the trigger track. Volume drops when the trigger is active.
/samples/duck/music.mp3audio/mpeg20sAudio Extractor
Launch countdown with audio
A 20-second NASA SLS launch sequence from T-12 seconds through liftoff, with countdown voice and engine roar on the audio track. Extract the audio as WAV, MP3, FLAC, or M4A.
/samples/extract/source.mp4video/mp420sAudio Fade
Instrumental clip
20-second instrumental march cut with hard edges — fade-in and fade-out effects are very audible at the edges.
/samples/fade/instrumental.mp3audio/mpeg20sAudio Merger
Clip A
Opening fanfare of Sousa's Stars and Stripes Forever. Will be joined with clip B, with an optional crossfade between them.
/samples/merge/clip-a.mp3audio/mpeg12sClip B
A second march (Sousa's Semper Fidelis) with a different energy. Try the crossfade slider at 1-2 seconds for a smooth blend.
/samples/merge/clip-b.mp3audio/mpeg12sMetadata Injector
Blank-slate MP3
A ~30-second spoken-word MP3 with no metadata whatsoever. Add a title, artist, album, year, genre, cover art, or chapter markers.
/samples/metadata/untagged.mp3audio/mpeg33sAlbum cover art
1000x1000 generated album art ("Midnight Signals — SonicVox Sessions"). Embed it as the cover image when tagging the sample MP3.
/samples/metadata/cover-art.pngimage/pngLoudness Normalizer
Under-leveled clip
Spoken narration deliberately leveled to ~-30 LUFS. A normalize pass boosts it dramatically to hit Spotify or YouTube loudness specs.
/samples/normalize/quiet-audio.mp3audio/mpeg30sFull-volume music
Brass-band march pushed to ~-9 LUFS with peaks at 0 dBFS. Normalize pulls it DOWN to streaming specs instead of boosting.
/samples/normalize/loud-music.mp3audio/mpeg30sVideo Resizer
Earth from the ISS (16:9)
A 15-second 1280x720 view of Earth from the International Space Station. Try Smart Crop to convert to TikTok / Reels 9:16 or square 1:1 while keeping the horizon framed.
/samples/resize/landscape.mp4video/mp415sPitch/Speed Shifter
Clean speech clip
15-second clean spoken narration. Shift up/down by semitones — formant-aware mode keeps the voice natural across big shifts.
/samples/shift/speech.mp3audio/mpeg15sAudio Splitter
Long-form music track
A 2-minute brass-band march for time-interval or equal-parts splitting.
/samples/split/long-form.mp3audio/mpeg120sStem Splitter
Song with vocals and band
Enrico Caruso singing "Over There" (1918) with orchestral accompaniment. Demucs separates it into 4 stems (vocals / drums / bass / other) for a live remix preview.
/samples/stem-split/mixed-song.mp3audio/mpeg30sThumbnail Extractor
Artemis I launch montage
A 16-second montage cut from NASA's Artemis I recap — night pad, engine ignition, ascent plume, and Orion in space. Plenty of hard scene changes for Smart Pick, Scene Cut, or a contact sheet.
/samples/thumbnail/source.mp4video/mp416sSilence Trimmer
Clip with silent gaps
A spoken narration with deliberate leading, middle, and trailing silences. Silence-removal trims the dead air automatically.
/samples/trim/speech-with-pauses.mp3audio/mpeg21svolume
Under-leveled clip
Spoken narration leveled to about -30 LUFS. A gain pass makes the change obvious end to end.
/samples/normalize/quiet-audio.mp3audio/mpeg30sBrand Overlay
Clean video (Earth from orbit)
A 15-second clean orbital Earth clip from the International Space Station. Add your logo with the image watermark, positioned anywhere on the frame.
/samples/watermark/source.mp4video/mp415sSonicVox wordmark
598x150 transparent PNG wordmark with a waveform glyph — white with a subtle dark outline so it stays readable on light or dark footage.
/samples/watermark/sonicvox-wordmark.pngimage/png