VTS

No phone app

5.9No. 22 of 28
in AI Sound Effect Generators
  • Recognised40% of the score32
  • Phone app26% of the score0
  • Documented20% of the score82
  • Free plan14% of the score30
Free plan
No
Runs on
self-hosted

Summary

VTS is ranked #22 of 28 in AI sound effect generators on Samsung Mobile US Press. It runs on Self-hosted.

Compared on AI sound effect generators

Free plan
Yesgithub.com
Text-to-sound
Yesgithub.com
Reference input
Yesgithub.com
Download formats
wavgithub.com
Commercial use
uncleargithub.com
API access
Nogithub.com

Facts

Purpose
VTS generates sound effects from a short vocal or audio sketch combined with a text prompt.github.com · 4 Oct 2026
How it works
It uses voice conditioning derived from dynamic audio features alongside text conditioning from a prompt.github.com · 4 Oct 2026
Intended users
The checkpoint is intended for research and creative sound-effect generation from vocal sketches or short audio sketches plus text prompts.huggingface.co · 4 Oct 2026
Local inference
The repository provides an inference-only package with local model and runtime code.github.com · 4 Oct 2026
Checkpoint
The inference code uses the pretrained checkpoint named dynamic_v3_0415.ckpt, which is available from the linked Hugging Face model repository.github.com · 4 Oct 2026
Generation
Generated WAV files are written to a chosen output directory, and the default output duration follows the input audio duration unless a duration is specified.github.com · 4 Oct 2026
Hardware
The quick-start inference example specifies the CUDA device.github.com · 4 Oct 2026
Limits
The model is optimized for short sound-effect style clips, and output quality depends on the checkpoint, input audio, prompt text, and sampling settings.huggingface.co · 4 Oct 2026
Training
Training code and dataset manifests are not included in the inference package.github.com · 4 Oct 2026
Integrations
The inference code uses google/flan-t5-base as its text encoder and includes local vocoder code.github.com · 4 Oct 2026
API availability
The model is not packaged as a Hugging Face Inference API pipeline and is not deployed by an Inference Provider.huggingface.co · 4 Oct 2026
License
The project and model checkpoint are listed under the MIT License.github.com · 4 Oct 2026
Support
The repository says to contact the maker at [email protected] with questions.github.com · 4 Oct 2026
Voice conditioning
Voice conditioning uses dynamic features derived from spectral centroid, RMS, and chroma-index signals.github.com · 4 Oct 2026
Text encoder
The inference path encodes text prompts with google/flan-t5-base.github.com · 4 Oct 2026
Output
Generated audio files are written as WAV files to the chosen output directory.github.com · 4 Oct 2026
Generation length
Output duration defaults to the input audio duration, can be configured, and the checkpoint is tuned for short sound-effect clips.github.com · 4 Oct 2026
Sampling
Sampling uses a local ODE solver and typically runs 64 steps with CFG scale 3.0.github.com · 4 Oct 2026
Deployment
The repository provides an inference-only package with local model and runtime code; training code and datasets are excluded.github.com · 4 Oct 2026
Hardware setup
The documented local requirements pin PyTorch and torchaudio CUDA 12.4 builds, with instructions to install matching builds for other CUDA drivers.github.com · 4 Oct 2026
Checkpoint access
If the Hugging Face repository requires authentication, inference setup uses an HF_TOKEN environment variable.github.com · 4 Oct 2026

Best VTS alternatives

See all 12

Where it ranks on Samsung Mobile US Press

Is VTS yours?

Claim it for free: prove the domain, then correct facts, plans and screenshots. An editor reviews every change.

Sources