The new SDK bundles the Niagara automatic speech recognition (ASR) and Nith text-to-speech (TTS) model families into a single API. By processing voice input and output locally, the system bypasses network dependencies, ensuring consistent performance even in environments with poor connectivity. According to CEO Kevin Conley, the architecture prioritizes immediate response times, a critical requirement for functional edge-based voice applications.
The current release offers a Python library with C and Java bindings expected shortly. Developers can integrate English, Spanish, Mandarin, Japanese, and Korean, with the software supporting Linux x86-64, Linux ARM64, and Android ARM64 platforms. The models are designed for efficiency, with Niagara ASR producing initial text in 115 milliseconds and Nith TTS generating audio in 147 milliseconds on embedded CPUs. For specialized use cases, the SDK includes options for voice cloning and custom vocabulary management, allowing companies to tailor pronunciation and recognition to specific brand or domain requirements without needing to retrain the underlying models.





Comments (0)
No comments yet. Be the first!