0.17
Q&A about TTS (Text‑to‑Speech)
Q: Why was the TTS model changed?
A: The official Qwen3TTS model is quite large—the model alone is 4GB, and even when compressed to the limit, it still takes up 7GB.
Q: I checked the VoxCPM2 official page and saw that it supports CUDA. Why isn't it used in this software?
A: The official model is also 4GB. A third‑party optimized version, including its dependencies, totals about 3.4GB, and the audio quality is almost identical to the original.
Q: Why do you try to reduce the program file size?
A: Overseas users experience slow download speeds from Chinese cloud storage services, so I try to compress the installer as much as possible.
The only drawback is that OneDrive only hosts the latest version.
Q: What are the differences between Qwen3TTS and VoxCPM2?
A: They are largely comparable. Qwen3TTS offers a better user experience and faster generation speed, whereas VoxCPM2 supports a wider range of languages, though its output may sometimes sound unnatural.
Q&A about RVC (Voice Conversion)
Q: Why is RVC only provided in the DirectML version?
A: The CUDA version only supports NVIDIA GPUs. DirectML can leverage hardware acceleration on any device running Windows 10/11 with DirectX 12 support, so I prioritized making a universal version. If there is demand for the CUDA version, feel free to contact me, and I will continue maintaining it as needed.
Q&A about the Software
Q: Why were the two previous Beta versions removed?
A: The old versions only supported CPU mode and had a large file size, so they were removed. (If you don't believe that, just assume it was to save cloud storage space.)
Version 0.15 will remain available, as it is the last version to use the official Qwen3TTS Safetensors model.
Some users may prefer the original, uncompressed synthesis quality over the optimized model.
However, this version sounds almost identical to the official version, so I still recommend trying the official original first.Feel free to ignore this if you are low on storage.
The official Safetensors model will not be updated in the future.
Q: If you want to reduce the program size, why not let users download the TTS core themselves?
A: Because an out‑of‑the‑box experience is the best.
Q: I'm concerned that my chat history and API keys might be leaked.
A: They are both encrypted. However, the encryption mechanism is open‑source on GitHub, so please always download the official version and avoid any unknown third‑party builds.
Other Questions
Q: Can I repost or share this software on other platforms?
A: Yes, you may repost or share it. When doing so, please comply with all requirements in the Character Usage Guide and Terms of Use (Voiceger:Zundamon). Posting pirated links is prohibited.
You do not need my permission to repost, but please include the project URL and the author's name in your post—preferably both.
Q: Why did you modify Voiceger?
A: The official version is simple and easy to use, but the UI is somewhat plain and has no Chinese‑language localization. I understand English, but not all users do, so I modified it to address these shortcomings.
Q: Can I repost or reuse the system prompt you provided? A: Yes, you can. When using the prompt, please comply with all requirements in the Character Usage Guidelines and the Terms of Use (Voiceger:Zundamon), and make sure to clearly credit the source.