- 25comments
- 23comments
- 103comments
- 54comments
- 8comments
- 37comments
- 47comments
- 146comments
- 199comments
- 804comments
- 116comments
- 12comments
- 85comments
- 1036comments
- 3comments
- —discuss
- 419comments
- 275comments
- 10comments
- 69comments
- 590comments
- 233comments
- 176comments
- 65comments
- 540comments
- 160comments
- 194comments
- 32comments
- 120comments
- 442comments
Of possible interest is my Emotive Audiobook Creator, KeenLore. Here's a video of the web app showing how it works:
https://www.youtube.com/watch?v=WAeHgE94rVo
Locally hosted, no cloud, no tokens to pay. Reads a book using a full cast of characters. Quotation attribution detection (for my novel) is at 97.2% accuracy (485/499 quotes identified and assigned correctly). The autofill of character voice descriptions uses the prose to determine how the character sounds.
<grumble grumble people putting in links they expect you to follow to arbitrary goatse youtube videos for all I know>
The title of the video is 'KeenLore - Emotive Audiobook Creator Demo' and it appears to be a web UI and some local stack that reads text files.
It's cool technology and I read a lot of audiobooks, even hundreds of hours of TTS. I feel like my brain can fill in the character voices from the text - on the page it's not like they're different fonts.
I understand audiobook narrators often do it, and that's fun. But it's not so critical in my opinion
Our imaginations and minds continue to rot under the weight of endless and effortless entertainment
I guess voice cloning is widely enough available now from other providers that Google are no longer hesitant to ship it.
They probably do something similar to GPT-Live where they expect a given voice profile to send them a sample saying 'This is the owner of this voice and I consent for synthetic samples to be made of it'
and/or local voice cloning is good enough as is so Google doesn't grant a uniquely liable ability?
"Don't be evil... unless other companies are doing it first"
voice cloning is a tool, it is not necessarily evil, even though the scenarios it can be used for nefarious purposes outnumber the legitimate ones.
Gemini 3.8 "Flash" says hello
"Super tinny monotone robotic voice" does not sound neither tinny nor monotone. Compared to what TTS from 90s sounded like. Or even how actors impersonated robots in movies. Has the model been eating too much hype DJs?
I direct my own extended daydream Star Trek fanfic (okay, I'm on season 2 episode 17) and recently I looked to see if I could have each scene file be read aloud a la an audiobook or radio drama.
Getting GPT-Live to have unique enough voices and to be expressive with how I imagine the voices going in my head is hard to direct, there's not enough control there.
So this Gemini 3.8 specific large voice library and ability to tightly control (if you are willing to write a script) is nice to find, and while I'm not sure which of the 5,286 Gemini products this is, nor how to onboard and get started feeding this my own text files, nor what training will happen to my data if I did somehow use it, I love that the state of the industry is such that Google can do this and release it publicly, because that means eventually an equivalent product can come from someone else and be used locally / confidently that the generated audio or inputs won't be retained and misused.
Original series or Next Generation?
You might find this interesting, it seems to be exactly what you need to make an audio drama: https://github.com/Finrandojin/alexandria-audiobook
Also the Qwen3-TTS demo is cool, you can describe the voice you want: https://huggingface.co/spaces/Qwen/Qwen3-TTS
I came across both on this subreddit, it's very active: https://www.reddit.com/r/TextToSpeech/
I'm personally using this locally: https://github.com/mateogon/pdf-narrator (it's a Python frontend for Kokoro) on my M1 Macbook Air (from 2020, with 8GB RAM) and it's incredible. I make my own audiobooks now - for free!
My favorite voice is am_michael and here's a sample: https://voicerankings.com/voice/kokoro-82M/male/am_michael/s...
Great. Now in additional to AI email responses I will get AIs impersonating my contacts on the phone too. Lovely.
Sorry if this is in that article, but I am on my phone and can't see it. How much would this cost to batch generate an audiobook? Right now I just listen to things in the 11 labs app which is free, but I would rather just generate audio files.
Roughly 5-10$ for 10h, assuming you few-shot it.
Price per hour:
- 3.8 Flash TTS, standard: $0.81
- 3.8 Flash TTS, batch: $0.41
- 3.8 Flash‑Lite TTS, standard: $0.54
- 3.8 Flash‑Lite TTS, batch: $0.27
Would be great if this would power the Google Books app feature. The voice system there is pretty out of date.
Yes, I am quite disappointed by seeing all this cool AI stuff and yet the same Play Books. Come on, it is the best place to apply AI, in my opinion.
the ratio of new voice models i see on hackernews to the number actually deployed in any product i use is approximately infinity.
I’ve never found a TYS that does convincing British accents.
They all sound like Americans putting in their best fake British accent.
Related, for embedding small models, this lib is incredible.
Having a voice under 1Mo is crazy, even if it sounds robotic.
https://tts.ampixa.com/sanoTTS/
Great price at least until December 31 too.
Seems like voice actors are safe for now. This is technologically incredible, but the results are really not very good, and usually not particularly close to the prompt. In basically all of these examples some core part of the prompt is completely ignored.