Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Fixing the Portobello Police Station Clock(pointinthecloud.com)
    25comments
  2. Gemini 3.8 text-to-speech says hello(blog.google)
    23comments
  3. GPT-6 Astra has gained the ability to drive a car(drivingbench.com)
    103comments
  4. Stripe's Knowledge AI Platform(stripe.dev)
    54comments
  5. Radicle: Disclosure of Vulnerability in the Network Protocol(radicle.dev)
    8comments
  6. Jev Can't Be Calibrated(alexmolas.com)
    37comments
  7. Strands Harness(strandsagents.com)
    47comments
  8. Jev in 25 Lines of Python(nobodywho.ai)
    146comments
  9. Claude Code reads AGENTS.md only when telemetry is on [fixed](szypowi.cz)
    199comments
  10. GPT-6 Sol and Luna(openai.com)
    804comments
  11. I don't want the details(michaelheap.com)
    116comments
  12. Z80 REPL (2018)(abagames.github.io)
    12comments
  13. Tokens Too Cheap to Meter(jyn.dev)
    85comments
  14. Claude Opus 5.5(anthropic.com)
    1036comments
  15. Web-based IBM 1620 emulator and IPL-V from 1963(github.com/pkimpel)
    3comments
  16. QuestDB (YC S20) Is Hiring a Sales Engineer(questdb.com)
    discuss
  17. OpenAI GPT–6 Astra breaks Enigma message that has resisted solution since 2005(cryptocellar.org)
    419comments
  18. Transit rewards(waymo.com)
    275comments
  19. Seattle City Council votes to ban surveillance pricing in sale of groceries(consumerreports.org)
    10comments
  20. The GitHub wiki is an anti-pattern (2022)(michaelheap.com)
    69comments
  21. What California is learning from solar panels built over irrigation canals(kqed.org)
    590comments
  22. Microsoft killed FoxPro in 2007. Anyway, here's FoxPro revived(foxscript.org)
    233comments
  23. How did AMD Ryzen get 50% faster in two years?(lemire.me)
    176comments
  24. ReBarUEFI: Resizable BAR for almost any UEFI system(github.com/xcuri0)
    65comments
  25. 'We hacked the FBI:' Hackers say they have data on all FBI employees(404media.co)
    540comments
  26. SAML: A fractal of bad design(trailofbits.com)
    160comments
  27. Samsung accidentally freezes its smart fridges with a software update(androidauthority.com)
    194comments
  28. Data-only attacks are easier than you think (2024)(usenix.org)
    32comments
  29. WordPress: Unauthenticated path traversal leading to conditional RCE(github.com/wordpress)
    120comments
  30. Pentagon says overreliance on AI contributed to missile strike on Iran school(bloomberg.com)
    442comments

Gemini 3.8 text-to-speech says hello

53 pointsby 1h agoblog.google
23 comments
34m agoHN ↗

Of possible interest is my Emotive Audiobook Creator, KeenLore. Here's a video of the web app showing how it works:

https://www.youtube.com/watch?v=WAeHgE94rVo

Locally hosted, no cloud, no tokens to pay. Reads a book using a full cast of characters. Quotation attribution detection (for my novel) is at 97.2% accuracy (485/499 quotes identified and assigned correctly). The autofill of character voice descriptions uses the prose to determine how the character sounds.

32m agoHN ↗

<grumble grumble people putting in links they expect you to follow to arbitrary goatse youtube videos for all I know>

The title of the video is 'KeenLore - Emotive Audiobook Creator Demo' and it appears to be a web UI and some local stack that reads text files.

28m agoHN ↗

It's cool technology and I read a lot of audiobooks, even hundreds of hours of TTS. I feel like my brain can fill in the character voices from the text - on the page it's not like they're different fonts.

I understand audiobook narrators often do it, and that's fun. But it's not so critical in my opinion

26m agoHN ↗

Our imaginations and minds continue to rot under the weight of endless and effortless entertainment

33m agoHN ↗

Voice replication: Recreate consistent vocal profiles from just a 30-second audio sample of your voice or a voice you have the rights to use, backed by built-in consent verification, SynthID watermarking, and C2PA credentials to protect both developers and their vocal talent.

I guess voice cloning is widely enough available now from other providers that Google are no longer hesitant to ship it.

31m agoHN ↗

They probably do something similar to GPT-Live where they expect a given voice profile to send them a sample saying 'This is the owner of this voice and I consent for synthetic samples to be made of it'

and/or local voice cloning is good enough as is so Google doesn't grant a uniquely liable ability?

17m agoHN ↗

"Don't be evil... unless other companies are doing it first"

4m agoHN ↗

voice cloning is a tool, it is not necessarily evil, even though the scenarios it can be used for nefarious purposes outnumber the legitimate ones.

30m agoHN ↗

"Super tinny monotone robotic voice" does not sound neither tinny nor monotone. Compared to what TTS from 90s sounded like. Or even how actors impersonated robots in movies. Has the model been eating too much hype DJs?

28m agoHN ↗

I direct my own extended daydream Star Trek fanfic (okay, I'm on season 2 episode 17) and recently I looked to see if I could have each scene file be read aloud a la an audiobook or radio drama.

Getting GPT-Live to have unique enough voices and to be expressive with how I imagine the voices going in my head is hard to direct, there's not enough control there.

So this Gemini 3.8 specific large voice library and ability to tightly control (if you are willing to write a script) is nice to find, and while I'm not sure which of the 5,286 Gemini products this is, nor how to onboard and get started feeding this my own text files, nor what training will happen to my data if I did somehow use it, I love that the state of the industry is such that Google can do this and release it publicly, because that means eventually an equivalent product can come from someone else and be used locally / confidently that the generated audio or inputs won't be retained and misused.

5m agoHN ↗

You might find this interesting, it seems to be exactly what you need to make an audio drama: https://github.com/Finrandojin/alexandria-audiobook

Also the Qwen3-TTS demo is cool, you can describe the voice you want: https://huggingface.co/spaces/Qwen/Qwen3-TTS

I came across both on this subreddit, it's very active: https://www.reddit.com/r/TextToSpeech/

I'm personally using this locally: https://github.com/mateogon/pdf-narrator (it's a Python frontend for Kokoro) on my M1 Macbook Air (from 2020, with 8GB RAM) and it's incredible. I make my own audiobooks now - for free!

My favorite voice is am_michael and here's a sample: https://voicerankings.com/voice/kokoro-82M/male/am_michael/s...

27m agoHN ↗

Great. Now in additional to AI email responses I will get AIs impersonating my contacts on the phone too. Lovely.

24m agoHN ↗

Sorry if this is in that article, but I am on my phone and can't see it. How much would this cost to batch generate an audiobook? Right now I just listen to things in the 11 labs app which is free, but I would rather just generate audio files.

21m agoHN ↗

Roughly 5-10$ for 10h, assuming you few-shot it.

Price per hour:

- 3.8 Flash TTS, standard: $0.81

- 3.8 Flash TTS, batch: $0.41

- 3.8 Flash‑Lite TTS, standard: $0.54

- 3.8 Flash‑Lite TTS, batch: $0.27

20m agoHN ↗

Would be great if this would power the Google Books app feature. The voice system there is pretty out of date.

14m agoHN ↗

Yes, I am quite disappointed by seeing all this cool AI stuff and yet the same Play Books. Come on, it is the best place to apply AI, in my opinion.

13m agoHN ↗

the ratio of new voice models i see on hackernews to the number actually deployed in any product i use is approximately infinity.

18m agoHN ↗

I’ve never found a TYS that does convincing British accents.

They all sound like Americans putting in their best fake British accent.

8m agoHN ↗

Great price at least until December 31 too.

4m agoHN ↗

Seems like voice actors are safe for now. This is technologically incredible, but the results are really not very good, and usually not particularly close to the prompt. In basically all of these examples some core part of the prompt is completely ignored.