Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. GPT-6 Sol and Luna(openai.com)
    455comments
  2. Claude Opus 5.5(anthropic.com)
    685comments
  3. Microsoft killed FoxPro in 2007. Anyway, here's FoxPro revived(foxscript.org)
    8comments
  4. 'We hacked the FBI:' Hackers say they have data on all FBI employees(404media.co)
    104comments
  5. OpenAI GPT–6 Astra breaks Enigma message that has resisted solution since 2005(cryptocellar.org)
    344comments
  6. SAML: A Fractal of Bad Design(trailofbits.com)
    39comments
  7. Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)(artificialanalysis.ai)
    51comments
  8. WordPress: Unauthenticated path traversal leading to conditional RCE(github.com/wordpress)
    64comments
  9. What California is learning from solar panels built over irrigation canals(kqed.org)
    31comments
  10. Native apps written in TypeScript and CSS(github.com/geastack)
    9comments
  11. The UV index is not the warm sensation of sunlight on bare skin(asciitweezers.com)
    discuss
  12. OpenAI is well positioned to fast-follow Jev(arcturus-labs.com)
    173comments
  13. Unreal Agent(unreallabs.ai)
    48comments
  14. Markdown in /src(htmx.org)
    18comments
  15. An update on how we confirm your age group on Discord(discord.com)
    12comments
  16. Show HN: JevBench, a reproducible benchmark for typed decision models(benchmarkheaven.com)
    1comments
  17. MUNI Heritage Weekend in San Francisco(lawrence.lu)
    32comments
  18. Did OpenAI solve the wrong Navier-Stokes problem?(scientificamerican.com)
    28comments
  19. Rabbit Hole: Minimum L-seams(fractalkitty.com)
    5comments
  20. How did AMD Ryzen get 50% faster in two years?(lemire.me)
    29comments
  21. The Trouble with 'Ntile()'(djnavarro.net)
    discuss
  22. Show HN: Training a model to identify AI web content from structure alone(arxiv.org)
    7comments
  23. A Faster Shortest Path Algorithm(vals.ai)
    5comments
  24. 16-bit Intel 8088 chip (c. 1985)(allpoetry.com)
    13comments
  25. The JavaScript Midlife Crisis(maroun-baydoun.com)
    5comments
  26. George Lucas Returns to Earth, Bearing Gifts(commonedge.org)
    24comments
  27. Pentagon says overreliance on AI contributed to missile strike on Iran school(bloomberg.com)
    126comments
  28. There's a high chance of devices being sold with GrapheneOS preinstalled in 2027(grapheneos.social)
    93comments
  29. Apple has added persistent 'ads' to iOS, and it's driving users crazy(techradar.com)
    398comments
  30. People hooked on vapes try a new way to quit: cigarettes(bloomberg.com)
    49comments

Navigating the Challenges and Opportunities of Synthetic Voices

75 pointsby 2y agoopenai.com
15 comments
2y agoHN ↗

We recognize that generating speech that resembles people's voices has serious risks, which are especially top of mind in an election year. We are engaging with U.S. and international partners from across government, media, entertainment, education, civil society and beyond to ensure we are incorporating their feedback as we build.

People had been theorizing the slowdown in releases was because of this, I didn't expect they would come out and admit it.

I've found the quality of Elevenlabs' voice generation is at least twice as good as this, which is surprising since OpenAI has had more time and money to work on this. There seemed to be an issue with inconsistent accents in the multilingual samples.

2y agoHN ↗

The quality of these samples would be more than sufficient to fool anyone I know.

2y agoHN ↗

I thought everyone (11Labs, open source, OpenAI even) already has human-level TTS models? Is there still an open challenge somewhere (e.g., is there a use case where a better model would make any difference?).

2y agoHN ↗

I haven’t seen any TTS take 15 second audio samples and capture the speech so well.

It’s getting clearer — in the future, if you didn’t see it happen in real life then it didn’t happen.

2y agoHN ↗

Even that...Everybody better get their own Certificate Authority and a family password.

2y agoHN ↗

Idk training a model like this is kinda trivial for most shops. Afaik there are open source ones that do the same thing. We trained one with similar audio quality to this by accident while working on music generation.

2y agoHN ↗

It’s getting clearer — in the future, if you didn’t see it happen in real life then it didn’t happen.

I'm convinced that all these AI technologies have legitimate use cases. But it's far too easy to use them to generate endless streams of noise, spam, and mistrust and drown out anything valuable.

2y agoHN ↗

Expressivity and naturalness are still not good enough. It's getting there, but today you'd never have a voice conversation with an AI and come away thinking "boy, what an engaging, charismatic personality that AI had!" This is partly due to latency, partly due to imperfect speech-to-text which is too slow and doesn't understand conversational dynamics or nonverbal signals, but yes partly due to text-to-speech.

When it comes to naturalness for dialog, OpenAI actually has the best voices I've heard—significantly better than Azure or 11 Labs which are second and third best. But at least the voices they expose through ChatGPT are fairly muted and inexpressive.

2y agoHN ↗

Voice Engine preserves the native accent of the original speaker

I sort of understand this as a goal, but the American-accented German is weird. Like they nail some difficult "r" sounds but botch the easy ones. An odd hybrid between a fluent speaker and a total beginner.

2y agoHN ↗

The Japanese has the same problem. It’s sounds like someone who is fluent, but maintained a beginner level accent, with a slight tinge of robot.

2y agoHN ↗

Companies doing speech synthesis where one can clone a voice using a small recording should add digital watermarks into audio itself, akin to steganography. Companies that allow speech communication can check for these marks to make sure the voice was not generated. Many problems, including quality drop resistance, but this can be done and be a valid defence a year or maybe a few years. Not sure if carriers or phone manufacturers can check for these marks in default phone calls.

2y agoHN ↗

I think most of them do… but they are too easy to remove.

It isn’t just quality drop resistance, it is also filtering and post processing resistance. It is kinda like antivirus signatures: those who want to abuse it just filter and tweak it till it no longer triggers the watermark detection.

2y agoHN ↗

Yes! Your business accepts speech input. So you pay me - your average shady security vendor - a good wad of cash to use my universal watermark detector. It's guaranteed to detect AI stegotech, or your money back!

We achieve a whopping 30 coverage over known voice clone services. It's better than nothing!

2y agoHN ↗

Specifically, we encourage steps like: * Phasing out voice based authentication as a security measure for accessing bank accounts and other sensitive information

I've been trying to tell people this for years. My bank still tries to get me to sign up to voice id whenever I call in. Voice is a terrible method for authentication.