Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. GPT-6 Sol and Luna(openai.com)
    455comments
  2. Claude Opus 5.5(anthropic.com)
    685comments
  3. Microsoft killed FoxPro in 2007. Anyway, here's FoxPro revived(foxscript.org)
    8comments
  4. 'We hacked the FBI:' Hackers say they have data on all FBI employees(404media.co)
    104comments
  5. OpenAI GPT–6 Astra breaks Enigma message that has resisted solution since 2005(cryptocellar.org)
    344comments
  6. SAML: A Fractal of Bad Design(trailofbits.com)
    39comments
  7. Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)(artificialanalysis.ai)
    51comments
  8. WordPress: Unauthenticated path traversal leading to conditional RCE(github.com/wordpress)
    64comments
  9. What California is learning from solar panels built over irrigation canals(kqed.org)
    31comments
  10. Native apps written in TypeScript and CSS(github.com/geastack)
    9comments
  11. The UV index is not the warm sensation of sunlight on bare skin(asciitweezers.com)
    discuss
  12. OpenAI is well positioned to fast-follow Jev(arcturus-labs.com)
    173comments
  13. Unreal Agent(unreallabs.ai)
    48comments
  14. Markdown in /src(htmx.org)
    18comments
  15. An update on how we confirm your age group on Discord(discord.com)
    12comments
  16. Show HN: JevBench, a reproducible benchmark for typed decision models(benchmarkheaven.com)
    1comments
  17. MUNI Heritage Weekend in San Francisco(lawrence.lu)
    32comments
  18. Did OpenAI solve the wrong Navier-Stokes problem?(scientificamerican.com)
    28comments
  19. Rabbit Hole: Minimum L-seams(fractalkitty.com)
    5comments
  20. How did AMD Ryzen get 50% faster in two years?(lemire.me)
    29comments
  21. The Trouble with 'Ntile()'(djnavarro.net)
    discuss
  22. Show HN: Training a model to identify AI web content from structure alone(arxiv.org)
    7comments
  23. A Faster Shortest Path Algorithm(vals.ai)
    5comments
  24. 16-bit Intel 8088 chip (c. 1985)(allpoetry.com)
    13comments
  25. The JavaScript Midlife Crisis(maroun-baydoun.com)
    5comments
  26. George Lucas Returns to Earth, Bearing Gifts(commonedge.org)
    24comments
  27. Pentagon says overreliance on AI contributed to missile strike on Iran school(bloomberg.com)
    126comments
  28. There's a high chance of devices being sold with GrapheneOS preinstalled in 2027(grapheneos.social)
    93comments
  29. Apple has added persistent 'ads' to iOS, and it's driving users crazy(techradar.com)
    398comments
  30. People hooked on vapes try a new way to quit: cigarettes(bloomberg.com)
    49comments

VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild

55 pointsby 2y agojasonppy.github.io
5 comments
2y agoHN ↗

Model weights were just released and people on Reddit have successfully run locally. The results are impressive!

You should let your family members know that the scammers are going to start sounding like people they know.

2y agoHN ↗

That is a truly terrifying reality. Scammers are already having a field day doing cold calls on senior people. Just imagine how confusing it will be when they think it's a relative

2y agoHN ↗

It's very impressive, but the samples sound slightly garbled throughout, as if some heavy noise suppression was applied to the whole clip. I wonder if that's pre-processing for the samples on this site, or if that's a side effect of this model.

2y agoHN ↗

It sounds to me like the edits are accomplished by taking the original text and feeding back through the model, so that the edited tracks contain zero original audio. Of course when your text precisely matches the training material it comes out pretty well, but not perfectly.

In late March 2024, this results in some noticeable distortion. Don't ask me whether you'll be able to notice by, say, September 2024 though.

It probably comes out less jarring than trying to insert the edits overall, but you still really want a clean voice sample going in. For instance, the Pulp Fiction voice had a background music track, and it does interesting things to the resulting voice once it has to start synthesizing novel text. (Second sample in the first table, Irving Ramses.) I wish they'd tacked on a bit more text to that one, I can just barely hear it's doing "weird things" but there's not a lot of time to analyze it before it's over.

2y agoHN ↗

Sorry for the derail, but I really hate that published papers with audio/video clips often don't work at all on iPadOS or iOS. I do most of my leisure reading on my iPad Pro, purposefully separating that leisure time from when I'm on my MacBook, which I try to use exclusively for working.