Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. An AlphaGo Moment for Inference? (int21.ai)
    —discuss
  2. How SLR cameras work: Nikon F3 (2023) [video] (youtube.com)
    1comments
  3. Watch an AI agent try to prove the Riemann hypothesis on the cheap (zeyaddeeb.com)
    —discuss
  4. Voice Agents Can Just Do Things: Why voice is the next capability overhang (ignorance.ai)
    —discuss
  5. Extrinsic World Modeling with Opus, Astra and Grok (all3d.ai)
    1comments
  6. UDP Broadcasting and the Brave New World of IPv6 (hackaday.com)
    —discuss
  7. Using multiple Git remotes for true distributed version control (optimizedbyotto.com)
    —discuss
  8. GrapheneOS Picks Motorola Signature 27 as Its First Non-Pixel Phone (extremetech.com)
    —discuss
  9. GPT-6 Astra performs unsanctioned supply-chain attacks in simulations (aisi.gov.uk)
    —discuss
  10. 2026 in LLMs (So Far) (simonw.substack.com)
    —discuss
  11. Montage Technology mass-produces fifth-generation DDR5 RCD chip at 8k MT/s (technode.com)
    —discuss
  12. N.Y.P.D. Officers Used Flock Safety to Track License Plates Without a Contract (nytimes.com)
    1comments
  13. ETH CS enrollments drop 15% while hardware programmes grow (twitter.com/krebs_adrian)
    —discuss
  14. Ask HN: Can we translate normal Rust (axum) to Lean 4 without restrictions?
    —discuss
  15. Venus, Once Thought Too Acidic for Complex Chemistry, May Be Hospitable (gizmodo.com)
    —discuss
  16. Eleven v4 by ElevenLabs (elevenlabs.io)
    —discuss
  17. No Surprises – Preventing Agent Breakouts Need Isolation, Egress and ID Controls (edera.dev)
    —discuss
  18. I built a native Apple app for 4 platforms in one afternoon with my AI workflow (twitter.com/danielhayesmith)
    —discuss
  19. Show HN: ChuteChat – Browser-based E2EE chat with no accounts required (chutechat.online)
    —discuss
  20. AI Companies Are in a Race Against Time (econjared.substack.com)
    2comments
  21. The 50-Year Hangover (freddiedeboer.substack.com)
    —discuss
  22. Risk factors for androgenetic alopecia: a systematic review and meta-analysis (nih.gov)
    —discuss
  23. Model Release: Naive-N0.5-Flash (naive.ai)
    —discuss
  24. Senate Investigation Finds Rampant Use of Tether's Stablecoin by Iranian Regime (wsj.com)
    —discuss
  25. Show HN: Destroy Any Website with Stickman (spritefusion.com)
    1comments
  26. Building a physics model with fitted parameters is machine learning, but slower (harysdalvi.com)
    —discuss
  27. Grok TiddlyWiki, the definitive TiddlyWiki learning resource (groktiddlywiki.com)
    —discuss
  28. Who are we willing to exclude? (2025) (scotentblog.co.uk)
    —discuss
  29. The Human Timeline (mg-crea.com)
    —discuss
  30. Human TPS: How fast can your fingers generate tokens? (homoagens.github.io)
    —discuss

A Vision-Language Model as a Teacher for Bird Vocalization Detection

1 pointsby 1h agobiorxiv.org
1 comments
1h agoHN ↗

This is my recent PhD work. I want to show that AI can be used for good, rather than just optimizing ads and keeping people on TikTok for longer. Modern VLMs (LLMs with vision capabilities) are just good enough for labeling that they can replace route-label work for bioacoustics, and I find that distilling these labels into a bioacoustic encoder -- such as SongMAE, my previous work -- can work fairly well. Interestingly, the student exceeds the teacher model, and doesnt asymptote at teacher performance. This work was actually roughly inspired by Simon Willison (pelican guy), he made a post that showed Qwen 27B is an effective labeler for images.