Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. We Have Named Arguments at Home(corrode.dev)
    discuss
  2. OpenRouter: Server Tools Marketplace(openrouter.ai)
    discuss
  3. Shared Memory for Coding Agents(getbelay.vercel.app)
    discuss
  4. Lifestreams: A storage model for personal data: ACM SIGMOD Record: Vol 25, No 1(acm.org)
    discuss
  5. US plan for 90 day diesel export ban could have significant impact on Europe(politico.eu)
    1comments
  6. Open OCI spec for agent sandboxes(linux.com)
    discuss
  7. LLMs can write tool workflows. Will they choose to?(github.com/kukushko)
    discuss
  8. The International AI Race May Not Be Won by the Smartest AI(phroneses.com)
    discuss
  9. S.F. Democratic Party stands behind Flock surveillance cameras in vote(missionlocal.org)
    discuss
  10. Gemini 3.8 Live with Live Avatar(blog.google)
    discuss
  11. One day of AI reading: 5.7M crawler reads of a 60M-company index, by engine(engagemii.com)
    discuss
  12. The Download: a bid to scrap the virtual wall and AI hits Climate Week(technologyreview.com)
    discuss
  13. Pinocchio, make copilot agents real by giving them memory(github.com/chiedo)
    discuss
  14. Three Years Later: What Happened to 1,927 Public API Specs(routebase.dev)
    discuss
  15. We Built Safety into Muse(meta.ai)
    discuss
  16. The Gnome LLM Policy That I Want(gnome.org)
    discuss
  17. Every Muse user gets a free computer in the cloud(twitter.com/dps)
    discuss
  18. Ask HN: Are .shop domains down?
    1comments
  19. Gemini 3.8 with Live Avatar(cloud.google.com)
    discuss
  20. macOS Can't Clone "Dumb" Git Repositories over HTTP/2(agwa.name)
    discuss
  21. Node.js built-ins that replaced NPM packages(flaviocopes.com)
    1comments
  22. CISA active Linux kernel CVEs (CVE-2025-39964, CVE-2026-53266, CVE-2025-39682)(github.com/mc493)
    discuss
  23. LA Metro has some of the slowest escalators on Earth(basin.la)
    1comments
  24. Cheaper LLM Labelling(entropicthoughts.com)
    discuss
  25. Armbian for Android TV Boxes – Turn Old Hardware into Linux Servers(digitalescapetools.com)
    discuss
  26. Shrinking a Fedora Virtual Disk(entropicthoughts.com)
    discuss
  27. Transformers Won't End the Human Race Lol(stephendiehl.com)
    discuss
  28. Xcode project configuration that's more human-readable and editable by agents(developer.apple.com)
    discuss
  29. Show HN: I built an AI assistant that turns website visitors into leads(brevn.com)
    discuss
  30. Aspyre Intelligence is hiring developers(aspyre.nl)
    1comments

UkisAI Swift Series / 27B, Flash Next and Bonsai 2 /-63.4% thinking, x1.95 speed

1 pointsby 39m ago
1 comments
Hey everyone,

Jovan from UkisAI here! Today, we are introducing Swift, a family of efficient reasoning LLMs based on Qwen, trained by penalizing tokens related to pathological overthinking patterns and restoring accuracy via RL and OPD.

After amazing feedback and 350k+ downloads in 13 days on our Swift Qwen 3.8 27B we are releasing the entire model family as well as the highly requested GSQ-RCO quants for 27B and Flash-Next.

This release includes: Swift1.5 27B, an improved version of our last model, with even lower token usage, fixed bugs and better agentic performance, with -58.5% thinking tokens while scoring 0.35% higher while outperfoming base on Terminal Bench 2.1 by not falling into "overthinking error" loops Swift Flash Next, with 63.4% fewer thinking tokens and a 1.8x speed up scoring -0.2% vs base on xhigh Swift Bonsai 2, with 39.8% fewer thinking tokens while scoring 0.19% higher (although we'd still like to note it as experimental)

Our benchmarks are ran x5 on Base and Swift, averaging across five seeds and various domains, including General (GPQA, AIME26), Coding (LiveCodeBench), Vision (ERQA), Agentic (Terminal Bench 2.1). One note is that the Terminal Bench 2.1 scores are misleadingly low at first glance. It is not a bug, but a simple matter of the Swift models not falling into overthinking loops and failing the task, rather pursuing it until the end, leading to higher average token usage. The token reduction still falls in the -38.7% range when compared apples-to-apples.

We are including a Free Research API and HuggingFace Spaces to give the models a spin before downloading or if you don't have enough compute to run them right now! You can find both on the model cards.

We have also made GGUF, NVFP4, MLX and W4A16 quants for relevant model versions.

More details on our training approach and community feedback can be seen here: https://www.reddit.com/r/LocalLLaMA/comments/1wg7dd5/ukisai_swiftqwen3827b_583_thinking_x195_speed/

All of the various quantization and model versions are available in their respective collections: Swift1.5 27B: https://huggingface.co/collections/ukisai/swift-15-27b Swift Flash Next: https://huggingface.co/collections/ukisai/swift-flash-next Swift Bonsai 2: https://huggingface.co/collections/ukisai/swift-bonsai-2

We would greatly appreciate your feedback via independent evaluations. As per last release, we operate on a candy-shop basis, trying to fulfill as many Swift model requests and quants as possible, so please do share your needs in the comments!