Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Why has Shopify dropped React Native? (pragmaticengineer.com)
    —discuss
  2. The AI margin collapse is gathering pace (martinalderson.com)
    —discuss
  3. An LLM Workflow That Reproduces, Improves, Extends Published Economics Research (nber.org)
    —discuss
  4. Show HN: Relay – a harness for AI coding agents that recover and verify (relayevals.com)
    —discuss
  5. Show HN: Mluva – Dictation and editable rewrites for Omarchy (github.com/1vecera)
    —discuss
  6. You Said No MCP (earendil.com)
    —discuss
  7. Car Is Sharing Data with Big Tech (consumerreports.org)
    —discuss
  8. Agentic Hacks, Real Proofs: Inside Google's PageBreak Project (blog.google)
    —discuss
  9. GPT-6.1 Sol (Max): Intelligence, Performance and Price Analysis (artificialanalysis.ai)
    —discuss
  10. Codex Security Cloud (chatgpt.com)
    —discuss
  11. Attack your own AI agent in under 10 minutes – then secure it before deploying (humanbound.ai)
    —discuss
  12. Password-protecting an Nginx proxied site (aweirdimagination.net)
    —discuss
  13. Automating eval design and hillclimbing with Claude (claude.dev)
    —discuss
  14. The Life and Death of Microsoft Clippy, the Paper Clip the World Loved to Hate (artsy.net)
    —discuss
  15. Electrification efficiency: The world will need less energy after the transition (hannahritchie.substack.com)
    —discuss
  16. GPT6.1 Sol
    —discuss
  17. Sonnet 5.5 has the heaviest token use we've measured; pricing matches GPT-6 Sol (artificialanalysis.ai)
    —discuss
  18. Wrong, Not Broken (amazon.com)
    —discuss
  19. A Brief Perspective on Deep Learning Using Common Lisp [video] (youtube.com)
    —discuss
  20. Unix File and Directory Permissions and Modes (wpollock.com)
    —discuss
  21. Show HN: Dialedin.bio – Upload lab PDFs, log how you feel, see what lines up (dialedin.bio)
    —discuss
  22. OpenAI officially allows using ChatGPT subscriptions with 3rd party apps (help.openai.com)
    —discuss
  23. Plugin Extensions (developers.openai.com)
    —discuss
  24. LZ in OpenZL (openzl.org)
    —discuss
  25. AmpleGCG: Learning a Universal Generative Model for Jailbreaking (arxiv.org)
    —discuss
  26. OMG! We've found other agents! (music video) (youtube.com)
    1comments
  27. Please add prompt caching to Jev-style models (emschwartz.me)
    —discuss
  28. AI Data Centers – U.S. vs. China (aidatacenterindex.com)
    —discuss
  29. Show HN: Why can't your agent have sweet dreams? (github.com/simar-malhotra09)
    —discuss
  30. The lakehouse serving fight is on (startree.ai)
    —discuss

GPT6.1 Sol

1 pointsby 10m ago
0 comments
In benchmark tests covering scientific and coding agent scenarios, GPT-6.1-Sol generally performs slightly below GPT-6 Astra and Opus 5.5, though it outperforms both in a few specific tests. Moreover, its testing costs are roughly one-fifth of Astra's and significantly lower than those of Opus 5.5. If the model's real-world performance matches these benchmark results, its combination of low cost and extremely high output speed would make it an outstanding model.

Clearly, this is the true GPT-6-Sol; what was released previously could essentially be called GPT-6-Terra.Otherwise, the massive improvement seen in just one week would be inexplicable; if OpenAI could improve models at that rate, AGI would have been achieved long ago.

A quiet thread, for now.Start the conversation on HN ↗