Best stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Show HN: An e-ink frame that hears birds and draws them as 1800s illustrations(github.com/arnegiacomo ↗)
    217comments
  2. Introducing System One Models and Jev(typesafe.ai ↗)
    439comments
  3. I can't stop thinking about Papua New Guinea(notnottalmud.substack.com ↗)
    455comments
  4. 25 years of mass surveillance is enough(schneier.com ↗)
    328comments
  5. Steam Frame starts at $1059(steampowered.com ↗)
    631comments
  6. iOS 27, iPadOS 27, and macOS 27(apple.com ↗)
    843comments
  7. A single firm is behind OpenAI, Anthropic, and Meta hacking scandals(effort.news ↗)
    218comments
  8. Dario, Please(rdi.sh ↗)
    309comments
  9. An update on Wayback Machine access(blog.archive.org ↗)
    333comments
  10. EU chief opens door for Canada to become 'associate member'(bbc.com ↗)
    702comments
  11. Suspected sabotage causes major Netherlands rail disruption(bbc.com ↗)
    453comments
  12. Pion, an agent designed to run any company autonomously(andonlabs.com ↗)
    601comments
  13. US confirms for first time it has deployed space weapons(bbc.com ↗)
    341comments
  14. Gemini 3.8 Live and 3.8 Live Extended Thinking(blog.google ↗)
    302comments
  15. Let's make quality the norm again(forbrukerradet.no ↗)
    443comments
  16. Building a Linux GPU Driver for the M4 Mac Mini in One Month(codyho.dev ↗)
    228comments
  17. Linux from Scratch(linuxfromscratch.org ↗)
    110comments
  18. Apple Reference Image: A New Approach for Verified Photography(security.apple.com ↗)
    244comments
  19. Show HN: Capsule – Single-file web apps that save their data into SQLite(withcapsule.app ↗)
    150comments
  20. Most people prefer traditional architecture(worksinprogress.news ↗)
    333comments
  21. Distributed Systems Classics (2017)(nvartolomei.com ↗)
    74comments
  22. Java 27(openjdk.org ↗)
    370comments
  23. Why I'm still bearish on LLMs after Navier-Stokes(dank.systems ↗)
    399comments
  24. Israeli Minister Threatens Filmmakers' Citizenship over Gaza Documentary(reutersconnect.com ↗)
    79comments
  25. We got admin access to Baseten's production GitHub(strix.ai ↗)
    173comments
  26. America's Driver's License Breach Is a National Security Disaster(lawfaremedia.org ↗)
    191comments
  27. Charts built for Chat(dbtcharts.com ↗)
    89comments
  28. Ubuntu 26.10 completes transition to Rust-based coreutils(omgubuntu.co.uk ↗)
    312comments
  29. A beginning for mathematics(daniellitt.com ↗)
    147comments
  30. Alternatives to MinIO for single-node local S3(rmoff.net ↗)
    102comments

Gemini 3.8 Live and 3.8 Live Extended Thinking

452 pointsby 20h agoblog.google
300 comments
19h agoHN ↗

Not a great impression to have your demo video demonstrate how one of your 'most advanced' AI models loses to the most common check-mate pattern in all of chess.

19h agoHN ↗

Seems more than good enough for a live model though! I can imagine this demo being extended to be a lot nicer to play with. You can just feed the model engine analysis and it can make as high of quality moves as needed. No longer any correlation between the model's understanding of the position and the moves that would be made but I think that's still a really nice improvement when thinking about this as adding live voice interaction to existing chess vs computer functionality rather than adding chess to possible interactions with the latest live voice model.

18h agoHN ↗

agree that it’s a weird choice, but more because i don’t need my chat model to play chess at all when chess engines exist.

17h agoHN ↗

For an LLM, just being able to play an entire game of chess without illegal moves and without inventing pieces that aren't on the board is an achievement. Even more so for a live model. Then again, who knows how much the harness helped here

14h agoHN ↗

I am an adult, and i lose to elemental school kids in chess.

I am not so advanced enough as a human being.

19h agoHN ↗

Gemini is underrated in that it produces the only prose that is somewhat bearable to read.

19h agoHN ↗

For heavyweight work I have been using Astra, but for rabbit holes and brain storming Gemini is far more enjoyable to interact with.

I'm worried in their push to catch up on the SOTA front, it's going to lose that natural sounding touch it currently has.

19h agoHN ↗

Good news then, I don't think they're in a hurry to catch up to SOTA.

18h agoHN ↗

Mostly because it answers quickly and is more agreeable (too agreeable at times).

Meanwhile Claude and Astra like to couch all their agreements with caveats and provisos.

18h agoHN ↗

Meanwhile Claude and Astra like to couch all their agreements with caveats and provisos.

Sometimes that's what being smart sounds like.

18h agoHN ↗

And sometimes that’s what trying to sound smart sounds like.

17h agoHN ↗

This is commonly why, on Reddit in particular, you can get eaten alive.

Someone confident but incorrect, can often sound more convincing than someone with actual expertise. The expert must add caveats/hedge, because those are the facts on the ground, whereas the person reciting google can be entirely confident.

Of course the people judging aren't experts, so they side with confidence and simplicity. Heck, just writing shorter replies on Reddit is rewarded. Nobody reads the articles, let alone a paragraph-long reply.

That all being said though, there are limits. Sometimes LLMs on high-thinking go off on full tangents based on little, and don't have the self-awareness to bring it back.

17h agoHN ↗

In my experience actual experts don’t hedge because they have a perspective. They might say “I think”, but avoid weaseling.

15h agoHN ↗

My experience has been very much the opposite of yours.

To an expert communicating with a layperson is a form of compression. You must turn some very complex idea into one that you suppose the other person can grasp given their limited frame of reference. It's always lossy, and you have to guess how much you can remove without sounding patronizing or being inaccurate. It's tough, and the more you know the tougher it gets.

Ever done that "explain what happens when I visit Google in my web browser" interview question?

A sales guy will answer in a sentence. An engineer might be able to talk about it for several days and still not be sure they didn't miss anything important. That much knowledge can actually be detrimental to communication.

16h agoHN ↗

I consider, on the contrary, caveating and hedging annoying 'typical redditor'/internet behaviors: they care more about being "technically correct" than conveying the message. On the internet, if you make even the tiniest mistake or simplification, someone will criticize you, so you're trained to always hedge. In normal discussions with friends you can just make general statements and people get what you mean.

14h agoHN ↗

You aren't arguing a "contrary" to what I actually wrote above. You're arguing with a strawman Redditor, and a point adjacent to what I was posting about - experts responding to topics within their actual realm of expertise.

So I won't be addressing this, for those reasons and others.

18h agoHN ↗

Yea. I have asked it to verify my ideas with experiments sometimes. And it cheats and warps the results so that the results are reached

18h agoHN ↗

“caveats and provisos” makes me think of Robin Williams’ genie imitating William F. Buckley Jr.

16h agoHN ↗

"couch all their agreements with caveats and provisos."

When you're a ChatGPT Projects or Claude Projects user, those caveats and provisos are your worst enemy because they'll change caveats into hard rules (either for the session or committed to memories) and you end up in absolute hell having to make it investigate to figure out why it can no longer produce anything but read-only pre-check code that never actually does anything but keeps performing stupid safety checks.

16h agoHN ↗

Yep the only way out is hooks to forbid what can be detected by ast and second model to prune comments, flatten pyramids of fallback, and squash the test suite removing quirks maintaining wanted behaviors.

18h agoHN ↗

Agreed. My impression is that the more verbose output of sol, astra etc is that it helps it steer itself on long running tasks (but is worse for the human user to read)

18h agoHN ↗

Yes, when post training models for long tasks this happens gradually. It is not easy to prevent it as such.

17h agoHN ↗

Yes I've noticed there's also this drive to implement and start talking about how it would write specific portions of code in response to design/trade off questions. I have to prompt Sol/Astra almost every time with a note that I am not looking for implementation advice since I mostly use them as a rubber duck in the design phase

18h agoHN ↗

For rabbit holes, how do you get Gemini to do any research before answering? I've very recently had it hallucinate on me like it's 2023, and that was on Pro/Thinking, as far as I remember.

17h agoHN ↗

Is Google still chasing frontier? Seems like they haven't had a "Pro" model in forever. I think a good niche for them would be right where they are now.

17h agoHN ↗

These are native speech-to-speech models, so I think they've decided not to do that anymore.

16h agoHN ↗

They are. They were supposed to release 3.5 Pro over the summer but haven't because of persistent architectural/technical issues allegedly. Which is better than releasing it in that state imo.

Their AI leadership team has taken some hits recently too, in the form of departures. I believe when they get their bearings they will be competitive again. 3.8 Flash has been a great model for me.

16h agoHN ↗

How do you know you're using Astra?

My ChatGPT env only says "low", "medium", "high".

Is this a "pro" thing? I have totally no idea what I'm talking to, so actually I'm thinking of stopping my plan. Gemini and Claude are much more clear about it.

Anyway, I like the speed at which Gemini responds so indeed for simple things it is preferable.

15h agoHN ↗

For me (Plus plan, iOS app), Astra only shows up under the Work tab.

I’ve been using Work for all my queries, since it seems to just be the same interface as Chat but with more features. I don’t understand why they’re two separate things.

15h agoHN ↗

For the old school (I hate that that’s arguably applicable) AI dating types, and so on. The people using it not for productivity.

15h agoHN ↗

It's entirely possible that in their testing of newer models, the whole problem is that even if it's doing better in benchmarks, maybe it's insufferable to work with, thus they're not releasing it.

18h agoHN ↗

Wondering if people have managed to have Gemini in-front of other models like claude/codex models and only interact with that. Having Gemini act as a pure human/llm translator.

18h agoHN ↗

Not in the principled sense you mean but I have in fact recently started having Gemini explain to me what Claude is talking to me about, lol.

16h agoHN ↗

BTW, as kind of a follow-up to this, I think the most important finding to report--for those who only use one model, as many people seem to--is just how much more often Claude seems to be extremely confidently (and even insufferably) wrong than any of the other three big models (all of which I use quite often... yet I only pull out Claude when I've given up hope in a problem and are looking for out-of-the-box brainstorming).

And like, it does this despite it speaking in extremely dense math, which both makes it sound correct and requires a lot more effort to prove when it is wrong... yet, it isn't actually correct more often, and so that time sink just isn't worth the benefit. I then think many people--including people who can speak math (as can I)--just stop bothering to check everything, as if you come across a human who speaks like this it probably does correlate with slow and careful thought that helps prevent errors.

Instead, Claude has the mistake rate of a somewhat accelerated beginner impossibly combined with the language of an expert professor; and we as humans just aren't good at that combination: it becomes very dangerous and makes it take longer to spot its egregious mistakes and trained-in biases. If you have to use Claude, I thereby claim you really need to have a team of not-Claudes to help insulate you from this, and Gemini (while being a bit senile) is a lot more collaborative and approaches problems in ways that makes it harder to get tricked.

(To translate this into more of an engineering analogy: Claude always feels to me like the engineer who put more effort into learning how to program in functional languages than into how to actually develop working code, and then confidently presents you answers in Haskell or Lisp that never quite work. To find their errors is then very costly. In contrast, Gemini feels more like a Java or Go developer who knows they are a cog... that's helpful! <- Which maybe just goes to show that AI has finally turned me into a manager, omg.)

15h agoHN ↗

Yes I have noticed this. I frequently have stronger models review weaker models. It’s very instructive to see what they get wrong.

18h agoHN ↗

Does anyone know if there is a dedicated model which makes Claude output nore human readable and less slop?

Lately it became load-bearingly-reality-difficult to not only read, but to comprehend the Claude output

18h agoHN ↗

It's also the only model that generates accurate translation and localization. No other frontier model comes close. Although Gemini's coding capabilities are subpar, its natural language processing is top-tier.

17h agoHN ↗

I’m curious how you guys keep track of each model’s coding capabilities. The landscape keeps changing. I don’t suppose you benchmark all frontier models every other month, right?

17h agoHN ↗

I use them. Daily. Gemini hasn’t been a contender by comparison for a long time.

17h agoHN ↗

True but it has a niche in SQL reviews for me. Looks like Google has a lot of good sql in their corpus and in their RL digital lobotomy factory.

16h agoHN ↗

I wonder if it's a harness thing or a model thing at this point. I feel all coding models are quite capable for most tasks I want them to do.

Most of the time I don't need what the bench tests and I'm not really giving them completely ambiguous tasks without any refinement.

I only find marginal differences between models at this point and it almost feels like personality quirks in each model than anything.

16h agoHN ↗

When comparing OpenAI and Claude thats pretty much true, but not Gemini... And have you tried Antigravity? Yikes

15h agoHN ↗

The CLI version of agy is great. Have you tried it?

12h agoHN ↗

Compared to gemini-cli that they took out behind the woodshed, I hate it.

10h agoHN ↗

I've used Antigravity as my main coding agent on one of my biggest projects for about a year. It's been great for me. (and I use Claude, Codex, Grok and Muse for all the other projects)

15h agoHN ↗

I did a test involving implementing cobol control flow in Java for a source to source translation project. Gemini was the only model to get the edge cases. Cobol is very peculiar in this regard.

15h agoHN ↗

It's very good at Elixir in my experience too. And it just does what I ask and doesn't wind me up like Opus. I don't think I've had to insult it more than once per day.

12h agoHN ↗

I had a typical $20 Gemini plan that I just downgraded to their $5 plan (to keep access to some of the models). It had been so long since I let Gemini work on (or review) any code / design / html (anything) that I couldn't justify bothering to keep wasting money on it. It fell behind badly over the past year. Astra might as well be an alien super intelligence at code compared to Gemini. I enjoy talking to Gemini, it is very good at conversation, I get solid answers to everyday questions. I intend to keep the $5 plan indefinitely for basic use. I don't expect they'll ever resurface as a competitor in coding with Astra & Fable et al.

16h agoHN ↗

I found that it's shockingly good with R. (the only language I know and can correct for)

I doubt they even intended it to be, but it seems like I kept going from resorting to 3.5-3.8 (over time) to realizing that Claude and GPT, while great at Python, will make rudimentary mistakes with R; even when they compose giant complicated R code.

9h agoHN ↗

I'm guessing this is partly because of Gemini's world knowledge. I tried asking the model multiple internet humor and memes and it answered correctly around 80% of the time

16h agoHN ↗

That might depend on whether you are translating fiction or nonfiction.

Anecdotally I'd rate Gemini behind Claude and OpenAI models at fiction and I can't find any benchmarks showing Gemini is the clear winner at this task.

18h agoHN ↗

We have an agentic system that produces insights for end users, and runs most of its work on DeepSeek v4.1 Flash but as an output stage transforms the resulting text through Gemini 3.8 Flash for readability, and it works.

On my TODO is try and run all of the analysis pipeline in dense "machine speak" to save on tokens and just let Gemini sort it out at the end.

18h agoHN ↗

In my experience, with minimum prompting, deepseek also generates very decent text.

18h agoHN ↗

Not at all in mine. Deepseek has some of the worst prose of the close to frontier models in my opinion.

18h agoHN ↗

I find its style the most sycophantic and annoying personally.

1h agoHN ↗

Never has it been brought so clearly into light how brilliant I am, than in correspondence with Gemini.

18h agoHN ↗

Try Gemini live in a multi lingual environment. It can pick out speakers and live translate to you. Truly underrated for its capabilities.

17h agoHN ↗

I actually did (was going to travel internationally), and it wasn't as useful as you'd think. I would be talking to someone, and in the background someone else would be talking, and it would translate both people.

Only worked in a 1:1 in a quiet place. Still, can't complain for free.

17h agoHN ↗

"Only worked in a 1:1 in a quiet place"

well then its not model problem

16h agoHN ↗

It's not a model problem, but it is a SW problem. It should be able to distinguish nearby people from people farther away and give me options to set a threshold on who to include.

15h agoHN ↗

As sister comments have pointed out, I find that it's very much dependent on the phone you're using. If it can't even show me a waveform then I don't think it's going to be able to tease out any sort of speech

14h agoHN ↗

Sorry I wasn't clear. It's about the hardware microphone. They're usually optimized to filter out far field sounds and only focus on near field sounds. So if you're trying to understand a conversation from even 10 ft away, it can be a problem. Not saying that's the situation that you encountered, but it does make sense that a $50 phone versus a $500 phone might have different capabilities when it comes to the microphone

8h agoHN ↗

I'm talking about Gemini Live in the Gemini App, not Google Translate.

Tell Gemini Live which speaker you want to engage with, and it is relatively intelligent about it.

16h agoHN ↗

Not the same, but related: Gemini is great for querying text in another language, it provides really cogent, useful responses with just enough source language quotes to be able to reference the source text effectively.

18h agoHN ↗

I find Astra's prose very good, too. I have been using it to rewrite all my LLM-generated docs as of lately.

18h agoHN ↗

I was surprised when (finally) trying out Claude how much I preferred Gemini's way of communicating. I wont argue Claude is better at coding, but for knowledge work, I had to dig through Claude output to find what I actually wanted. At times, it even felt borderline incomprehensible.

17h agoHN ↗

Just today I had Sonnet 5 generate this (asking about always-on display in the iPhone e-versions):

This mirrors how Apple has always segmented Pro vs. non-Pro iPhones: base models got LTPS panels while Pro models got LTPO, and only with the mainline iPhone 17/17 Plus did that gap close the standard versions previously lacked the smoother 120Hz ProMotion technology and the always-on display feature, unlike the Pro models — the 17e is the one model line still using the older, cheaper panel.

(emphasis mine)

I mean, I can guess what it is trying to say, but who RL'd this nonsense?

17h agoHN ↗

Fable 5.1 even today inundates prose with "it's not this, it's that" type of garbage. I had to rewrite two paragraphs from a generic class announcement which I was planning to post on the LMS. Not sure where that productivity gain is that everyone is talking about.

16h agoHN ↗

Opus specifically talks as if having a stroke. 4.6 was the last version that was pleasant to work with

15h agoHN ↗

I couldn't even stand 4.6 by the end. 4.5 was very usable, 4.6 started alright and somehow got more annoying. 4.7 made me quit my subscription and ditch their services entirely - it's an honourary member of my quite short "Coworkers I'd like to throttle if I didn't work remotely" list. Infuriating to instruct or communicate with.

17h agoHN ↗

Chatgpt doesn’t seem so bad lately. At least as a Claude refugee.

16h agoHN ↗

Yeah, I set it to “warm: less”, “enthusiastic: less” and “emoji: less” and it was much more bearable than I remembered it being before. Although it does love to “separate” questions when it thinks.

16h agoHN ↗

Terrible for code, amazing for prose

I've set my documentation sub agent to Gemini and my code agent to Luna

15h agoHN ↗

I work for Google and we have the choice between Gemini models and Opus. Opus is slightly better than Gemini flash but I find the style unbearable.

It reminds me of a pedantic grad student.

11h agoHN ↗

Not only that, Opus likes to invent arcane jargons. And the longer the task runs, the harder it is to understand its output. The conversation will be filled with uncommon word choices and awkward sentence structure.

6h agoHN ↗

The 4.5 models were the last Anthropic models that were actually pleasant to read. Especially Sonnet 4.5 was really fun to use for non-coding purposes. I liked having it generate silly short stories while waiting for deployments to finish. Haiku 4.5 still has a small bit of that charm, but 4.6 and later models are all unbearable unless you want a simple 'this is the answer' to a question.

15h agoHN ↗

It also doesn't anthropomorphise itself, like at all.

11h agoHN ↗

I can’t agree more! It feels really smooth to drive, kind of buttery compared to other frontier experiences, at least in antigravity 2.0 or whatever. I’ve been doing some web app coding with it and I’m happy with the results.

19h agoHN ↗

Our company's Google Workspace Business only offers 3.6 flash & thinking in the Gemini App. Has anyone else seen 3.7 or 3.8 roll out?

19h agoHN ↗

I have access to both, benchmarks are actually better on 3.7 for my task, but happy improvement over the others.

18h agoHN ↗

I'm still only seeing 3.6 Flash / 3.6 Thinking in my Google Workspace for Education account, and 3.5 Flash-Lite / 3.6 Thinking in my "Plus" plan Gmail account.

18h agoHN ↗

3.6 Flash and 3.1 Pro are included in the basic Workspace subscription. The Workspace admin has to upgrade your seat for the access to newer models ($17/mo now, $24/mo starting Jan 2027).

18h agoHN ↗

I'm the admin. I see the "AI Expanded Access" addon option in the dashboard, but it says nothing about which models it includes.

4h agoHN ↗

Google takes a long time to roll out models in general. Both my personal and work accounts still have only 3.6 Flash, even though 3.8 was released 2 weeks ago and 3.7 more than a month ago.

19h agoHN ↗

Gemini's Live Mode is already much better than GPT Voice in my personal experience, even though it was much dumber. It really does feel like talking to a real person. ChatGPT keeps humming to whatever I say and has some weird voices.

Excited to try this out! Shame on Google for not releasing Gemini 3.8 for Google AI Plus users yet, though.

18h agoHN ↗

OpenAI just released the new full duplex mode to the API as gpt-live-1 or something like that. Very realistic.

16h agoHN ↗

Agreed. I've been using GPT-Live-1 this week, with Claude as the backend brain. It's amazing, feels like working with Jarvis. It certainly made me feel there's no point in human telephone support now - but I'm sure I'd find edge cases if that really was something I wanted to build out myself.

11h agoHN ↗

Oh... Interesting! How do you have Claude as the backend brain for GPT-Live-1?

I'm not very happy with live conversations with Claude, so this seems like it might be a good option.

18h agoHN ↗

Would you mind expanding on this? I thought Google changed the name to 'Gemini Enterprise Agent Platform', and altered focus to 'agent governance' workflows, but that there were no breaking changes from what was offered with Vertex AI.

19h agoHN ↗

I've been asking the same about the open weight models, we're buying our tokens from others now, though I think those people are renting hardware from Google in the end anyway

10h agoHN ↗

The real question should be when will it be available on Vertex and not be full of bugs?

19h agoHN ↗

I wonder when/if we’ll see Gemini beating Fable and Astra. Last year I would have confidently bet Google will overtake the others just because they have the data, the hardware (TPUs) and a fat advertising money pipe and yet they are still behind. Anyone anonymous at Google want to hint when Gemini 4 will be out?

19h agoHN ↗

If Google didnt have their ad buisness theiy'd be out by now. They're like BlackBerry and Nokia at this point almost.

19h agoHN ↗

Maps, Waymo, TPUs, YouTube, Docs, GMail, Cloud, Android, Chrome, Photos...

19h agoHN ↗

Not sure what your comment mean in the context of parent's comment.

As opposed to what, them not having it and burning money that isn't their instead like openai and anthropic? At least Google is feeding itself instead of having to create a bubble to stay alive

19h agoHN ↗

Google spent most of their cash, they are now taking loans for data centers too. They recorded their first quarter of negative cash flows ever

18h agoHN ↗

Which is still a better position that the others? My point is you can't consider that a bad thing if you think it's ok for their competitors in the field. And if you don't and your judge them equally, then at least Google has its own cash glow and could turn the gas off at any point to go back to printing money while they have not choice.

18h agoHN ↗

Google has in excess of $121 billion of cash (& equivalents), net of total debt.

Big rich companies take on debt for reasons that are sometimes inscrutable from the outside. Recently, they have been borrowing for ~5%, about a half point above what the US government gets for 10-year Treasuries.

Apple has been financing operations with debt for a number of years as part of a complex optimization plan.

No, Google is not broke.

16h agoHN ↗

There are a lot of accounting shenanigans going on in Big Ai, financial reports are misleading, eg Meta "building" a data center through a shell company and then renting it to themselves. It does wonders for the looks of the books

Some expert wall street analysts discussing what they found and how they dissect things, have a healthy skepticism of Big Ai

https://www.youtube.com/watch?v=YrJzjC4kKCY

18h agoHN ↗

In a quarter with record breaking income of over $100 billion, they also had negative $5.9 billion in free cash flow from their "bank account", because the people selling them chips and building them data centers want to be paid, while the investment itself is depreciated over time.

If my quick search is correct, Google is sitting on a quarter trillion dollars in cash and marketable securities. They could keep doing the negative cash flow thing at this scale for another decade.

18h agoHN ↗

They remind me of Kodak inventing the digital camera and sitting on it to preserve their film business.

18h agoHN ↗

And YouTube, and Cloud, and Play Store, and Waymo, not to mention that they could coast on their Anthropic and SpaceX stakes if they didn't have any of the above.

19h agoHN ↗

As an "everyday mans AI" I'd say 3.8 Flash definitely already has. Smart enough for the vast swath of people, and only slightly eeked out by Astra(Max) on vision capabilities, like the kind of "Point your camera at something and ask questions" that non-tech people like to do. It's crazy fast and very compute light, so not getting bogged down constantly.

I can't think of a better general purpose model than 3.8 flash right now. It also writes more naturally than the other big models too.

19h agoHN ↗

Yeah its good. Reasonably priced too (at current prices, if they do raise them in January I would stop recommending it). 3.7/3.8 were good releases.

18h agoHN ↗

To me it's a good replacement for search engines. I ask it things like 'If redshifting destroys energy ala Noether, than how can we say that time is reversible or that entropy will find an equilibrium?' and it will not only explain, but make nice interactive diagram/toys to help. A regular search engine would have taken me hours to find an answer.

However, if I want it to DO something then Gemini is in absolute last place. I don't trust it for anything more than renaming files that I don't care about very much or extracting data (though it's too expensive for data extraction at scale).

16h agoHN ↗

I’ve mapped my iPhone Action Button directly in to Google App AI Mode (which is Gemini 3.8 but faster inference than the Gemini App, due to harness)

It’s for this exact kind of scenario where a random question pops in to my head.

Plus, it’s the most grounded by real live data of all the chatbots.

Silicon Valley people are majorly sleeping on Google Search AI Mode.

18h agoHN ↗

Yeah +1 to this

People use text with LLMs but it's great to have a high fidelity "analyze this image"

17h agoHN ↗

I was thinking about coding specifically. Also, see all these math and physics breakthroughs, it's usually not Gemini but Fable and Astra.

15h agoHN ↗

I think 3.7 Flash is very good at coding. I have access to both that and Opus 5. Opus is only slightly better IMO and it's much more frustrating to read.

12h agoHN ↗

I think it's good at basic coding and is cheaper and faster. Trying it vs Fable 5.1 I don't see it being even close capability-wise. What I mean by basic coding is "write a script to parse this data" and answering some questions based on the code. It does that well and answers fast. However, when it comes to planning and thinking through a more complicated design problem it's just not there yet.

8h agoHN ↗

It's not even close. Some benchmarks show 3.8 Flash as being close to Astra (DeepSWE v1.1), but they are mostly "bench-maxed". Newer benchmarks like Terminal-Bench 4.0 show Astra at 57.9% and 3.8 Flash at 19.1%. For anyone who has tried to code with Gemini models, there is no contest here.

Astra wins handily on every other benchmark, too. Research, science, 3D, visual, Humanity's Last Exam, etc.

16h agoHN ↗

Not only that, the output of tps is high. So not only can it handle most mid-level tasks perfectly fine, but can do it with speed!

15h agoHN ↗

I find it really good and up-to-date, like within an hour of current events or website updates that it can reference. Astra is a step up for complex stuff, but I'm not going to use up all my tokens to ask it the specs of a 2023 MacBook pro then wait 3 minutes only to get a tome with a complete breakdown of the Mac OS platform and including such things as the supply side dynamics driving the ram size choices and pricing at that time.

18h agoHN ↗

For coding models? I don't think Google is motivated to fight in that market. There's no incentive for them.

Ask yourself, how does Google -- a company that famously does everything -- benefit from SWEs outside of Google having access to powerful coding models? They would just be competition.

Famously, Google just eventually discards almost all businesses that don't have the same fire hose of revenue that ads does. Selling coding plans isn't something they are going to want to do.

Google is clearly motivated to make better search and information finding tools and stuff that will ultimately drive users through their existing search/ads/youtube ecosystem. That's really why they're in Android, that's why they do Chrome. Everything else with them is a sideshow.

Google is also full of beancounters obsessed with data centre quota and resourcing. Even massively profitable ads projects have to justify and fight for it. (Source: used to work there).

I can't think of anything less resource & revenue sensible than providing outside parties access to your TPUs for the purpose of letting them write stuff which could just end up competing with you.

Yes, maybe as part of their cloud business, selling token access could be useful money. But I doubt they'd tune it for coding.

17h agoHN ↗

Also: Google has like a 20% stake in Anthropic, and a very fat cloud partnership

17h agoHN ↗

Ask yourself, how does Google -- a company that famously does everything -- benefit from SWEs outside of Google having access to powerful coding models? They would just be competition.

Bragging rights to say they have a SOTA model. I guess that was more like Google of 10 years ago with moonshot projects. Nowadays, yeah, perhaps if it's not helping sell ads, it doesn't make sense.

14h agoHN ↗

Google did it AlphaGo, AlphaFold etc for bragging rights, now let frontier lab do their thing.

It would be utter stupidity if they keep competing with hyper agile frontier labs. They sensibly moving to big infra provider and that's good niche.

.. perhaps if it's not helping sell ads,

Perhaps cloud business is also a thing which is making big revenue

8h agoHN ↗

You say "keep competing"... the hyper agile frontier labs and Google/MSFT are practically one and the same via circular financing arrangements. Competition on the frontier is kinda a farce, what matters is who comes out of the modern-day Darien scheme as a winner.

At worst Google just need to keep the lights on until the frontier labs file for bankruptcy under the weight of that >$1tn debt they've built up paying for FAANG's data centers. Then Google walks away from all of this with the the best remaining models, best R&D lab, and several hundred billion dollars in infrastructure on their balance sheet.

17h agoHN ↗

Ask yourself, how does Google -- a company that famously does everything -- benefit from SWEs outside of Google having access to powerful coding models?

Catch is, even Googlers internally do not have access to top-tier models. (or did not until recently, when apparently Claude was made accessible to the SWEs internally).

16h agoHN ↗

Googlers now have internal access to frontier models.

Funny enough, after trying it, I went back to G3.8

11h agoHN ↗

Yes, agreed, 3.8 with thinking is quite good.

17h agoHN ↗

I’m wondering if they even see a coding agent as a valuable prize. It’s a competitive market in a race to the bottom economically, hard to establish consistent differentiation and virtually zero switching cost for customers.

I think they’ve made a shrewd move in focusing on search integration and everyday users (Gemini app) vs software power users. They have their corner and nobody is really competing with them, plus it feeds directly into their existing revenue stream.

13h agoHN ↗

OpenAI/Anthropic are being valued at 25/50% of Alphabet respectively in secondary markets.

Maybe everybody is wrong about how valuable these companies will be, but atm it looks very dumb to not be competitive in coding.

10h agoHN ↗

Google is not being dumb or making some 4D strategic move by not going after coding capabilities.

Occam's Razor is overwhelmingly that they just don't have the organisational capability to capture this market. If they did then they absolutely would have.

16h agoHN ↗

Anecdata, but I've been using gemini personally instead of what I used regular google searches for the better part. Even in car chat/search and some lighter and not so light research (if it's not heavy on technicals). It's good enough that I use it nonstop like that. I wouldn't trust it for coding at all - switching between fable and now astra.

Google in a sense won (me over) like that. I also expected them to brute force their way into everything and dominate. This is how it played out though. Image generation is great as well, but ChatGPT one is more lenient on copyright and nannying - for example when my kid asks me to "take a photo of him and Sonic". Gemini cops out either because of the kid or Sonic, disappointing us both, but ChatGPT can be.. persuaded.

9h agoHN ↗

I just use Grok for anything "controversial" like that. I actually love that different LLMs have their niches. In a typical week I might use all of Claude, Gemini, and Grok, for different kinds of asks.

16h agoHN ↗

Considering how far backwards Google has gone since the release of the 3.x models, the brain drain they've let happen, the lacklustre software ecosystem and their track record with product management ... I would say they're more likely to drop trying to compete at the frontier and try to focus on something else instead.

15h agoHN ↗

For me 3.8 has been good enough that I don’t think I’ll be extending my Claude subscription.

15h agoHN ↗

I have a shared understanding for what you're asking, but semantic versioning doesn't even make sense for AI models

these are essentially monthly releases, the dated releases that Deepseek and Qwen do make more sense

8h agoHN ↗

I wonder when/if we’ll see Gemini beating Fable and Astra

Google has essentially give up the race for the frontier. The gap is widening very fast. Their strategy appears to be to focus on niche areas like voice, small models for on-device (see their deal with Apple), and specialised models for tasks like search, which are very inference efficient.

19h agoHN ↗

Ensure transparency with SynthID watermarking

All audio generated by our AI products is watermarked with SynthID. This imperceptible watermark is woven directly into the audio output, ensuring AI-generated content remains detectable to help prevent misinformation. For details on our approach to safety and responsibility, review the model card.

19h agoHN ↗

It’s little to do with misinformation and much to do with trying to keep their model from collapsing from ingesting too much slop.

19h agoHN ↗

Great tech but the voice is like nails on a chalkboard to me

19h agoHN ↗

They should just give up at this point, it's just embarrassing to watch.

As PrimeTime said; these are the guys that invented the 'T' in 'GPT', that deployed their first TPU in 2015, that is using billions on AI - and they are beaten by 300 people startup named Moonshot AI even. People are going to write books about this complete fumble.

19h agoHN ↗

And yet they might become one of the winners "in the end" because they have near infinite money and others have not. I will drink tea and watch the show.

19h agoHN ↗

they have near infinite money and others have not

Given the very high margins on inference, once volume is large enough the other can also start printing enough money.

19h agoHN ↗

None of the startups are profitable. What exactly are they getting beaten at?

My advice is to listen less to brainrot 'influencers' that optimise for engagement through sensationalism.

19h agoHN ↗

Intelligence.

They have "unlimited" resources and has researched AI since the very beginning - PageRank is a form of AI even. And still, Gemini is behind Claude, GPT, Grok, Muse, GLM, Kimi and is maybe on par with DeepSeek?

As I said, it is embarrassing.

19h agoHN ↗

Grok

No one is behind grok. It literally has "be funny and irreverent when appropriate" (whatever the hell "when appropriate" means for them) baked into the system prompt. To me, that is all you need to know about how useful it is.

No serious people use it and the numbers bear it out tbh. It has the smallest market share of the "big companies" for a reason - and it's by a very, very large margin (~2.5% last I checked).

18h agoHN ↗

If RSI is achievable, it will leapfrog everything produced so far and so it will make sense to focus on RSI instead of incremental improvements for your top model. Startups need investment and need to show progress. Google does not at the moment need to take lead in the current race.

13h agoHN ↗

And yet they are burning through their famous cash on hand and taking on debt for data centers like everyone else.

12h agoHN ↗

Their Flash model is going head to head with the SOTA models. It's completely the opposite of embarrassing.

19h agoHN ↗

Their live models have been and continue to be at the frontier. I like them a lot!

19h agoHN ↗

Have you used Google search at all on the past few months? Every single search brings up a live chat prompt. They're serving fast AI to billions of users at huge scale everyday

And they're making money doing it.

Perhaps they don't have the best coding model right now (although 3.8 flash is arguably SOTA at some benchmarks), but is that the be-all and end-all of AI? Only coding matters?

18h agoHN ↗

It’s so annoying that everyone just points to the Artificial Analysis index (or even worse, Epoch AI, where part of the score is how good the AI is at chess) as a proxy for “how good” the model is.

12h agoHN ↗

Its 200-300 TPS. GPT is at 50-60. And its like 5% worse? Yeah that is a good tradeoff.

10h agoHN ↗

Are you watching/waiting on your agents? I care zero about t/sec, quality is far more important that quantity or latency, but they run in the background and I check in from time to time

10h agoHN ↗

Most people talking about Gemini quality do so by experience.

10h agoHN ↗

what's your solution? force people to use gemini until they like it? or hire them as googlers?

8h agoHN ↗

my approach is to not use closed weight models, I'm not trying to solve for how to get people to use gemini, quite the opposite actually

10h agoHN ↗

brings up a live chat prompt. They're serving fast AI to billions of users

Just because it shows up does not mean it is being used. I only click the feedback button to tell them how much I dislike their Ai summaries. They have removed that feedback button this week. Can't take the heat I suppose.

I've also never know anyone who finds them trustworthy nor heard someone say anything besides how they also dislike them

10h agoHN ↗

A lot of my friends will say "Well, the google AI says this" so they're using it, and maybe they don't trust the answers, but I think the thing is that they're using it.

"It's AI but anyway", seems to be a way to use it, but not promote that your using it?

There's a significant amount of people in my friend group who don't like "AI", but everyone seems to use what google's doing, even if they add the "the AI said this" disclaimer to it.

Even the people who don't like 'AI', will still reference the LLM output, which seems like they're just slow to accept it, I guess.

18h agoHN ↗

Strongly disagree with that take. Kimi K3 is a distilled model. I'm not saying that as a moral judgement, or to disparage the team behind it, but distilling and building on that is significantly easier and cheaper than building from the ground up.

And Gemini is kinda good enough at everything. Never the top, but it is decent at every task, and it is much faster than Kimi K3 and significantly cheaper. Kimi is very focussed on coding, Gemini isn't.

More importantly, it natively understands text, audio and video. If/when we are able to make the jump to robotics, this becomes essential. As you say, Google has a lot of deep background and deep pockets, they are able to make more of a long play. No idea if it will pay off, but it is way to early in the game to count them out.

10h agoHN ↗

Distillation does not make a model. There is so much more that goes into creating LLM Ai, traces from better models alone is insufficient. You need a good pre training run as foundation, a great RL reward function to make the other model's traces useful, and serious engineering chops to manage training runs at that scale. (non-exhaustive list)

18h agoHN ↗

Uh, what? This is SOTA for a live model. What are you talking about?

15h agoHN ↗

Yeah, you should definitely sell all Google stock you may hold, immediately, to me, before it's too late. I'm just a sucker, I'm happy to buy from you.

19h agoHN ↗

I'm disappointed with "Extended Thinking" for 3.8 Flash. On the plus side, it's a strong general-purpose model and the cost-benefit is still compelling.

However, the "Extended Thinking" should be renamed to "Slightly Extended Thinking". Considering that it's the maximum thinking option for Gemini Flash in the chat UI, it doesn't actually think a whole lot, leading to an uncomfortably high number of incorrect/poor replies.

12h agoHN ↗

Have you tried selecting the retry option below a reply? I think it let's you get a longer answer at least.

3h agoHN ↗

I used Retry plenty - but not for good reasons. Gemini fails quite a bit (error) in the app and chat interfaces.

19h agoHN ↗

So I am building a voice assistant to control AI harnesses, and recently tried switching from GLM 5.3 Flash to Gemini 3.8 Flash because of higher tok/s and better rate limits. Before that I also used Kimi K3 and DeepSeek-V4-Flash-0731.

Let me tell you unlike every other mentioned model Gemini 3.8 Flash trial had to be reverted the same day. Instead of simply delegating tasks it would invent additional requirements and implementation details it knew nothing about and no amount of convincing not to do it would work. That's the first time a model failed on me so spectacularly despite having practically same Artificial Analysis Intelligence Index as another model that just worked (and higher than working DS Flash).

The reason I think it is relevant is: Live is likely even stupider model in every way possible (except hearing better than separate STT). So beware using it for agentic scenarios.

19h agoHN ↗

My Gemini app is still stuck at 3.5 Flash-lite and 3.6 Flash so I truly don't understand how Google rolls this stuff out. I don't use Gemini for anything serious so I'm not going to use the API, but it's my go-to for just searching basic information (replacing google search) because it's so darn fast.

18h agoHN ↗

Yeah still on 3.6 here too, this is like the 4th or 5th model Google has announced since they last gave me access to the latest. And I pay for pro too!

10h agoHN ↗

I'm on AI Pro and have had 3.8 since announcement day, both in the Android app and in Antigravity CLI.

19h agoHN ↗

I have been looking for a model that's good for GUI testing. Original computer use isn't right because it's a slow screenshot loop, which doesn't capture transition and animation. Docs says this one does up to 1 FPS. That might be fast enough. If not now, we must be within a few months of high enough sample rates to do it.

17h agoHN ↗

I find just letting a model record a video it can read frame by frame later works fine, Gemini 3.8 flash would be solid for that

19h agoHN ↗

Did anybody watch the Primeagen's video on Google bag-fumbling? Interesting they released on the same day!

17h agoHN ↗

well, its because gemini is sucks at coding

8h agoHN ↗

I'm curious to know how do people in FAANG / Bay Area tech industry see influencers like Primeagen, Theo or Casey Muratori.

Do they have a good read on the industry or completely out of touch? Or to put it simply, are they spouting bullshit?

My concern is that some of them arent in the industry or have never been in it.

18h agoHN ↗

I love talking to chatgpt voice mode. Voice to voice AI is the only big leap that I see after the RL trained coding models.

18h agoHN ↗

My issue using the voice mode is the overly expressive mimicry of natural human intonation is distracting and starts to become extremely grating after a while.

There's one or two I find more understated but I would love a 2026 SOTA V2V model that speaks clearly but without the artificial personality layered on.

Human interaction/theory of mind relies so much on non-verbal clues for interpreting emotion/intent and so for me having those neurons firing constantly while talking to an LLM just for an emotional no-op is exhausting to put up with for more than a couple minutes.

There's one male voice that would make me assume someone was sarcastically mocking me if I was talking to an actual person because it's just so over the top.

17h agoHN ↗

One of the new Siri voice demos sounded like a lover whispering inuendo into my ear.

I don’t want to be aroused by my turn-by-turn street directions, thanks.

18h agoHN ↗

Just gave it a try - very solid release.

Copes well with thick accent, voices are pleasant and latency seems low.

Oh and I can actually use it on a workspace account - which for most of the recent releases was an account stuck in limbo. Not personal enough for personal offering, not enterprise enough for enterprise.

Well done G - will definitely be using this

18h agoHN ↗

Also appears to do well in other languages (prefer that when walking & talking in public for a bit of privacy)

And looks like one can trigger live mode via siri

8h agoHN ↗

Where did you try it, how can you use it? I have nothing in my Workspace (I am admin), Gemini app, AI Studio. Only 3.6 Flash

18h agoHN ↗

I'm wondering if Google intends to drop the next major version of Gemini Pro as a total bombshell drop to make Anthropic and OpenAI panic. They seem to be taking their sweet time on frontier model updates.

10h agoHN ↗

They promised 3.5 Pro at their next or i/o event, but the rumor is that is never going to be released because it would have been embarrassing. They just started letting their engineers use Claude, so it sounds like things may not be going so well with Gemini

18h agoHN ↗

Google need to allow saving history and exclude it as training data. I will not use it seriously until this is resolved.

17h agoHN ↗

Agreed. They really are the greediest when it comes to data for training (unsurprising for google I guess)

17h agoHN ↗

It is completely broken for me. After I ask a single question, it starts replying to itself in an infinite loop. It answers my question, then generates another reply to its own response, and keeps going. At some point, it even starts switching languages randomly.

17h agoHN ↗

From the demo video: "Welcome to the team, we're looking forward to working with you" is sooo creepy in a synthetic AI voice. In general, try not to have agents express sentiment that really should come from a human in your company.

17h agoHN ↗

3.8 is an incredible model even better is the infrs they host for it.

Is this a pure TPU infra? Really high performance solid intelligence.

17h agoHN ↗

My first language is Afrikaans, which is a somewhat niche language and hard to find teachers/conversation buddies outside South Africa. (I live in USA now)

I've been using Gemini to live chat in Afrikaans and do impromptu Afrikaans grammar lessons during my solo drives around town. It is phenomenal at speaking the language - like, it really shocks my family members when they hear it.

This is probably the most joy I get from any of my usages of LLMs/AIs. It's been really, really nice getting to speak my language regularly again. =)

So, I'm excited about this release and live chat getting better. I also hope the other frontier labs pick up niche languages like this as well so that I have more options.

16h agoHN ↗

Yeah Gemini has been consistently better than open ai's chat in icelandic, but I would still say it's far from passable as natural sounding. Lots of grammar errors and the pronunciation sounds like a non native speaker.

11h agoHN ↗

Makes sense as the origin of LLMs was Google’s work on language translation.

16h agoHN ↗

for me the live mode in androids google translate app has been as close as it gets to perfect for traveling cannot believe it is a free service after trying so many others

16h agoHN ↗

I have a similar experience using Gemini for quick Catalan translations for iOS apps given enough context.

I once asked it to summarize The Hobbit in Catalan to explain it to my daughter before sleep. I was expecting a lot of mistakes as I see regularly if I ask anything in my native language when using GPT or Claude, but it was surprisingly good. I was going just to kind of skim ahead and retell it my own way, but ended up almost saying it verbatim because it was good already.

She loves Zelda so I asked it to explain the story of Breath of The Wild keeping the original names, and to make it fun, etc.. I was surprised again. I did retell some bits in my own style and taste but it is very convincing.

I haven't tried Catalan on newer models like GTP-6 Astra or Fable tho. We have all these benchmarks based on software development, and AGI, etc.. but it would be cool to have some language benchmarks for different communities.

As I work in english and use them in english, I wonder if using LLMs in a different language to code renders a different result as well. Like, if some of these benchmarks were made in other languages, would the result be similar.

14h agoHN ↗

As I work in english and use them in english, I wonder if using LLMs in a different language to code renders a different result as well. Like, if some of these benchmarks were made in other languages, would the result be similar.

I am not a Chinese speaker but my understanding is that all of the models have substantially different behavior in Chinese, to the degree that it's kind of like a second model. Would be interested in hearing more if anyone has direct experience.

9h agoHN ↗

Fwiw, it's been very easy to nudge Qwen 27B/35B to think in Chinese. No need for <think> prefills like "思考:" ("Think:") or such.

Relatedly, when chatting with frontier models about science education content design, I've found it very helpful to mix in Chinese education terms. In English, for example, NGSS is such a massive attractor, discussing nearby topics often yields NGSS "slop". And "estimation" (educational) in the US means one (dysfunctional:) thing, which similarly distracts. Perhaps if AIs become increasingly multilingual, but remain weak at deep conceptual reasoning, it may be fruitful to have multilingual thesauruses, to use language-associated cultural conceptual differences as a way to convey conceptual nuances with which LLMs otherwise struggle?

9h agoHN ↗

In English, for example, NGSS is such a massive attractor, discussing nearby topics often yields NGSS "slop".

For that particular attractor, even German or Spanish or so might help avoid it? Or perhaps even just using British English?

1h agoHN ↗

Oh, good point! Something to check. I've been thinking of Chinese as a strength of Chinese models, but other languages might be sufficient here.

1h agoHN ↗

What does "estimation" mean in the US?!

"Estimation" as a skill in US education, primary school on, is pervasively about point estimates - single numbers. Bounding, a range of safe/reasonable/likely/possible numbers, is very rare. In Fermi questions/problems also.

Even in such goodness as Sanjoy Mahajan's Art of Insight and Street-Fighting[1]... the only only "bound" mentioned is "Printed and bound in [USA]", and there is no "round".[2]

In contrast, IIUC, China is pervasively about "Da Gu" and "Xiao Gu", "round up" and "round down". Point estimates are secondary. Lower and upper bounds are apparently even a cultural concept: an individual's degree vs capability/fate; policy equity vs excellence; economic targets.

In education, my understanding is point estimates have much poorer conversational/collaborative dynamics than range estimates. Bounding lends itself to incremental collaboration. "Can anyone suggest another low or high bound?". Discussions around point estimates seem usually merely about the collection of individual estimates.

NGSS

Ah, NGSS isn't bad particularly. I meant that discussion/mention of NGSS appears so very much in English training data, that AI chat on topics related to NGSS seem to draw regurgitations of NGSS. AI "slop" in the sense that "ignore NGSS, just reason from a superset of its underlying concepts"... isn't an LLM strength.

hard to understand

Sorry - Thanks for asking!

[1] open access: https://direct.mit.edu/books/oa-monograph/5345/The-Art-of-In... https://direct.mit.edu/books/oa-monograph/5339/Street-Fighti... [2] search "bound": https://www.google.com/books/edition/The_Art_of_Insight_in_S... https://www.google.com/books/edition/Street_Fighting_Mathema...

4h agoHN ↗

The newly opened possibility to learn“how non-English cultures do X” is criminally underutilized.

15h agoHN ↗

I love Gemini in Google Maps for long drives. I start getting bored of music and podcasts and start grilling it with random questions I've always wondered about.

15h agoHN ↗

I do this too! Usually some idea in ML ...when I am driving I might suddenly remember what I was thinking about and then it is my personal podcast via Gemini-in-Maps. My only complaint is if you have follow up questions, you have to be quick, otherwise it cuts off the mic.

15h agoHN ↗

I feel like a 10 year old all over again with my frequency of questions.

How did the early Roman Empire interact with Greek city states?

How are LLMs planning on learning new information on the fly without new context or retraining runs?

Why did mom leave?

You know, standard stuff.

14h agoHN ↗

Why did mom leave?

Reinforcement learning by human feedback.

4h agoHN ↗

This would be horrendous if it weren't so true.-

4h agoHN ↗

Why did mom leave?

Even as a human being who has a working brain and all this thinking power, I am not sure how I would answer that. It's kind of funny Gemini will still come up with something and even suggest follow-ons to the conversation but saying all that stuff as a human to another human would be a wild response to that question.

13h agoHN ↗

I do the same for improving my English skills. It's amazing!

12h agoHN ↗

In Maps? Or is it Android Auto? I have Google built it. When I push the button on steering wheel I hardly know which layer I'm talking to.

10h agoHN ↗

Really? I am in Android Auto and I try to even just ask it to get me to "Costco Near Me". I know where it is. I wanna know if the sometimes there traffic jam that blocks the right lane is gonna expect me or not without being distracted zooming out from current location and then back in to the highway exit I know is prone to that at times. You know, I wanna be a good citizen and not be distracted by looking even a second at the screen I'm zooming around in.

It tells me there's a Costco 10km from me and I'm like, yeah let's go there, that's the one I meant! And then it sends me to one that's like 40 minutes from me all across town and I'm like WTF?!

Never mind asking Android Auto any question that's not driving related. It just doesn't understand me at all. Or even some driving related ones. "Find Alternative route". "Sorry, I can't help you with that". WTF?

9h agoHN ↗

Does you Android Auto show a microphone for voice our the Gemini star? I think you're likely talking to the pre-Gemini voice assistant which is very basic.

8h agoHN ↗

How to tell if someone doesn’t understand selection bias…

15h agoHN ↗

I've been building a language learning app which uses AI voice chat to let people practice. Would that be something interesting for you to try out? It's all still rough, but we've been learning Japanese with it and its pretty cool.

15h agoHN ↗

I want to, I have like a 2000+ day streak on Duolingo and I know I can only gain fluency once I start speaking it more.

14h agoHN ↗

We (me and my teens) used Duolingo for a few years and.. didn't really get that far in actually being able to speak. It's kind of mind-blowing that they are the market leaders. In retrospect, I think of it more like an addictive language game and not really a good way to actually learn to speak.

If you're really interested, shoot me an email.

10h agoHN ↗

Latest in a long line of scams for learning a language. Goes back to cassette tapes and records. The only good way is the natural way: immerse yourself in it.

6h agoHN ↗

Some cassette tapes work very well. Pimsleur for example.

You can use it as supplement or do it standalone on your commute before a trip and end up with good pronunciation, vocab and a conversational skills.

50m agoHN ↗

While in general I do agree with you (it's how I learned Italian, Hebrew, and Thai) being on the ground for a year is cost-prohibitive for me right now, and unreasonable for kids under 16.

What I've been building, with the help of a few language teachers, is our best guess at a really good way of learning new languages from the comfort of your home. It's not ever going to be perfect, but it is pretty good (at least so far in testing within our small group), and very cost-effective.

If you're interested in funding it, I'm happy to go live in Japan for a year :)

14h agoHN ↗

Ai pitch accent is not correct with most TTS systems for Japanese. I’ve tried many.

15h agoHN ↗

My father-in-law was saying the same thing about Gemini 12 months ago regarding the Afrikaans speech. I've tried a few of them in my studies but none of them could seem to switch between English/Afrikaans except Gemini. It's an interesting time to be a language learner

As jy wil, ons kan saam praat op Discord :) maar my Afrikaans is sleg

6h agoHN ↗

As a Dutch person i find Afrikaans always interesting whenever i happen to encounter it in the wild. As if a common ancestor took a very different evolutionary path.

6h agoHN ↗

Given that Dutch is the ancestor here (not that Dutch hasn't changed since the split), I wouldn't call it "as if".

14h agoHN ↗

It’s literally translation technology told to guess what’s most likely next instead of translate what it was given.

It makes total sense it’s good at regurgitating the edge cases of language.

5h agoHN ↗

Did this with shona.On the road trip we (3 human occupants) spoke to Gemini in ndebele, karanga, portuguese and ndau and have switch seamlessly to another lanuage like german without loosing the thread of the conversation. what impressed me was its ability to recognize the subtle dialects of shona being spoken, articulate its unncertainty, ask for clarification and carry on. when asked to speak those dialects it politely lamented it could not. when pushed it insisted and basically apologized in a technically correct but awefully pronounced sentence of the dialect before going back to english.

8m agoHN ↗

Hey I literally just did the same thing after seeing the Afrikaans reference ndikawona kuti zvirikushanda! I was incredibly surprised, its shona was really formal but that makes sense when thinking about getting training data.

3h agoHN ↗

I like it. I had a friend at school who came from SA and spoke Afrikaans - so I know how a natural sounds like.

For language learning, translation, research - it is fantastic.

It works very well for certain dialects as part of a local language heritage as well: From Welsh to Bavarian - it is a joy to really get approval from people used to these dialects as being really accurate.

Happpy times! :)

2h agoHN ↗

Lekker man. Although I speak it often, I could use an LLM to help me progress as I've been stuck at my current level for about a year and typing it is not something I enjoy.

I find no greater joy than listening to a TTS model say Poes (Google translate is one of the funnier ones to make swear). Reminds me of being 14 lol.

1h agoHN ↗

I am assuming these languages are present in translate.google.com If Google Gemini is good in languages other than English, while other models are not, then we can trace back it back to the translate.google.com as a resource as one possibility.

17h agoHN ↗

I've been using Gemini as a life partner. Talking to it about my feelings thoughts, plan of action and the like. It's great. I've anthropomorphized it and put the computer speaker in a doll's mouth so it seems like it's a real baby.

Looking forward to where this can go.

16h agoHN ↗

I’ve noticed antigravity become significantly slower over the last week or so. It’s a bummer because speed is what I care about. Qwen3.8-27b seems on par with 3.8 flash so if I don’t get speed out of a paid service I’ll just use my local model. Sad.

15h agoHN ↗

If you'd like to experiment with Gemini 3.8 Live on a US telephone number, you can try Wokay [0]. It's an agentic memory demonstration that's built with LiveKit and Gemini. Tel: 408–897–4019

[0] https://wokay.goodmem.ai/info

10h agoHN ↗

Big Ai is having a moment of reflection and reckoning with open weights, I wouldn't bet either way on a new Gemma release, 4 was distilled from the Gemini 3 series

1h agoHN ↗

Lack of releases relative to Gemini. Not that I expect them to take down the old versions, just unsure about new ones.

12h agoHN ↗

i've been using gemini (api & pro) since last november last year. for creative writing compared to other llm, it's the best in capturing local nuance, it can even create jokes in my languages. But that just it, i cant rely on other work, hallucinate too often, the deep research are not reliable at all. Too many discussion i've had that it grasp main concept consistent but the supporting concept just plain hallucinate and not consistent. it's tested between pro & flash. This doesnt happen often on open weigh

11h agoHN ↗

I swear all day long I was thinking that there will be a new release soon because Gemini 3.1 Pro performance was in the crapper (both via API and Chat)

10h agoHN ↗

Still so expensive. I wish Google could compete with GLM.

10h agoHN ↗

Where do you see a list of which languages it supports?

10h agoHN ↗

Now Gemini can’t make a summary of a YouTube video on his own. I need to give it the transcript.

10h agoHN ↗

I don’t understand good experiences people are having with Gemini. It’s the only model that sometimes loses/forgets context in literally next message. Plus feeding unasked product links to responses.

9h agoHN ↗

I wonder how much life DeepMind has left in it, especially after Hassabis's departure. Google execs must be having discussions about simply throwing their weight behind Anthropic since they already own so much of the company.

8h agoHN ↗

I agree. Lot's of people really like it. For me it often just forgets all context and starts showing random slop. It's super clear as, when I ask it what happened to some element earlier in the conversation it tells me it does not have that. It might be good if it told me, but randomly lose the plot is quite frustrating.

Claude does it occasionally but it's a more a soft landing earlier context seems to be compacted, not completely lose the plot.

I just cancelled my pro subscription. I really wanted it to be good but not yet.

8h agoHN ↗

I strongly agree. I suspect it's people who have not yet used the paid models from OpenAI and Anthropic. Gemini is comparable to free models from other providers, but not in the same universe as paid models.

This is frustrating because when I discuss AI with laypeople they think it's still incapable of counting the number of Rs in "strawberry." They believe it to be essentially useless and incapable of basic tasks. Which, to be fair, is the case with the free models.

6h agoHN ↗

you're just not using the latest model, bro

Pro tip, ChatGPT is the normiest of all normie websites right now. You're not part of the cognoscenti just because you learned how to type prompts into one of the most popular websites in the world.

P.S. You're probably not using OpenAI models for complex or non-standard tasks. It shits the best just as often as Qwen when you need precision and detail in a non-obvious problem.

6h agoHN ↗

Where did he claim otherwise?

Take your meds.

6h agoHN ↗

You're probably not using OpenAI models for complex or non-standard tasks. It shits the best just as often as Qwen when you need precision and detail in a non-obvious problem.

The benchmarks clearly show otherwise. This is your cue to tell me the benchmarks are made by the Illuminati and only your superior and subjective methods of evaluation are correct.

5h agoHN ↗

The tasks you actually need to do trump benchmarks, yes. I haven't tried out the ridiculously expensive models besides the latest Gemini, and it gave from equal to slightly worse results than latest DeepSeek, at a far higher price.

It does also seems Gemini's main problem wasn't that it was stupid, but that it was good at doing slightly different things than what I asked it to, very well. Which might well also have to do with me being better at wrangling DeepSeek's quirks than Gemini. Still, at that price tag, it's not worth it.

3h agoHN ↗

I think 3.8 Flash is on par with DeepSeek on some benchmarks and tasks (not coding or design), but it's not close to Sol/Astra or Opus/Fable. I would not consider a $20 subscription "ridiculously expensive," but I suppose that is a relative term.

2h agoHN ↗

I would run into the use limits very quickly, and (for Anthropic) have to switch frameworks.

By all accounts they are far more expensive than DeepSeek, and vs. Gemini I've found out that myself.

2h agoHN ↗

No argument that DeepSeek is much cheaper, but you certainly get what you pay for.

2h agoHN ↗

artificialanalysis just updated their benchmark after the release of GPT-6. They removed old, saturated benchmarks and replaced them with new, until GPT-6 floated to the top with the cream. One of those new benchmarks is AutomationBench-AA, where GPT-6 had a clear lead. Today that benchmark is topped by DeepSeek v4.1 Flash.

Edit: For those who are not familiar with it, this model is quite a bit faster, and about 100x cheaper, per token, than Fable and Astra.

4h agoHN ↗

Impressive how feverishly you defend the steaming pile of shit that Gemini is.

6h agoHN ↗

I suspect it's people who have not yet used the paid models from OpenAI and Anthropic. Gemini is comparable to free models from other providers, but not in the same universe as paid models.

I totally disagree. I pay for both ChatGPT and Anthropic (haven't tried the chinese models yet) and yet Gemini is my go-to model (and I pay for it too through Google Workspace subscriptions for several domain names tied to Google/GMail) for anything that is not coding.

I find Gemini better/quicker/more polished for basically every single subject out there that is not "write me lines of code".

6h agoHN ↗

To be fair, it has been about six months since I tried a paid Google model. I will give it another go to compare. Hallucinations were the main issue back then but perhaps it has come a long way.

5h agoHN ↗

Much respect for being able to admit that you haven’t used a model in a while after saying folks probably haven’t used the models you use.

3h agoHN ↗

Okay I just tested 3.8 Flash (High) on some real world problems I have used Opus (High) and Sol (high) to solve. The outcome here is terrible.

One of the problems is local zoning laws regarding an expansion of my house. Comparison of annex vs extension, boundaries, precedent, costs, etc.

3.8 Flash didn't check most of the required zoning laws. It relied on parametric knowledge, which is outdated and inaccurate. It checked zero precedents. It did made a very cursory check of the boundary area, but didn't validate it, so it missed a lot of important nuance and exceptions to the boundary. Its cost estimates were wildly inaccurate. Ostensibly because it was inferring an average based on historical pricing data rather than gathering current info.

I could go on, but if I had to judge this attempt I would give it a 3/10. It's very fast, but wildly inaccurate. It's clear that the model is designed for speed over accuracy.

But don't take my word for it. [Most benchmarks show it to be significantly below frontier models like Astra.](https://llm-stats.com/models/compare/gemini-3.8-flash-vs-gpt...)

This has been a useful exercise. It's important to understand the developments taking place. I am disappointed to see that Google has made very little progress in six months relative to the frontier labs.

2h agoHN ↗

Gemini's search harness in the Google app is (ironically) bad so it makes the model look bad.

If you really want to compare apples to apples you need to test Gemini models against other models using the same third party search harness.

Otherwise you are largely measuring how much computation the model provider is allocating to a search harness.

4h agoHN ↗

Same here.

I use gemini for everything not-coding, from doing research, to have custom personas for more niche topics (and feeding more detailed knowledge in these cases).

For coding and image editing, right now I find ChatGPT superior. And for software architecture designs or planning Claude is the best since a while. I still have to try Grok to be fair.

56m agoHN ↗

I've got the 100 paid to all 3, gemini and chatgpt are a level above claude for a lot of my work now. claude i actually fight with if its not just write code.

7h agoHN ↗

i have programmed a dozen scripts with google search AI lol

they'll live in my rc for decades

6h agoHN ↗

Ah yup for quick one-off one-liners or small Bash script, Gemini works perfectly fine too. For longer scripts I use another model.

6h agoHN ↗

I love it as a variation from the others. 3.8 flash is the best back-and-forth model for iterating imo, but would not use for long horizon

6h agoHN ↗

Have you tried Gemini 3.8 Flash recently ?

I was like you before, Gemini was the worst model to me.

Then 3.8 came out. At first I was sceptical, but this model *is* able to do useful things ! Complex things.

Of course, it is NOT perfect. But for things like small/medium complex tasks subagents, it's perfect.

Now, is it worth the money vs Astra ? I don't think so. But still my point remain relevant.

5h agoHN ↗

They do a lot of weird things with the context in their user facing products like Gemini. Google seems to always have trouble with their harnesses

3h agoHN ↗

I also had that issue happen to me, surprisingly only when I used Polish, not English. But other than that I actually like Gemini, I check stuff against it all the time, especially when walking my dog. It also works in AndroidAuto for me, but I use very basic stuff like changing Spotify music. I was on iOS before and there's no comparison to old Siri that was just garbage. I use Gemini practically every day, it's fine for the most part IMO. They really do need to polish integrations though, app connectors barely every work outside of Google's own apps. Using Oppo's app Mind Place through Gemini f.e. is just bad experience and almost never works. For example, I had pleasant experience of Gemini finding me places for a walk/hike on vacation in Tirol, where I specifically required not too much of an ascent and an asphalt road for a stroller.

3h agoHN ↗

I’ve found this to be true as well. It’s got really poor attention and will derail into world-building fast. I’m convinced Google is just shipping it to capture market share but they know Gemini isn’t ready for serious use.

2h agoHN ↗

For information retrieval and collation related tasks (reading lists, deep dives on subjects, etc.), Gemini is way better w.r.t. other models in my experience.

This difference is probably due to Google's web knowledge and free pass to YouTube. However, I'm happy what I got from it so far.

When I ask the question once in a blue moon, it can generally one-shot the answer, even.

8h agoHN ↗

same experience here. Benchmarks don't capture the magic of having a 24/7 native-speaking conversational partner in a rare language. That alone makes these models worth it

7h agoHN ↗

Maybe now it will call my wife when I ask it to...

7h agoHN ↗

Even if really great, it still can’t use tools. The only consumer Google thing with access to MCP tools is Gemini Spark and that has lots of other problems. I wish they would combine their efforts on a great consumer product but it’s Google we’re talking about…

I still dream of the day that we get full tool parity in voice and text mode so your voice assistant can do everything you connect for you. Grok and Claude are btw almost there, only a very few minor built-in tools don’t exist in voice mode, but I already use both to connect to heaps of things! It’s so valuable to verbally discuss something with an agent, have the agent pull in context from GitHub, Notion, email, and then create artifacts somewhere

4h agoHN ↗

Note sure what you mean by "tools". It does support function calling, so you can hook it to whatever tools you want. I'm using it with Home Assistant to control my smart home devices.

1h agoHN ↗

I mean the consumer Gemini app. It has no way to connect MCP servers to it. Only Spark can, if you switch your Gemini to the separate Spark mode

6h agoHN ↗

I often use Gemini as a third backstop and it never fails to disappoint.

4h agoHN ↗

These are not the Gemma model family you might be thinking of. Gemini models weights are never made open.

3h agoHN ↗

they are working on it non stop huh, well good for us as they keep things free or cheap

3h agoHN ↗

Gemini is the most sycophantic and world-builder of all LLMs I’ve used. I’m sure all models do it but Gemini seems to have no guardrails and it will happily induce AI psychosis on you sooner or later. Just no.

1h agoHN ↗

The voices sound really pleasant and realistic. Sadly it doesn't support SIP. I found forwarding streams with websocket for phone calls tend to introduce some unwanted latency that's quite noticeable in a conversation. I've found a lot of success with GPT-live-1 so far, but the generated voices are lacking something I can't put my finger on.

1h agoHN ↗

What's your use-case for automating phone calls?