Show stories

Live mirror
30 storiesupdated 0s agoView source snapshot
  1. Show HN: An e-ink frame that hears birds and draws them as 1800s illustrations(github.com/arnegiacomo ↗)
    168comments
  2. Show HN: Capsule – Single-file web apps that save their data into SQLite(withcapsule.app ↗)
    113comments
  3. Show HN: Pizza Bot – An inbox for AI agents that work in the background(github.com/pizza-bot-app ↗)
    3comments
  4. Show HN: Agenttik – work on multiple projects in parallel with AI agents(github.com/pausan ↗)
    discuss
  5. Show HN: Hacking a $20 4G wireless hotspot into a texting device(bkovac.github.io ↗)
    30comments
  6. Show HN: Panel – A research workspace where the agent can build its own panes(github.com/greentfrapp ↗)
    11comments
  7. Show HN: DeadLock OS – A Linux terminal puzzle game(crazygames.com ↗)
    discuss
  8. Show HN: Check if your IP has appeared in a residential proxy network(haveibeenproxied.com ↗)
    34comments
  9. Show HN: Ordewell – turn one goal into an ordered plan of coding-agent tasks(github.com/ordewell ↗)
    29comments
  10. Show HN: HonestSky, a free iOS weather app without ads or subscriptions(apps.apple.com ↗)
    discuss
  11. Show HN: Redis City – Explore how Redis works in an interactive 3D model(poltora.dev ↗)
    25comments
  12. Show HN: DaiDocs, AI memory as a plain-text file format, not a service(github.com/kerneta ↗)
    1comments
  13. Show HN: Loss. a tiny satire about AI progress(workatloss.com ↗)
    7comments
  14. Show HN: SCIP MIP solver bindings for Go, ported from russcip(github.com/egoisutolabs ↗)
    discuss
  15. Show HN: farseer.space – fly anywhere in the universe in your browser(farseer.space ↗)
    1comments
  16. Show HN: Macros with a Behringer FCB1010 MIDI Pedalboard in macOS(github.com/jamesryanatx ↗)
    21comments
  17. Show HN: Sass – Rust and WASM(github.com/zoosky ↗)
    discuss
  18. Show HN: Warp – Run DeepSeek v4.1 Flash with 5 GB of RAM at 3.77 tok/s(github.com/sqliteai ↗)
    2comments
  19. Show HN: Kinesis – Control your Mac with the Meta Neural Band(github.com/callbacked ↗)
    45comments
  20. Show HN: Pelican-bicycle alternatives(gally.net ↗)
    45comments
  21. Show HN: Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost(narilabs.com ↗)
    31comments
  22. Show HN: Open-source passive NFC tag that signs with ECDSA, verified on-chain(github.com/mwbpnftechnology ↗)
    discuss
  23. Show HN: Let agent read files with secrets while redacting values for LLM contex(github.com/daniel-sc ↗)
    1comments
  24. Show HN: AirmailAI, a BYOK LLM chat app with a browser extension backend(airmailai.net ↗)
    6comments
  25. Show HN: TrailVid – Cinematic Travel Animations(trailvid.com ↗)
    discuss
  26. Show HN: Omni – Open-source workplace agent, built on Postgres
    discuss
  27. Show HN: TabPFN-3.5, a Tabular Foundation Model for messy real-world tables(priorlabs.ai ↗)
    1comments
  28. Show HN: I ported COLMAP (photogrammetry) to the browser in WASM and WebGPU(offlinetools.io ↗)
    discuss
  29. Show HN: I built a tiny camera that knows where it is(mightycamera.com ↗)
    4comments
  30. Show HN: Neobrutalism.dev – Just added Base UI support and added new color theme(neobrutalism.dev ↗)
    76comments

Show HN: Pelican-bicycle alternatives

128 pointsby 1d agogally.net
45 comments
In November and December 2025, inspired by Simon Willison’s pelican-riding-a-bicycle benchmark, I had some then-current LLMs create SVGs from thirty similar prompts, such as “Generate an SVG of an octopus operating a pipe organ.” Simon mentioned that experiment on his blog [1].

Nine months have passed and much stronger models have been released, so I tried the experiment again today. The linked site shows the results.

Running ten of the prompts through six models at OpenRouter cost about twenty dollars, so I stopped there for now.

[1] https://simonwillison.net/2025/Nov/25/

1d agoHN ↗

Feels like google has a different training set than the others?

1d agoHN ↗

All models are pretty good now at generating these images. Back in the day, I remember experimenting with the pelican images and most of the models couldn't align the legs with the wheels. Right now as well, GPT messed up an octopus leg by originating it through the instrument rather than the octopus itself.

I think that intertwining two entities (living/non-living) is still challenging but overall they're pretty sound.

1d agoHN ↗

All models are pretty good now at generating these images.

Not zebras. If you want to see how bad SVG output still is, ask for a zebra riding a scooter.

1d agoHN ↗

Man I thought the giraffes were bad enough

1d agoHN ↗

All models are pretty good now at generating these images.

That's pretty generous.

1d agoHN ↗

Website looks very cool, Fable's octopus-organist looks very cute, but I feel like this benchmark (generate an SVG by a short and slightly ridiculous description) in general has been completely Goodharted [0].

I think they all just added a bunch of similar tasks to their training sets, so we cannot judge true emergent capabilities of the models anymore.

[0] https://en.wikipedia.org/wiki/Goodhart%27s_law

1d agoHN ↗

I don't think a bunch of similar tasks can really saturate the "create a SVG of X", because the model should have a quite good spatial understanding of the world and how everything interacts.

For example Gemini 3.8 Flash seems very impressive at first glance but the results are not actually very coherent, this shows that its "world model" is not particularly great (compare to SOTA models).

1d agoHN ↗

“I think they all just added a bunch of similar tasks to their training sets, so we cannot judge true emergent capabilities of the models anymore.”

This has been the plan since the start of all this, they regurgitate code in ever better forms but they still aren’t inventing new things yet.

19h agoHN ↗

In my experience LLMs have become generally a lot more useful if you need to create something like a company logo or even 3d scenes.

1d agoHN ↗

Interesting, out of all examples Gemini 3.8 is the best for me. Also the image style is different and more vibrant than others

1d agoHN ↗

Better than Astra and Fable? It looks quite pretty and even impressive at times if you squint, but look closer and it falls apart in terms of coherency. And I say that as somebody who mains Gemini 3.8.

1d agoHN ↗

An elephant typing on a typewriter

A monkey, surely?

1d agoHN ↗

Asking to animate it add an interesting layer of difficulty

1d agoHN ↗

I noticed in the "A penguin juggling chainsaws" prompt that Qwen created an animated svg

1d agoHN ↗

I'd say at least half of Qwen's 2026 runs are animated.

The only other one I've spotted is animated is Gemini 3.0's 2025 run of an elephant.

1d agoHN ↗

I'd like to mention the Little Dorrit Benchmark [1] which I have been running for a couple of years now. It has a few nice features:

1. It tests visual reasoning and structured output in a single task.

2. It seems to sort correctly on advancing general intelligence. As a counterexample, if I'm not misremembering, artificialanalysis.ai made some changes to their benchmark recently after Astra ranked below several older models.

3. While models have gotten significantly better in the past 2 years, the top model is still at 0.78 F1, so the test is not yet saturated. As a reference point, when I started, the top models were in the [0.1, 0.2] range.

[1] https://dorrit.pairsys.ai/

1d agoHN ↗

Does a test of instructions how to fold origami figures in a SVG/jpeg exist? Or could be useful?

1d agoHN ↗

Well that benchmark is now saturated, what next. How fast you can hack the pentagon?

1d agoHN ↗

Gemini 3.8 flash seems to (subjectively) be the outlier in terms of performance to cost ratio?

1d agoHN ↗

Gemini 3.8 Flash results are often not very coherent but it does put a lot of shading and details to hide the fact.

1d agoHN ↗

Its giraffe / grandfather clock one is pretty bad... (two necks? wearing a suit?)

Weird as well, it's clearly pulled out some 1884 patent on clock designs, and a quick ddg/google doesn't show it as anything to do with grandfather clocks.

1d agoHN ↗

It’s very likely they all add svg generation into the training data. It’s part of the reason it’s no longer a good benchmark (unless you need to generate SVGs).

1d agoHN ↗

It's about more than just training data. Every model since GPT-2 had SVGs in the training data, because they all used a scrape of the Web and the Web is full of SVGs. What's unique about Gemini is that they actively worked to get better results for their SVGs - maybe RLHF, maybe RLVF of some sort.

1d agoHN ↗

The 3 US models have their own style.

Qwen3.8 is very clearly distilled from Claude models.

17h agoHN ↗

How do you support this claim? I am curious about tell tails for distillations.

1d agoHN ↗

It is interesting how similar the designs are across the models.

1d agoHN ↗

Because it's one of the few ways of comparing models that lets you instantly evaluate them visually. That makes it more comprehensible than a numeric score on a benchmark.

1d agoHN ↗

I notice none of the octopi seem to be actually facing the organ.

1d agoHN ↗

Try asking an LLM to draw you the cool S.

1d agoHN ↗

Or something like, "Draw an S, then a more different S, close it up real good here, then using consummate V's, add teeth, and scales, and eyebrows, and legs. And then add smoke, and fire, and some wings, and one of those big beefy arms for good measure." :D :D :D

1d agoHN ↗

Would love to see Qwen3.8-27b here, since that is the model most people are running locally.

1d agoHN ↗

Anyone else surprised the generations look so remarkably similar? All of these models have “independently” generalized that the moose should roughly be standing at the same position (left) or that the giraffe should have a certain color palette.

1d agoHN ↗

With respect to bikes (with or without pelicans) there's a strong natural bias because people displaying bikes tend to want to show off the side with the gears.

More-generally, I suspect an influence from how left-to-right languages (i.e. English) affect comic layouts. Overcoming that bias often means using vertical space to exploit the top-to-bottom habit instead. (Consider the rarity of an English-language comic panel where action is from bottom-right to top-left.)

1d agoHN ↗

Yes; I found it fun they all decided telescopes should be repaired at night.

1d agoHN ↗

Love the output from Qwen 3.8 it seems very impressive for the cost! Why does Gemini 3.8 flash blur everything? What are Google playing at!

1d agoHN ↗

Gemini 2.5 Pro is the only model with a sense of where a ferris wheel operator would be.

1d agoHN ↗

I hope there would be parameter iterations allowing especially cheaper models on inspect and fix their output via rendering

15h agoHN ↗

What's interesting to me is that the SVG versions don't have the "AI Image generation hates negative space" issue as badly as generated images do. They kinda stay on point and don't fill every single empty bit with some pattern.

9h agoHN ↗

Qwen 3.8 Flash-Next is really missing here (because it's small enough to run it locally on a (formerly) affordable machine). This is what Qwen3.8-Flash-Next-UD-Q4_K_XL gives me for "an octopus operating a pipe organ": https://imgur.com/a/zHyHIqI Seems very similar to the output of Qwen 3.8 Max to me