Curious about people’s experience here. I am working on a small model, verify by jev, and escalate to big model. Some cases, the small model is not a model but some regex.
I'd be interested to see if using DiffusionGemma-as-Jev helps as you can feed the image directly into the model and it'll make decisions based on the image embeddings.
I had a long talk with chat gpt about this today as well. I think its duable and prolly not too hard either, also you could do lotsa funky stuff with stitched frames of a video in one 4x4 grid for example and send that as one image for analysis. that way temporal understanding can be had for fractions of a second by jev... also because vlm works in pixel space you can get around the whole state machine issue as well, so many possibilities...
And every comment on HN calling out vibe-coded slop has this same complaint comment in turn. OP has an actual complaint, your complaint is just "stop complaining".
I appreciate it when people say something is slop. Saves me from wasting my time looking at it.
How many slop ideas have come to the front page, never to be heard from again because the execution isn’t actually any good? I’m guessing most of them.
Super cool. Hope waitlist will move soon. I have a use case for it too.
are you the author? If so - what are your notes on using Jev in this scenario?
Did you do the follow up questions? I was invited within a few hours of joining today.
It's also now on openrouter and cloudflare
Same here. 8h hours later I was in.
Curious about people’s experience here. I am working on a small model, verify by jev, and escalate to big model. Some cases, the small model is not a model but some regex.
cheap-confirm-escalate
Using jev as the confirm step.
How does it do on OSWorld-verified? Recently read that even Fable 5 is just at 85% .
This probably throws a spanner in the wheels there:
EDIT: Not to shit on this though. I totally believe that some smart mixture of LLM-reasoning + Jev-style + determinism is going to be pretty amazing.
I'd be interested to see if using DiffusionGemma-as-Jev helps as you can feed the image directly into the model and it'll make decisions based on the image embeddings.
I had a long talk with chat gpt about this today as well. I think its duable and prolly not too hard either, also you could do lotsa funky stuff with stitched frames of a video in one 4x4 grid for example and send that as one image for analysis. that way temporal understanding can be had for fractions of a second by jev... also because vlm works in pixel space you can get around the whole state machine issue as well, so many possibilities...
I trailed off a few lines into the README. No human ever edited any of this. « LLM detected, project rejected ».
I stopped reading at “The honest caveat”…
Every third post in HN has this same complaint comment. Everybody knows.
And every comment on HN calling out vibe-coded slop has this same complaint comment in turn. OP has an actual complaint, your complaint is just "stop complaining".
I appreciate it when people say something is slop. Saves me from wasting my time looking at it.
How many slop ideas have come to the front page, never to be heard from again because the execution isn’t actually any good? I’m guessing most of them.