Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Linus demands real users before hazard pointers land(freenode.net ↗)
    discuss
  2. Here's how we're (actually) all going to die(siliconpalace.substack.com ↗)
    discuss
  3. Want to Entice New Residents? Offer Cash, for a Start(nytimes.com ↗)
    discuss
  4. Relativistic Raytracing(publish.obsidian.md ↗)
    discuss
  5. Memoization application that you have used but might not know(github.com/prashant2400 ↗)
    discuss
  6. ZCode, embroiled in a controversy over stealing user code, is now open source(github.com/zai-org ↗)
    discuss
  7. Show HN: Ambits – agentic grep/rg tool will history tracking(github.com/joshlong145 ↗)
    discuss
  8. Coding Theory: A Playful Introduction(paramrathour.github.io ↗)
    discuss
  9. Scammers found a way to make people drain their own wallets(twitter.com/wyckoffweb ↗)
    1comments
  10. Kelvin Wave(wikipedia.org ↗)
    discuss
  11. AI Weekly Warns Firms on Google AI Studio Data Retention Fraud(bitu79.substack.com ↗)
    discuss
  12. Art of the Problem Launches $99 AI bot(artoftheproblem.com ↗)
    discuss
  13. Have a Question? Ask Jev(askjev.net ↗)
    discuss
  14. Ax: Google's Open Agentic Orchestrator(github.com/google ↗)
    discuss
  15. The Cultural Leadership Fund (2018)(a16z.com ↗)
    discuss
  16. Farhud memories: Baghdad's 1941 slaughter of the Jews(bbc.com ↗)
    discuss
  17. Farhud(wikipedia.org ↗)
    discuss
  18. Meta launches fresh legal challenge over UK's Online Safety Act(ft.com ↗)
    discuss
  19. 'Ask for what you want' is a key skill for the 21st century(robinsloan.com ↗)
    discuss
  20. A Brief History of Light(psu.edu ↗)
    discuss
  21. What Is Coastline Index and Why Focus Only on Confirmed GTA VI Facts?
    discuss
  22. World model companies are keeping a lot of secrets(techcrunch.com ↗)
    discuss
  23. AI is breaking the academic sorting machine(lemire.me ↗)
    discuss
  24. A Fresh Start: My Minimalist Blog Redesign After 14 Years(bcastell.com ↗)
    discuss
  25. Reviving Deserts with Syntropic Agroforestry(wikifarmer.com ↗)
    discuss
  26. Delta (zed.dev) – Help understanding ToS(zed.dev ↗)
    1comments
  27. Can I Let My AI Agent Run on Shabbat?(chabad.org ↗)
    7comments
  28. Deterministic grounding checks for LLM agents, no LLM in the hot path(github.com/polarisbuiltinc-wq ↗)
    discuss
  29. When Intelligence Becomes Abundant, Architecture Becomes the Constraint(henrymurphy832344.substack.com ↗)
    discuss
  30. Deterministic Core, Non-Deterministic Shell(outdata.net ↗)
    1comments

I turned Jev into a (lousy) chatbot

98 pointsby 8h agogithub.com
33 comments
8h agoHN ↗

It's the digital equivalent of Morty speaking with the death crystal: https://youtu.be/YjepJlvkdKs?t=51. The crystal shows him how he will die, so he iteratively determines his speech based on whether he sees himself dying with the life he wants.

7h agoHN ↗

Sometimes I think I'm Morty speaking with the death crystal but I don't even have a death crystal and all I fear is life itself.

8h agoHN ↗

Try returning a short list of word options by lookup based on the current word completion. Would save some turns.

8h agoHN ↗

> write me a short story

a story

I think this is the first time I’ve knowingly laughed at a model’s joke.

6h agoHN ↗

In the olden days when the text-davinci models would just wholesale make shit up they did some pretty funny completions.

An early ChatGPT model made me a pretty funny track list for my imaginary Indian cover band The Needful Dead. It's no longer in my chat history though so clearly Sam retconned anything his models did that could be considered racist.

5h agoHN ↗

Reminds me of the very early LlMs. I also sometimes miss the sheer demented horror of early image models told to generate images of biology.

8h agoHN ↗

Codex and I were trying to infer why Jev gave a certain parameter a given score. So I asked codex to produce a list of like 30 plausible reasons Jev might've selected that and then presented Jev with the initial prompt, followed by "You scored this with XYZ. What was your reasoning for doing this?" And then allowed it to do a noul value for each of the reasons codex generated. Felt like those people that give their dogs the buttons to push.

7h agoHN ↗

Use one black box to debug another black box. Neat :)

7h agoHN ↗

Jev has taught me the same lesson three times over now.

When it first came out, I thought "this weekend, I'll do a little open-source Jev based on single-token prediction and the token logit output", but of course when it came to it, there were at least 5 that had already been done between me thinking that and getting around to it.

So I wrote up[0] what other people had done, but wasn't happy with how weak the benchmarks were, but in the time between writing the first word and the last few, two excellent sets of benchmarks had been written, so I was able to incorporate those. I published the article, and one of the authors of one of the implementations commented that I'd beaten him to doing the write-up he'd wanted to.

This morning I thought "huh, you could have some fun giving Jev a single letter or token at a time, turning it into a chatbot", but as the time of looking two people had already done this (and taken the gag further than I would have), and ... this is isn't either of the ones I'd found. I bet if you scratch the surface there already at leat 5.

Time from idea to output has dropped off a fucking cliff.

0: https://sgnt.ai/p/jev/

7h agoHN ↗

Time from pointless idea to bad output, anyways. We aren't seeing good software, and now neat hobby project ideas are getting harvested pointlessly when the only purpose of those ideas was the fun and learning of doing.

7h agoHN ↗

LLMs are now like major highways, and everyone thinks theyll solve software jams by just adding one more lane; but that just induces demand, and doesnt increase efficiency because the traffic jam is about how people evaluate usage and fill the voids.

Similar to how we upgraded computers for decades and the software bloated to fill the specs

1h agoHN ↗

Also known as Jevons Paradox and Wirth's law.

7h agoHN ↗

Good ideas are the survivors of lots of bad ideas.

Slack is glorified IRC yet they're worth billions.

Dropbox can be trivially implemented via rsync yet they're worth billions.

5h agoHN ↗

There is a canonical hacker news post about Dropbox being useless and a pointless idea at one point :)

4h agoHN ↗

I always have a Mandela-effect moment about the Dropbox and iPod dismissive comments with Hackernews and Slashdot respectively.

God I’ old

2h agoHN ↗

The post is specifically about how it can "trivially be implemented using rsync" which is why I chose that verbiage :-)

5h agoHN ↗

when the only purpose of those ideas

I quite enjoyed handwriting my article to be honest.

7h agoHN ↗

I had the same sort of thing going on ahah. But I was convinced some hoje must have already done it and I decided I’d research when I got home (I’m out today). Didn’t expect it to reach hacker news so soon, though!!

6h agoHN ↗

I'll do a little open-source Jev based on single-token prediction and the token logit output

How is this idea generally working out in comparison with Jev? I'm curious, from what I read so far it seems like Jev is still beating this kind of thing.

But it's curious because it's not entirely clear why, from an architecture point of view for all we know that's exactly what they're doing. So it must come down to the quality of those logits, ie., model size and training details.

It seems to me that what most of these single-token-prediction projects are missing is that Jev seems to be claiming they predict well-calibrated probabilities. This is an incredibly valuable thing that LLMs simply can't deliver unless they are trained specially for it.

5h agoHN ↗

How is this idea generally working out in comparison with Jev?

Jev clearly has _some_ secret sauce compared to doing the dumbest thing that could possibly work with Qwen. It's not clear how durable that advantage is against OpenAI wiring up Luna-5.6 and doing a minimum amount of tweaking, but I presume we'll know in a week or two.

5h agoHN ↗

Everything old is new again, huh. I remember people doing this back in the early llama days, restricting grammars to yes and no tokens or 0 and 1 and then classifying questions. Usually it was rather ass in terms of performance cause no model is tuned to reply that way and it was WAY out of distribution, and yet then it got turned into the main way to run multiple choice benchmarks, and then everyone benchmaxxed it. Doesn't the normal MMLU/Pro also just do the same thing, restrict the output to one token, top-k=1, and it has to be one of the choice letters?

I think the real difference Jev makes is the fast parallel decode, it just seems rather bizzare how that works.

5h agoHN ↗

Given Jev does System 1 thinking, this would be equivalent to your ADHD heavy friend.

5h agoHN ↗

The worst part about this is someone still using poetry for python development.

4h agoHN ↗

i think the next step is make Jev a emoji bot... The architecture and its limitations would work well in that regime imo better then human language.

3h agoHN ↗

I think it would work well as a multi-emoji compiler too. Some phrases pair well with N emojis, and you’d be able to produce arbitrarily many emojis per phrase

4h agoHN ↗

i am having a bad day and the examples really cheered me up

3h agoHN ↗

I saw a lot of ppl think about what jev could use under the hood and could someone explain why this can't just be an embedding model where we just embed all the input + decisions and give back the cosine (or whatever) similarities?

2h agoHN ↗

if you compare an embedding model to something like Jev which asks 100 questions and use the answers as the embedding you will be able to get move mileage out of the latter. especially because you don't need to train any classifiers for your task, you can work directly on the answers.

that said, I don't understand the hype. I have been doing what Jev does for 2 years now by just forcing json tokens onto an LLM. you can even get the LLM to think. and you can ensemble multiple LLMs.

I suppose the appeal of Jev is how cheap and fast it is, but then it's entirely unsuitable for anything but the most cursory extraction. using it to play games seems like a waste of time especially when most of those games will be played better by an algorithm written by an LLM (just give it the state and ask it to write a bot).

2h agoHN ↗

Embeddings just convert the tokens to a vector that represents the text in an abstract semantic space. JEV goes a step further and actually processes the instructions/meaning of those embeddings to produce output, just not the usual series-of-tokens output we expect from an LLM.

1h agoHN ↗

ha i just tried making a sort of ouija board for jev this morning. it struggled. i look forward to reading how this one works.