- 23comments
- 48comments
- 3comments
- 94comments
- 10comments
- 14comments
- 122comments
- 1comments
- 932comments
- 2comments
- —discuss
- 14comments
- 36comments
- 169comments
- 12comments
- 2comments
- 31comments
- 89comments
- 339comments
- 187comments
- 22comments
- 4comments
- 45comments
- 81comments
- 5comments
- 426comments
- 30comments
- 242comments
- 327comments
- 70comments
To demo this technology for Hacker News, we built HN.watch. It’s like HN, but with explainer videos instead of articles. We create them on-the-fly the first time someone clicks on a link.
While there are obvious visual drawbacks of using HTML instead of diffusion models, there are three big benefits: - Speed: Much faster to generate than pixel-based videos (just a few seconds from click to playback) - Cost: Our cost per video is ~$0.04. (Excluding image generation, which some videos utilize. Quickly blows up the cost) - Easy editing: the above benefits also make AI-assisted editing cheap & fast
Our hypothesis is that if video creation goes from “dollars and minutes” to “cents and seconds”, a bunch of new use cases will be unlocked. Here are some we see already: - A video explanation of every single Pull Request (we do this internally) - Give every page in your internal/extrernal docs a video - Turn a complex article into a video in ~4 seconds (via our Chrome extension) - Course creators can quickly draft lessons before recording the real thing - People also create a lot of personal stuff stories for their kids, wedding invitations, birthdays, etc
The stack is based on an open-source programming language (Imba) created by our CTO, Sindre Aarsæther. It compiles to JavaScript, so it interoperates fully with the npm + node ecosystem. You can learn more here: https://imba.io/
We’ve also built our own sync engine (OP), and a context management system for agents (Q). We feared this would make the LLMs struggle when writing code for us, as neither is in their training data (there’s very little Imba in there too). However, we’ve been pleasantly surprised to see that LLMs actually are really good at our stack. This is probably because the stack is extremely dense. Imba is compact, and so is OP, where a single declaration sets storage, sync, permissions, UI, and what the AI sees. This means there’s no translations between frontend, API, db and JSON where the model can get confused and get things wrong.
Simply said, instead of using React.js, Express, Supabase, and LangChain, we built it all from scratch. Definitely suffering from the “not invented here” syndrome, lol! As for the models, we use Gemini, GPTs, Inworld, ElevenLabs, and a few others.
If you want to try it out, just take your pick: - The Web UI (scrimba.com/explain) - MCP (add it to your coding agent) - ChatGPT Plugin - Chrome Extension
You can find a link to all of the above in our docs: https://docs.scrimba.com/explain/introduction
And finally, a real pixel-based video of the tool: https://www.youtube.com/watch?v=k6rbHmBxSEs
Would love to hear your feedback and if anyone has ideas for other use cases.
PS: I expect quite a bit of pushback from HN for this launch, given how fan of text the HN crowd is. This kind of tool is not for everyone. But there are a lot of people today who prefer videos over text, especially in the younger generations.
Thanks, I hate it.
Kidding, I already sent it to a friend with ADHD who has been struggling to remain anchored to the tech world in any way besides Shorts.
But, culturally, I definitely hate the trend it implies!
Neat. Thank you for sharing.
loved scrimba. thanks for building that. was anout to pushback bc of the torrent of show hn sloppy pists but this one is fun.
yo this is wild
It's actually very good...
Please, nobody click on the link to this hn item within hn.watch as it would be even more dangerous than typing 'google' into Google
I'm an instructional designer, and I think this is awesome!
I'm going to be looking to reproduce this.
These AI explainers are taking over. I tried getting this working back in May. The models were okay but not quite there and it was a ton of effort. Opus 5.5 seems like the tipping point.
I made an OSS framework for these for when you want to go beyond one-shoting it: https://github.com/scosman/videowright
- Voiceovers: aligns animations to the voiceover, can generate voiceover with elevenlabs, or will transcribe and timestamp a real voiceover
- can reorder scenes both in code, and using ffmpeg for audio.
- interactive controls during authoring, can ask for micro edits or re-builds
- MP4 export/encoder
- Generates a video from a prompt (obvs)
This is great - many of us who are technical but not working in hardcore tech probably scroll by quickly on stuff we have never heard of. Getting a quick synopsis like this may drive more traffic and understanding.
Thanks!
This is actually quite impressive from an engineer viewpoint. I just have the feeling that the videos quickly become very monotonous and rather boring due to the monotonous AI voices. If somehow you could bring dynamic variation in these videos that would be fantastic.
This is awesome!
I like it! I bookmarked it. But I wouldn't pay for it.
This feels like when a decent book is made into a movie, and then the movie does well so someone will take the movie and summarize the plot in a YouTube video. TLDR: this is TLDR for TLDR. Just read more from the source.
This is pretty cool, I'd want to see the visualizations a little different but that's a personal preference.
Because we all contain multitudes, I can simultaneously accept that:
1. I hate everything about this because I vastly prefer text over video for the same content, especially AI generated video
2. There are a lot of people for whom video is their preferred medium and so this will be valuable to them.
It's a technically cool project and your cost-per-video is impressively low. Best of luck!