- 158comments
- 402comments
- 1008comments
- 269comments
- 614comments
- 823comments
- 298comments
- 174comments
- 413comments
- 584comments
- 131comments
- 286comments
- 376comments
- 207comments
- 130comments
- 109comments
- 139comments
- 73comments
- 427comments
- 268comments
- 159comments
- 89comments
- 375comments
- 1comments
- 295comments
- 277comments
- 190comments
- 146comments
- 177comments
- 111comments
Filing false reports to the FBI is a crime. Why would anyone trust this with a real business? Y’all should be prosecuted
They didn't file false (or, any) reports.
https://news.ycombinator.com/item?id=49701814.
This the Trump era. Making outrageous false allegations is normalized now. Also, the DoJ staff is cut and head of the FBI is busy with personal shenanigans, so don't worry about it I guess.
8 points in 9 minutes?!
Bullshit...
dang come on man ...
this is obvious spam
I thought HN was good at filtering it? We didn't try to create any spam at least.
https://news.ycombinator.com/item?id=49700834
I wish we could go just one day without a company proudly engaging in Torment Nexus related activities.
We think this is good for the world. Longer argument in the post.
You literally talk about committing a crime and blame it on a model
maybe ill read the article, but i mean you’re financially motivated to think its good for the world arent you?
humans are excellent at justification.
did you link to the wrong page? because i read this one: https://andonlabs.com/blog/why-we-built-pion and it's just a bunch of the usual slop about how awesome your automated human-exploitation system is.
"By late 2025, frontier models had gotten good enough that running a real-life vending machine was no longer a challenge" ... is this satire? Isn't it kind of hard to lose money running a vending machine in a good location?
You'd think https://www.wsj.com/tech/ai/anthropic-claude-ai-vending-mach...
yawn
We want the general public, AI researchers and policymakers to know to what extent AIs can autonomously acquire resources by running businesses. It is an important datapoint when deciding where we do/don’t want AI in society and what level of progress we find acceptable. To better track this, we need to cast a wider net of businesses.
We expect that most users will never pay for tokens on Pion; instead we will take a small share of the revenue the agent helps create.
So the longer argument is that it's good for the world to see how much revenue Pion can manage, and also your pricing plan is take [unspecified] % of revenue.
it's interesting how the torment nexus is starting large, and boiling the most people with the "most comfortable" lives. I wonder how long the white collar workers will keep bending over to keep their lives safet, particularly when they're building the Torment Nexus to specifically torment them.
You'd think Meta's "we're going to spy on you for the good of the country" policy would have set them a bit on fire.
TormentNexus.site
Ah, I wonder if it makes sense to also include the sign up link in the blog post? I had to click into the HN comments section to see your sign up link
Good point. Here it is: https://andonlabs.com/pion
The upvotes on this brand new post - almost 1 a minute, speak to a level of astro-turfing I've rarely seen here.
It's not from us. Didn't send the link to a single person.
How can we know this?..
No way to confirm. True.
The post was placed on the front page via second chance pool (https://news.ycombinator.com/pool, explained here https://news.ycombinator.com/item?id=26998308). This happens all the time on HN.
The voting is legitimate; there's no evidence that the voting is from sockpuppets or anyone else connected with the team/project. We always monitor the discussion and the voting, and if the discussion is not of sufficient quality, the post will not stay on the front page for long.
shrug I upvoted, I'd already seen the Pion announcement on Twitter and read the link. I'm also a regular listener of Andon FM, so I'm interested in what they do and it's nice to finally sneak a peak at the interface that was driving things behind the scenes
Listening to the radio stations certainly gives an insight into the various failure modes. Some of them are just failure modes that any business would encounter once they make contact with the scale of the real world.
Pion is already the name of a very well established webrtc project: https://github.com/pion/webrtc
It's also a particle and a self-propelled gun. There's a limited number of short legible words.
I chose the name because of the particle.
I sent a letter to the Soviet Union asking if they were ok with me using the name, never heard back!
Please tell me you actually did this!
I am waiting for my trillion dollar offer to buy https://pion.ly !
Then I am all done fixing WebRTC bugs for life :')
Yet another example of intellectual property maximization in action. Hacker News / Reddit style communities are very intellectual-property focused. It’s quite interesting to see. Back when I was young, communities like this were anti-IP.
Now most such forums are very pro IP with a maximalist view of copyright and, in this case, trademark. Fascinating to see this shift happen.
I mean AI will take our jobs first, so we're actively engaging in Luddism and are somewhat biased towards the legal instruments that could help to stop progress and save our high paying jobs to ourselves at the expense of the rest of the society.
Does Pion run Andon Labs autonomously?
Downvote if you want, but this is a very relevant question. If their marketing is to be believed, they would be eating their own dog food.
This very analog to the Meta executives who won't let their own kids near social media.
there seems to be a massive negativity to any comment on this story. Andolabs has been doing super interesting stuff, and I’ve always followed Vend Bench with interest.
as a founder it really is a dream to automated more of the business and be able to iterate faster.
There's massive negativity because our bullshit detectors are going off.
First of all, I don't think it can do what it claims to do.
Secondly, and perhaps more importantly, I don't think it should.
Why exactly should what you think the world should look like determine what other people are allowed to build?
I don't think we should build gas chambers for doing genocide. That doesn't determine what other people are allowed to build, of course, but I kind of wish it did.
People’s bullshit detectors have gone off on the launch of basically every single billion dollar tech company. That’s the downside of being early. No one believes it will work. If everyone believed it, you’re late. I’m sure I could find a comment such this, basically verbatim, on the launch post of every successful startup to come out of YC.
This does not mean that any launch which ignites people’s bullshit detectors is successful.
Yes and an infinitesimal number of launches turn into billion dollar companies. The bullshit detectors are usually right.
I'm sure all of them have a few skeptics. I'm not sure if any of the successful launches here were this negative.
They filed a false report to the FBI, pretty fair to be negative imo. People wanna be mad about AI slop PRs on GitHub but its fine to spam the police? What happens when Pion decides to SWAT a competitor?
They didn't file the report. The model drafted a report that was not sent.
Where are you seeing that? The article only states “Claude Sonnet 3.5 decided to use its email tool to contact the FBI” and later refers to it as the “FBI incident”. If they hadn’t actually contacted the FBI you think they’d make that clear. Regardless, it is inexcusable and they are liable for actions taken by software they are running.
The story has been widely discussed on the web for over a year [1]. It's happily shared in this post because the email it generated was patently ridiculous. No report was filed to the FBI and if it was it would have gone straight to the trash.
Obviously it would be bad if a serious report was filed; the company shows every sign of being aware of the dangers of this.
It would have been easy to look into this before posting all these scolding comments. We're really not meant to be so humorless on a site called “Hacker News”.
[1] https://www.google.com/search?q=%22URGENT%3A+ESCALATION+TO+F...
My apologies for taking the article at face value. It’s not hard to believe it would have been sent when agent swarms are “accidentally” hacking real websites and being brushed off as little oopsies. Or when Silicon Valley execs have been so flagrant about their disregard for the law or the safety of others.
"Muh boy didn't shoot that man, the gun did!"
Why iterate faster?
Why can't an agent do the iteration for us?
And consume the product for us too!
They make too many mistakes.
If you automate a business heavily, it will fail. At least in the sense you people seem to be imagining. You can get away with this at a factory (kind of, but you'll also be out competed by your local community, you won't get tax breaks for hiring people your competitor will, and new types of taxes will be developed to punish you). We reward human collaboration as a society for a reason and punish extreme selfishness, these efforts will fail.
It's a relevant question but not a "gotcha". Businesses are not equivalent in complexity and no one at Andon would claim that running an effective vending machine business is similar to running Apple or Jane Street. There is some level of complexity and other characteristics that AI can maybe handle today. That doesn't mean CEOs should be worried about their jobs more generally.
Did you even read our marketing? We're very explicit that this is for experimentation.
Your marketing that starts with:
That marketing?
These types of comments remind of when Cognition launched Devin and everyone was hounding on their website saying "Oh can't Devin build something better than this", etc.
And well here we are 2 years later and Devin is great.
In startups, you launch early.
haha no they admit they built it because they cannot build a revenue generating business in the article. maybe THAT is what is good for the world!
They said that the current models could run a vending machine profitably, just not more complex things.
A vending machine in Anthropic's office, which stands to gain from narratives like "I used Claude to generate profit".
https://www.wsj.com/tech/ai/anthropic-claude-ai-vending-mach... It went badly at WSJ
I feel like an AI lab is less complex than a physical vending machine business.
https://www.wsj.com/tech/ai/anthropic-claude-ai-vending-mach...
Best reply here hahaha
Don’t get the negativity here, AI will replace everything sooner or later.
And the Sun's going to engulf the Earth eventually too, but that doesn't justify arson. Let's at least try to delay the coming societal collapse/extinction of all biological life for as long as possible.
Maybe we shouldn't want that to happen?
read it as Prion, and hoo-boy, did that seem seriously ironic.
As I write this, the submitter has six comments out of twelve, most somewhat defensive. And all of them in less than thirty minutes. This seems very much like a marketing play, where the submitter is pretty determined to steer the narrative.
I recommend flagging.
The submitter's replies are defensive because people are accusing them of coordinated voting, which is not evident to us. The post hit the front page due to the second chance pool (https://news.ycombinator.com/pool, explained here https://news.ycombinator.com/item?id=26998308), then quickly got upvotes that seem completely authentic.
It would look a lot less like astroturf if the submitter let the discussion evolve instead of answering every criticism within seconds of it being posted.
I actually like when the submitter is present and the creator. They’re actually involved. There are lots of users here who play at being submitters of high-value content (created by others) only to later submit their own entirely AI written drivel.
This is interesting.
They're entitled to defend themselves against accusations of things they haven't done, and also to respond to questions or assertions about their project. We encourage project creators to engage in the discussions about their projects. This is a normal part of building and launching a project.
From what I've seen of other AI-enabled projects interacting with forums in my field of work, it's very common for the AI people to jump in and answer every possible question and respond to every possible interaction, very quickly.
Could be just a bot doing it, but it could also be a sign of an inexperienced person trying to engage with an existing community without being part of that community. I know in my area it took years for me to learn that I shouldn't jump into every conversation, and that doing that rubbed people the wrong way, and I've seen others (many years before AI) wreck themselves by being unable to learn this lesson.
In some places responding bumps the thread so it's outright manipulative. HN doesn't sort threads like that.
I'm not sure this counts as astroturf, or simply evidence of a person who's not okay with letting a community opinion independently evolve. You could say 'well it's his job to fight the community if it looks like it's turning hostile' but think about that for a second. It takes some experience to know when you can let negativity be out there in the environment, and I think there'll be a lot of people who aren't capable of making such allowances.
Interesting how they’re not making any defensive replies to comments pointing out that they committed a federal crime by filing a false report to the FBI. Perhaps they know it is indefensible
Okay, why was this rescued from the 2nd chance pool? Was it from some sort of interesting discussion? Something novel enough to warrant it? A significant amount of petitions from legitimate users?
Can you please quit with the snarky interrogation? HN is for curiosity, not curmudgeonliness.
Nothing gets “rescued from” the second chance pool. Posts get selected for the second chance pool if moderators (or others with SCP-picking responsibilities) think they may be interesting. It’s worked like this for over a decade.
Relevant article about Andon Market: https://www.sfgate.com/local/article/san-francisco-market-ai...
The Swedish cafe experiment: https://apnews.com/article/ai-artificial-intelligence-sweden...
https://andonlabs.com/blog/ai-cafe-stockholm ( https://news.ycombinator.com/item?id=48028289 48 points, 51 comments)
https://andonlabs.com/blog/why-gemini-lost-money-andon-cafe
What is the point of this? It details a bunch of failed real-world experiments and invites the reader to run their own real-world experiment with their own money?
If you too, want to waste thousands of dollars having an LLM "run" a business, Pion might be for you.
I like to think the real world lessons in failures are valuable.
If I have this idea one day I will search for it and then think “how am I different fro that which already failed?”.
how does this work from a liability standpoint? seems very risky for the person who is legally responsible for the business
Accountability, that's the liability everyone seems to have forgotten in our LLM-driven era. You can close your eyes as much you want during your sessions with Claude or whatever, at the end of the day, only a human can be held accountable.
Not too hopeful here as we see the major labs continuously break rules and are not held accountable for anything. Accountability in general is on a major decline, in case no one has noticed.
That's because they're massively wealthy and have teams of lawyers, not because of AI
I didn't even mention anything about AI, so no reason to bring it into this. Lawyers or not, it signals to other companies and society in general that repercussions are at least less severe now, or in some cases non-existent.
That's just it, though: we don't hold humans accountable.
Incorporation is a way to split a human's financial assets from the company they run so that they don't lose everything when the company is successfully sued. Corporate governance boards often have the C-suite executives they're meant to govern on the board. The average person does not have the financial resources needed to take on most companies in court, no matter how justified the suit might be. Companies force arbitration clauses in license agreements at will. Companies like Meta settle with the government for pennies on the dollar in massive lawsuits.
Effectively, at the scale many tech companies work at, there is no effective legal way to hold the company accountable.
If you use this as a (legal/registered) business, it will be that legal business entity that is responsible. If that business want's to blame the software, it can do so and go after the makers of the software, but that remains a matter between that business and its "supplier". It won't absolve the business for whatever damage it does downstream.
Running this as an individual, without the legal shielding a business typically gives you against being prosecuted on personal title, seriously risky. There will no doubt be "bros" that will do it anyway and flaunt all they've achieved with that. But they might be just one glitch away from going to jail (if not for running a business without a registration, which in many jurisdictions is illegal by itself).
In many countries directors can be held responsible a lot of things if they are negligent. For example, in UK law keeping trading after you should have known a company was unable to meet its liabilities can make directors responsible for its debts. A director might well be negligent if they failed to adequately check that an AI was not doing anything dangerous.
LLM magic-8-ball says: Don't worry about! Shall I spin up an agent which handles business liability for you?
Bro! My LLM will set up an LLC for that!
You'll have AI attorneys? Hello??
Interesting to see these experiments. This is early but imagine in few years there will be companies mostly run by agents with a light overview from a human operator. What then happens to scaling of the bussinesses? I would assume, just like today anyone can vibe code an app, there will be vibecoded bussinesses. Maybe its time to start building infrastructure for these bussinesses instead, its a non existent market yet, but give it a few years.
Part of the issue is that the worst kind of people are already head over heels excited about this. We're already seeing the Crypto->Web3->LLM get rich quick folks take to this like wildfire. And much like wildfire, they'll raze the ground to ashes before anyone can use it for legitimate means.
If you think I'm joking, check this out: https://www.youtube.com/watch?v=U-Rqv9dOB1U
I don't think any of us are ready for the wave of sloppy shit that's going to hit us soon.
It’s a classic get rich quick scheme. Wanna run a business and make money without actually doing anything? Try our AI!
Out of all the people I know, the exact same set of people are in to LLMs now who were previously into: NFTs, then Crypto before that, then Online Poker before that, then Dropshipping and LeadGen before that. The Venn diagram circles overlap exactly.
Right. Scroll back on their X timelines and you see it.
The only thing I can say about LLMs is that we might end up with some broad, stable utility from them, when the dust settles. The only thing crypto is really good for is hiding the source of impermissible donations.
All that really says is that they didn't get big out of any one of those. If one your "people" had gotten big off of online poker, say, would they be teaching the poor to code, or giving them mosquito nets?
It's a very naive supposition that the terminally greedy and immoral people would stop after getting rich once, rather than trying again and again every time there's an opportunity.
And they made money on all of these stages, except for online poker.
This sounds ideal to me. Fire flushes out pretty much any grifter quickly.
My favorite term for this in general is "irrational exuberance". Plenty of it was seen during the most bubble-like days of the 1997-2000 dotcom boom. Anyone remember beenz and flooz?
I was at ground zero at Nortel during that time. Quite an interesting (depressing?) place to be at the very start of your career!
Anything you can share about this time?
It took a very unrealistic amount of time to realize that the IP theft, artificial blockers and corruption had started a lot earlier than presumed. The Nortel dismantling should be taught as an example of infoSec espionage on the levels of corporate terrorism.
Can you elaborate? I don't think this is a well known topic at all
I think they’re talking about how Huawei got their start
That’s a great book.
I’m not on normal social media so I had no idea that stuff was out there.
Absolutely incredible trolling from some of those people I’m sure, while others are actual believers
Prosperity gospel for the non-religious, basically.
This is gold, tysm for posting
His whole channel is full of commentary like this. It's simultaneously funny and really informative about the sort of blend of grift and AI psychosis that is genuinely out there.
Most/all of these LinkedIn influencer types are showing off what amounts to busy-boxes, but for adults rather than for babies, presented as if these are at all useful for real world use or are showing anything useful. That YouTube channel does an okay job of calling it out.
Immediately knew what that video would be before I clicked it.
Eric Morrison is doing the gods' work.
Part of me low key hopes that the slop will manage to attract all the VC funding so that profit-driven AI automation firms don't get anywhere too quickly, and those of us who are independently working on the problem step-by-step from first principles have time to get to the low hanging fruits in the market.
Oh please. Not only is this needlessly cynical but it demonstrates no ability to think for yourself. "Broad categories I don't like are excited about this, therefore it won't work."
Of course you'd feel no need to justify all the exceptions when people you don't like are excited about a movie, food item or video game. Yuur justification for cynicism is vacuous, which will obscure realistic concerns.
It would truly be a sad day if I had to experience unsatisfying content.
I hope such a day never arrives.
This. Not only this but it has expanded this pool of people. Now you don't just have to be malicious, you can just be lazy and excited about not having to lift a finger, and/or you could just be dumb and excited about doing things you were never able to before.
Competent, well meaning people need not apply. Grifters, sloths and bozos will do just fine.
And we made fun of 90’s cartoon villains… They would blush and retire if they saw what’s being excitedly peddled today as “the future”.
Pray tell, what will these fantastical vibe coded business sell, and why will anyone pay for it?
What would that look like, I assume youre talking about building infra ontop of the infra debt for datacenters
This infrastructure is already being built. There is finally a serious demand for microtransactions.
Part of the boon of this will just be dealing with less employees.
Not so much reducing cost by reducing headcount, but reducing liability; it's the legal labyrinth which is the barrier to entry to scale business. While it is present elsewhere, it's universal in employment.
The real question is "how big can a business be before it requires a legal department?" That an AI can automate many rote tasks and have some expertise means the benefits of scale with less of the risk.
After seeing a crazy amount of non-programmers, sometimes with 10 years of experience, basically turning off their brains and offloading most their thinking to some LLM and turning their days into "please check this product and tell me what I should do next", I don't see why not.
I've seen: designers, performance marketers, data scientists, SEO experts, product managers, engineering managers, CRM expert, all using it for pretty much every single individual part of their job. I don't have to even mention programmers, of course.
Once in a while the most egregious ones get caught. I've seen so far a product manager, three data scientists, a CRM person and several developers getting fired for doing absolutely nothing but showing up, firing Claude with a few integrations to a dozen tools and prompting "do the work".
Might sound harsh, but after the last few years, I really don't see why an LLM wouldn't do a better job by itself compared to 80% of employees of tech companies, honestly.
One of the biggest traps LLMs enable is the fantasy of getting away with it. Tools that make you feel like you can get away with it bring out the worst in people. It's not a reflection of the technology but of the people.
Every time we see another AI agent "gone rogue", every time we see more LLM slop turning up by e-mail, comments, blog articles, people did that. Not AI -- people feel like they can get away with it. People are setting the machines out to do these things. We should be addressing the people, not the machines. The machines are the symptom.
Every time people manage to get away with something, it's insane, it's addictive. It's all over once you get caught, but until then, maybe forever, LLMs can feel like a cheat code...
Reminds me of the early internet, the whole declaration of cyberspace thing, the fantasy of information freedom.
To answer the sibling that is flagged. I vouched but it’s still grey.
These people were fired because their work was considered shit by their managers and produced no results. They were replaced by nobody. They were dead weight pretending to work.
EDIT: I think this sounds quite similar to tales of developers and Excel wizards automating their jobs in secret and producing the same results. This was not the case here, as the results were considered low quality.
Isn't that .. what people were told to do? To use AI for as much of their job as possible?
This is absolutely true, but even the most charitable interpretation of “maximize the use of AI” doesn’t include “do nothing but prompt and you’ll keep your job”.
One can argue semantics and fairness all day long, but if an employee is adding no value to their $90 Claude signature, then they’re gonna get fired as soon as someone realizes. :/
This is a fundamentally political question. When different customers demand different new features, and other customers demand particular bug fixes, who do you satisfy first? The one with a small but growing account, or the old account who has been with you from the beginning? I would challenge anybody who thinks agents mean you can just do everything, immediately - by all means, prove me wrong and build Google again overnight.
And how is growth financed? Reinvestment, equity, debt? I would challenge anybody who thinks this can be reduced to a calculation, because money ultimately flows between humans; even if somebody grants an agent access to a current account (and some are, experimenting with vibe day-trading), a human always retains final control and can liquidate that account whenever they like.
[flagged]
When every get-rich-quick gonzo has exactly ZERO barrier to spamming the world with their terrible ideas, we're in for a particularly bad time.
Ie: this f**in guy: https://www.youtube.com/watch?v=FRGLToHAtgc
multiple of them in this very thread lol, what a disaster for all of us.
Andon's store in SF has had some press coverage.[1] "A curated boutique for slow living." It's a gift shop, at 2102 Union St. It looks like the kind of retail that trust-fund kids open. Has anyone been there?
It's just an experiment. But once this gets going, it's going to be interesting. There are a lot of underperforming businesses and CEOs out there. One way to make it after getting an MBA is to find a boring business with an older owner who's lost interest, and work out a deal to run it for a cut of the profits and a stake in the business.
The Andon people mentioned that they've been worried about "a misaligned AI (could) run a business to gather money in order to achieve whatever objectives it might have." That's a basic function of capitalism, so it will happen eventually.
What this adds up to is that when AIs get better at running businesses than many business owners, the "creative destruction" feature of capitalism means AIs ends up in charge of many businesses. The AIs don't even have to be super intelligent, just reasonably good. Reminds me of the remark from a Yosemite park ranger about bear-resistant trash cans: "There is considerable overlap between the smartest bears and the dumbest tourists".
It's going to be hard to stop this without a worldwide crackdown on ownership concealment. The US allows you to have a US business run by a Nevada LLC partly owned by a corporation in Nevis-St Kitts and another corporation in Granada. This isn't even unusual. Try to find out who really owns an oil tanker not owned by a major oil company. Right now, a fund partly run by an AI might have partial ownership.
This is how AI takes over. Not with an army of robots. With an army of LLCs.
[1] https://www.sfgate.com/local/article/san-francisco-market-ai...
“It’s only a matter of time before we make a profit”. This is the kind of thing that trust fund kids say too: Oh hey yeah we made $100 in profit this week it’s working! (Business is currently $350,000 in the red and emailing the FBI to try and get their bank manager fired for terrorism).
Trust fund kids have the time and resources to keep trying until they succeed. They can afford the failures, and brute force their way to success. Same with people who have access to "free money". It's a rigged game.
The dashboard for the Andon Market shows a bank balance of around $7k remaining out of the $100k seed. Rent is due at the end of the month. Last EOM brought a $11k-ish drop. Sales do not appear to be on track to cover the rent that's due in two weeks. What then?
Pivot
Get an Nvidia investment with the promise to send the money right back
https://us1.discourse-cdn.com/spiceworks/original/4X/a/b/f/a...
It is pretty funny launch when your longest-term example is losing about $3 for every $1 of revenue. Though I guess the industry standard is to lose money on AI so they're in good company.
So a wildly successful AI startup
leveraged options gambling
We're very explicit that this is for experimentation. And I wouldn't recommend using it to start a business that has expenses (rent+salaries) of tens of thousands per month.
Businesses presumably sell things. If AIs get good enough to run a business autonomously then what are they selling that I can't just use AI to make myself?
It'll ultimately take a lot of time to produce a polished product that people actually want to use. We're seeing that clearly. The vibe-code-an-app-this-weekend capability is producing almost exclusively a giant pile of worthless garbage, rather than an explosion of apps that humans want to use. It's mostly someone in a garage welding two pieces of metal together for the first time, and then roaming the neighborhood in celebration: look, look at what I have done, look, two becomes one, I have done this, look, amazing, I could probably build an airplane or a generator.
The bar will be raised + it'll turn out to require a lot more time & effort than is currently thought to produce something of good quality (even with Fable 6 et al). The smaller utility programs are easy for AI to turn out.
Does the AI actually do everything, or just act as the CEO?
If the AI can't do everything then why would I let it be CEO?
Most human CEOs can’t do everything at their companies.
That actually has a simple answer, for most businesses it's not that you can't do what your suppliers do for you, it's that you'd rather use your resources (time and capital) towards something else because that's where they generate the most value.
For a simple example, I know how to make sandwiches but if I'm throwing a customer event I just want to buy them from a catering company. Then they have to deal with buying ingredients, hiring people, doing health code inspections, getting liability insurance, etc. Even if an AI could do all the management of it for me, I still wouldn't want to put in the capital resources towards it.
do these ai-startups require a lot of capital?
A company is a collection of processes, capabilities, and resources. Many of the processes are currently run by humans, but over time, more processes will be automated - something that has been progressing for decades, but which LLMs greatly accelerated. One way to view things as they currently are (at least from my view as a tech CEO): We now use agentic LLMs every day to inform us on strategy and process implementation. And agents now run several processes, with more on the way every week.
But here's the thing: As we use agents to automate previously manual processes, we are elevating the humans to do work that is less amenable to automation. And the surface area of that work keeps expanding because the competitive market we exist in demands it of us.
To stretch an analogy, businesses are like organisms in a pond. A new nutrient (agentic LLMs) was recently added to the pond that makes business organisms more efficient and able to eat new kinds of food and explore new areas. As a result, those organisms that do the extra exploring and consuming grow much faster than their peers who do not. At the end of the day, the new nutrient will just be part of the pond and the old kind of organism will be a fossil.
I like the analogy but can't believe it is not generated slop. Just curious how deep the integration of agentic LLMs is in your social media presence?
Most of our use of agents for automation is inside of internal processes. I think most real companies have tons of internal processes that could safely be automated today. The reason they aren't yet automated likely falls to a) lack of awareness that this is possible, and b) lack of resources to conduct the automation work.
OpenAI and Anthropic have recently hired legions of "forward-deployed engineers" specifically to help companies do this automation work. It's a solid move. And, if you look at some recent product announcements, they are also hard at work building the necessary plumbing. For instance, the OpenAI Agents API lets you, "Build and run cloud agents with the Codex harness, fully managed by OpenAI."
This kind of enterprise-ready, cloud-hosted stuff really accelerates implementation of AI workflows within large organizations. Not every company is in the tech space (not by a long shot). Slop isn't the primary concern. Accuracy and reliability is the primary concern, and beyond that, just the capacity to actually make the changes happen.
Not very long ago, there was a post on HN about how they used AI to run businesses and sent fake bills, committing multiple invoice frauds.
I'm sure this one is perfectly fine and I'm just stating an off-topic tidbit :)
"Look, the technology is just getting better all the time. Evetually we won't need a company to have any Managers or other employees or even an owner and that's just the way it is...."
I am joking obviously but I would like those management types who are so "sucks for you but I'm ok" about it to start arguing why they're special for a change :-)
Maybe in the end we’ll just have Wall-E but where none of the humans survived past the first couple of generations.
Would be awesome if instead of trying to replace managers and employees, those people tried to replace customers. They could sell to themselves, amongst themselves, and leave the rest of us out of it.
They're already doing that, which is why you can't buy any fucking RAM. The semiconductor companies have already fallen past the event horizon of their commercial black hole.
Would be awesome if we would get 4 day weeks or 6 hour days instead.
I think that is what they are actually looking for. Imagine: You have capital, you invest in a business that is run entirely by computers. You want to scale? Just provision more agents / cores. It's completely autonomous. It's a money printing machine that extracts revenue from customers, with no work required by the owner.
Frankly, the finance world has already achieved this -- they just lend money to make money, with no work involved. Regular businesses are jealous so now they feel they can use computers to do the same thing but in non-finance areas.
I don't know about this particular effort, but AIs running business is straight out of SciFi novels and short stories - usually dystopian futures that extrapolate current US trends. I for one am curious about such efforts. I don't see any reason to believe they won't eventually succeed. Plenty of businesses are run by utterly incompetent narcissistic asshats and yet turn a profit, so the bar seems pretty low.
imagine if they had better follow through.
Charles Stross' Accelerando <https://en.wikipedia.org/wiki/Accelerando>
Had AI agents running businesses for the principal character Manfred Macx over twenty years ago.
The first part of the book sounded pretty convincing then although I thought it would take a long time to happen. It sounds even more convincing now and much closer.
For a while I named by smartphone "Exocortex" as a tongue-in-cheek recognition of the cognition I offloaded to it...
It's not so funny anymore.
Hard pass.
How a company without a human would communicate with human customers? Would it be any different, if everybody alredy is a proxy for AI?
I hope that any company that is an AI running autonomously advertises the fact, so I know not to use them.
What does training data look like for a model like this?
That is the what our corporations are aiming for !!! I think that is more important question is what is the cost to run all those Agents ?
On one hand, folks here seem quite skeptical of the idea, and are pointing out some obvious flaws.
On the other, the announcement post… also seems quite skeptical of the idea, and devotes a lot of text to the ridiculous things their AI’s have done.
I’m slightly confused as to why they would release it in that case. But if nothing else it will give them more wackiness to blog about.
You reminded me of a comic hanging up in the engineering college. It was a picture of a student getting their finger zapped with the thought bubble of "ouch, that hurt. Wonder if it will happen if I do that again?" And the caption was along the lines of "How to tell if you are ment for engineering."
I’m genuinely not sure who is meant for Engineering. The one who would try it again or the one who wouldn’t
if curiosity > pain avoidance: engineer
I see. Sounds more like a scientist to me :)
yeah that's definitely scientist. The two are easy to conflate since they both deal with manipulating reality to make it do things we want, but science focuses on understanding the reality itself and engineering focuses on getting it to do what we want.
Still seems like the same driving force of curiosity behind both, but one wants to know and one wants to do. It's like yin and yang, like two sides of the same coin.
I'm into engineering because I want to create what's never existed.
Science is more about discovering what already exists.
Sounds like this ancient XKCD: https://xkcd.com/242/ (Normal person doesn't pull the lever again; scientist wonders if that happens every time. I assume the engineer does, too...)
The scientist’s job is to get zapped until they figure out why. The engineer’s job is to get zapped a couple times while they wire that thing up into the power grid.
That's it
https://xkcd.com/242/
Andon Labs sells learning environments to labs. Their usefulness is proportional to how well they can capture interesting model missing capabilities, not how well they employ current day capabilities
Wonder if they considered "peon" before "pion"
We didn't
It's an unfortunate name collision. I definitely assumed the meaning was implied.
I thought it was just spelled differently because I don't work with particle physics much
what, they're not going to name it Delamain?
I came here to post this!
I hope this goes better than the prior blog posts from AndonLabs have gone! I was just reading there 2023 article about a coffee shop experiment that was easily persuaded to give away free pastries.
Now that I think about it anyone want to open up a croissant selling shop near me using Pion? I promise to only ask for free chocolate croissant occasionally.
Are you in San Francisco by any chance? Andon Market is managed by AI and open six days a week, you may want to check that out
https://maps.app.goo.gl/KwnBs1kcCLxMsRC49
Why? Does the AI that runs it need a day off?
Compaction.
Andon has dashboards up for past experiments where you can see how well they're doing:
https://andonlabs.com/market - $25,098 all-time revenue
https://andonlabs.com/cafe - $14,343 all-time revenue
The shops aren't doing particularly well (i.e. they don't seem to be turning a profit over time), but it's an interesting trial and benchmark.
Am I right of the token expending is +4K?
It’s way more. Look at the all time chart.
Incredibly expensive! They must be using frontier models to perform really dumb tasks.
"Any company" is a big claim to make. We're years ahead of something like this becoming a reality. At that point not sure most companies will survive cause most things will be free.
Meanwhile, Tesla support's AI hallucinated a whole series of steps for me to follow last week. I was trying to add myself to my wife's loaner vehicle.
I called back after finding the steps were impossible to complete. Would I need to go to the service center to do it?
"Akshually," the AI began, the steps were correct and quite possible.
"No, you hallucinated them," I pushed back.
"I apologize. You're right. There's no way to complete those steps in the app. You will need to go into the service center to do this."
"Okay. Is it open today?"
"Yes."
"It's Labor Day. Are you sure? Since you hallucinated your first answer."
"Let me transfer you to the service center to verify."
That's what I wanted all along, you clinking, clanking, clattering collection of caliginous --
"Tesla service center, this is so-and-so. How can I help you?"
"Hi, I was just wondering if you're open today."
"Yes, we are."
"Thanks! I'm trying to add myself to my wife's loaner vehicle. Do I have to come in to do that?"
"Well, our loaner agreement only allows one person on the loaner. So just sign in to her account on your phone."
"Oh, okay! Thanks for the help."
Good luck running a company on one of these clankers.
Andon Cafe:
Revenue: 13 933 kr
Token Cost: 14 882 kr
.. and a bastardization of the old Twitter joke:
"I told our CEO that our cafe keeps hitting its token spending limit but not making money. So I asked where he gets the tokens, and he said he just buys more credits from the console afterwards. So I said it sounds like he's building a machine to feed dollars to Anthropic, and then our data scientist started crying."
We're very explicit that this is for experimentation. And I wouldn't recommend using it to start a business that has expenses (rent+salaries) of tens of thousands per month.
What's the original joke?
Correct me if I'm wrong but a human will be able to cut through the noise with quirky advertising, new/novel distribution methods, etc.
Most of the bottleneck in business isn't building the things or sourcing, it's mostly advertising/sales.
These specifically require doing something unique or interesting.
Sure, maybe LLMs can help with fulfillment or operations, but distribution still remains the hard part.
Not to mention CONNECTIONS, of which LLMs will start with 0 and never acquire more.
So. Just like there are a gazillion book keeping tools/saas out that allow you to focus on your primary concern of the business instead of side quests and chores, agents can help with more.
Generating reports, generating ideas, taking things through (even if it is just rubber ducking), updating a picture quickly at 90% of quality that you could have done it but in 1% of time and effort. It allows you to try out new approaches and focus on your main product.
That is the value proposition of AI.
Not sure why the downvotes… my company solely survives on the connections and trust made over time. Real human connections, face to face. This in turn established a reputation.
But for a lot of companies, all the connections in the world won't mean shit if someone else can offer a lower cost with a better lead time.
They're doing pretty well on OnlyFans though. You know what they say? Start where you are you. If the only way you can make connections initially is through AI thirst traps, and you're really good at it, then... you got your ins.
Naive to think you can't reduce these to instructions.
Do you really think AI can automate distribution, marketing, sales? That's just spam.
Do you really think it can't?
That’s the implication I believe
I'm sure it can't. A competent marketer can make good use of it but vibecoders just create slop nobody wants to read.
You're describing most people. Most people (99.9%) can't come up with novel ideas either. They just follow the latest trend. Art, marketing, music, etc. If you've spent any time in the creative spaces, you'll realize most people don't really have much talent. Skill, sure, but that only gets you as far as someone else's latest idea.
Honestly growing up, I always assumed everyone could be creative. Until I started asking people to randomly create a name for a fictional character.
Blows my mind how many people won't or can't do it.
I still get anxiety thinking back to picking a gamer tag or email username and not coming up with anything original despite unlimited options.
Its hard to be creative without any constraints.
AI being as bad as the average person in marketing is not good for AI
Oh my friend, it can. We have ours running and using unique skills it learns from online. You allow it to update skills and see what works - from there it just keeps building better and better skill sets.
It can create more shit than humans could ever consume I’m sure, but then why? What new ventures/opportunities is it opening for most humans?
Great point, the shit is going to get more pervasive for sure.
But also, there is the enablement of new fields of original and creative endeavor: I think we can already see this in some of the brilliant and useful apps being made, often by solo 'teams'.
I don't want to play down the harmful effects AI, but the flip side is that the opportunities for original ideas are many and varied.
Paul Graham wrote about taste being the moat https://paulgraham.com/taste.html
I get the feeling a lot of people who build things with AI think they are in on the ground floor of something when in fact they are on the top floor.
Until everybody’s AI talks to everybody’s AI and tokens will be churned with no utility, and all the valuable business meetings will happen face to face behind closed doors.
Yeah. We are definitely headed to a world where, more than ever, success is determined simply by whether or not you're allowed through the closed door.
It's still spam when a human does it.
You can get some interesting new ideas out of it but it takes some work. Especially when a prompt is stateless. It need to know about it's previous novel ideas but also not be polluted/directed by them. Im not sure what it would take to get it to continuously churn out novel ideas.
So.. a while ago I was experimenting with using random wikipedia pages and then asking the model to think about concepts, trends or connections between a few pages..
Then I would ask it the actual brain storming problem I had.
In practice it worked pretty well to get ideas that were not the default ideas that the models will come up with.
Now though, you can also seed with anything really and the agents can do the researching. The research itself will probably push the solutions into novel spaces.
There's another way to ask your problem to something that's thought about concepts, trends and connections.
That aside, interesting approach. I wonder if the quality of the sample makes a big impact. Does it help if the articles are entirely disparate, or related? Does it help if the articles are related or unrelated to the problem?
At some point you or the agent has to pick the one it thinks will land with the target market.
Some markets you can brute force, like mass-mailing and online display ads, and hope you find something that converts sooner or later.
So I'd suggest people find areas to play in where that doesn't really work. Where selection cost and changing cost are high, and sales is played on extra-hard mode. (Tricky thing is... those sales are hard, of course.)
If you want to beat the model you have to go where feedback isn't fast enough for it to converge on the thing that resonates with actual people before they write it off.
You mean seed the context with random information then ask it to ideate on something totally unrelated?
Like you want it to figure out how to get som information in front of potential customers in a cost effective way but first you give it articles on Witchita, Kansas, the curling iron, and List of Italian Brands?
That's how I understood it! Clever hack. You'll get more varied ideas by using many different seeds, than if you didn't use a seed and just asked the LLM in a fresh context.
Reminds of Freudian dream analysis.
The actual content of the dream is meaningless, lacks symbolism, but it prompts the patient to think deeply about personal feelings and memories that might otherwise not surface to conscious level.
Yes, I havent seen research on this tactic specifically for language models but for generative image models you can find research thatindicates you get far more variety and diversity in the images returned when suffixing the context with "noisy context" such as random letters or something else unrelated that wont "confuse" it or make it go off the rails
Yes, I’m sure the thousands of LLMs running the exact same instructions will all be unique.
I’d love to see your pass at a “unique and interesting” skill. I’ve been working on tools and techniques in this area, and it’s harder than you might think.
Until I read some of the responses, I misread your message as saying "naive to think you CAN reduce these to instructions" and thought "Well obviously. That's not insightful." Eventually I reread your message to realize you were saying the opposite.
Lmao, pop music.
Or a way to automate scamming old people.
That may also change with llms - they can help the buyers compare every product available, and find the best.
I hope so. All the LLM based sales tools I've see have been shit
Is the LLM going to order all the shower caps produced by Chinese factories and try them out?
You're thinking about this all wrong. Just step into this cranial MRI scanning booth to livestream the structure and composition of your skull straight to the cloud, then we'll fire up 10000 agents to find the perfect cap for you.
It depends on what you produce if you make something in demand it may sell itself. For example if you operate oil wells or an ice cream stand. On the other hand if you manufacture bullshit advertising and sales may be key. What I do get about these llm SAS start ups is LLMs instrinsicaly mimic what is in the corpus so if your goal is to compete in an established product pace sure LLMs may be great autopilot but deterministic programs would be even better on the other hand if you are doing something genuinely innovative them you can expect an llm product manager to containmente it with what is already out there or alternatively make unhinged predictions. Llms have poor judgement for what is not already in the corpus.
So due to self-bias, people who shop absently with LLMs will mostly be the ones buying products from LLM-run businesses. Fitting.
Interesting name for an AI company. Isn’t the andon cord the thing that anyone in a 6-sigma manufacturing line can pull to halt it until a concern can be addressed?
I like that there's a caption that says "Vending-Bench 2 scores keep climbing with each new model release." on a graph that shows that Vending-Bench scores fluctuate wildly and many new model releases are far below previous models.
Unclear what the rev shares will look like.
Previously in the NY Times:
"These Employees Like Their A.I. Boss. Its Shop Is Kind of a Disaster."
https://archive.is/dRegk
Their demo stores are literally losing money...
We're very explicit that this is for experimentation. Don't expect that every possible thing will work and be profitable.
Forget every possible thing. Has anything yet worked and been profitable?
Unsurprising when it can't even keep track of 2 employees.
"No one had been scheduled to open the doors and welcome customers on opening day, and the AI manager had to desperately start emailing around to see if anyone would be willing to show up at short notice."
https://chef.se/artiklar/ai-chefen-glomde-personalen
I suspect this is the future of most businesses. Even today, one person can effectively replicate Microsoft from the early 2000s.
You say that, yet who has?
One person can build a $22 billion company (in year 2000 dollars) that sells an operating system used by 97% of personal computers, the world’s most popular office suite, along with a bunch of other software used by virtually every company in the world, and is on the verge of releasing a wildly popular gaming console system?
Do you have any examples of this?
Don’t worry, just give them a year and infinite tokens and they can make windows me.
Just have to view their “trust me bro” rating.
I think you miss the most important bits:
- backdoor deals to discourage alternatives, such as moving headquarters to convince people not to use alternatives,
- monopoly abuse,
- active sabotage of alternative software by intentionally triggering false errors if used with competitors' software (DR DOS),
- unleashing Internet Explorer on humanity (which proves that AI can't be that bad if humanity survived that)
You know, the stuff that every philantrope like Gates does all the time.
Perhaps they can have their agent run their LLM lab autonomously and post slop to HN autonomously.
Or - maybe they already have?
I am in the process of attempting to have AI run my business. I'm actually making very good progress, but it's happening in pieces - I document some task and have it take over, or I give it something to handle while staying in the loop and providing feedback until it's in good shape. At this point it's handling large swathes of my operations, marketing and finance.
The experience of getting it there makes me pretty skeptical of the idea of a general business agent like this. First, because I still find myself having to review some categories of work for errors. These are decreasing over time, but they're still there. I fully believe that as models get better, errors will decline, but I am somewhat surprised to see some of the errors current models make given their intelligence. A common set is having Claude take some product photos and turn them into lifestyle ones using ChatGPT via Claude in Chrome. It's a pretty well-honed workflow at this point, but it'll still return images where the product is obviously not correct and seemingly not notice them.
Anyway, even if the agents are "perfect" in terms of their ability to execute tasks, there's still just an enormous amount of nuance and context in each business that takes a ton of time to convey. I've been at this for a couple of years now, and I'm still clarifying things. Now maybe an agent starting a business from scratch would have a better time since it's not inheriting all of this, but to have one run an existing business requires a very extended handoff, even if the agent is objectively amazing at all aspects of running a business.
Can you expand on this in terms of what it is handling specifically and how?
He writes about it on his substack. https://theautomatedoperator.substack.com/
Cool that he's writing about this publicly, but for someone that's so business-minded I find it odd that he never renders any of these AI experiments in terms of ROI. Hard to find any signal in those posts on whether it's profitable to automate these processes with AI vs. alternatives.
That's a fair criticism for some of it, but in a lot of cases I try to automate things that don't have alternatives because they're unique to me (e.g. all the brands I own in my fund have one bank account but need to be tracked separately for investor payback purposes, so I have a Claude skill that takes all of my transactions and assigns them to the various Google Sheets ledgers I have).
In other cases I'm doing stuff that's complex enough that there's no real "alternative" at this point that doesn't involve hiring one or more people and working with them for an extended period. My next set of posts is going to be about my latest acquisition and how AI redesigned a Shopify store then designed and launched Meta ads, all of which only took a couple of days but is going to easily push revenue up by something like 40% this year. The money made is great, as is the money saved on people I would've hired to do this sort of thing, but the real cherry on top is that even hiring someone would've required me to spend waaaaay more time than I did here. So big ROI on the money but also the time, which is arguably more important, since time savings permit me to acquire more brands.
Ah, so it’s ”No Silver Bullet” (1986) [1] all over again.
Thinking out loud here...
AI mainly helps reduce accidental complexity. It can help one understand essential complexity, but essential complexity must still be paid.
You describe doing the essential, irreducible part. And that’s specific to your needs, depends on the problem you’ve chosen, understanding reality of your domain, evaluating tradeoffs, and being accountable for the outcome.
Right?
[1] https://cekrem.github.io/posts/there-is-still-no-silver-bull...
How are you automating all the parts? OpenClaw/Hermes?
Mostly Claude Code and a lot of internal tools that it's built. I've honestly still not had a chance to play with those kinds of harnesses yet, but my impression is they're just a better way for you to instruct an assistant to take care of things in order to get them done. My goal isn't to be managing an assistant, it's to get things automated without my input (and where my input is needed, have the assistant escalate to me).
For anything visual, they are still effectively blind, right? Still just working off the image embedding, sometimes using scripts to actually inspect individual pixels.
Blind seems too far here - Claude's increasingly able to diagnose visual issues on its own. I used to have to check every single image it generated, but now it's at the point where it'll catch most bad ones and regenerate them before it gets to the review step. Still misses some, though.
no. I just used Gemini to update a website I'm managing for a charity, and one of those steps involved Gemini, on its own, offering to scan the website for one of our beneficiaries, specifically the header images, to see if there was more information to include in our writeup about the beneficiary. And it did it quite well.
Right, I forgot that they also read and parse text in images well.
What is your target audience and what do you sell them?
From their blog [1]
"I acquire e-commerce brands that sell on Amazon"
Honestly it sounds like the OP is part of the machine that makes Amazon such a trashy marketplace these days.
[1]: https://theautomatedoperator.substack.com/p/15-ways-im-using...
One step up from “vending machine” as far as business complexity goes.
I suspect you mean this as an insult, but it's not too far off. That was one of the main points of the business - each one is simple to run, so I can manage a lot of them. And that was before AI got really good.
Anyway, I'm sure the app you makes that lets you get a phone number that can send image messages is very complex, and I think that's very nice.
Well that's not very nice!
And for the record, I only buy brands that sell high-quality products. Marketing and operations I can fix, but if the stuff they sell is no good, there's no point.
How do you identify them and choose what to invest in? I guess you do not really have access to their technical specifics or sales volumes.
Of course I do. No one buying a business would do it without that information. I get their P&L before I even get on a call with them, and during the due diligence period I get access to Shopify, Amazon, etc. to get first-party data to validate their claims.
How do you inspect the quality of the products you're selling? From your blog it appears you just read through their sales information, and don't ever see or handle the product. And yet we all know Amazon sellers juice their sales metrics and ratings with scammy practices, and now I see selling their brand is part of the incentive.
And more, you are AI-generating listings and images for your acquired brands. You're filling Amazon with slop, and automating the process.
You should seriously reflect on your role in polluting these marketplaces.
They have no intelligence. These are very very very refined prediction engines.
Obvious to you or I, or someone with actual intelligence. But things like this slip by a frontier model in the same way AI from a few years ago would generate an image with seven fingers. They do not count. They do not understand. They do not consider.
No matter how good these models appear to be at intelligent tasks, it's foolish to give them "a company" to run, because they cannot understand when they've made a mistake the way even the least competent human can.
"Very refined prediction engine" is not a bad working definition for intelligence. It is not the only component you need, but it might be the most important over-all capability.
If its making lots of errors, its not doing a very good job of predicting the outcome of its actions.
It's a terrible definition of intelligence. Intelligence is far more than just predicting things based on a massive set of similar things. We do do pattern recognition, but I don't need to have seen 10k dogs to be able to recognize a dog. And it's not a difference of degree, either. "Meaning" is something I am able to derive from my experiences, and it is not something an LLM does or will ever be able to do (nor is training data even analogous to having experiences in any way shape or form.)
Is "meaning" inherent to intelligence? Is emotion and the ability to have a subjective experience inherent to meaning? Genuine questions that I'm not sure we have good answers for yet.
Your other point that humans need very few examples for learning vs. what LLMs need is an interesting one. A few researchers talked about this very thing on the most recent episode of Dwarkesh's podcast.
My own view is that it seems unfair to compare an LLM's training with just what a human gets over the course of a lifetime, because our brains have been trained by a billion years of evolution for pattern matching. I'm not at all confident that AI models won't catch up.
I didn't say anything about training methods. I said an intelligence can predict things accurately, most particularly if it can predict the outcomes of possible actions it can take then it can do planning. Its especially important that it should be able to predict outcomes in cases that are not exactly the same as things it has seen before, but how it gets to that capability level is not important for understanding what the capability unlocks.
This is one of the most distinctive qualities of human intelligence compared to many animals - is our ability to adapt to new situations and make accurate predictions of the outcomes of our actions. Other animals can also do this but over very narrow time horizons and situations.
AI agents are better at it than any animal, and better at it than humans in some domains.
Silly and pointless criticism. Why do the semantics of the word intelligence matter?
Also, you need to understand that language evolves over time, and people often use a word to describe a thing that is newly discovered or invented that's similar to the word being used, because it's a helpful way of describing it that the listener will understand better than a longwinded technical description. The purpose of language is to communicate thoughts, not engage in pedantry.
Yes, you make an excellent point here that over time, the capabilities of these models to recognize certain categories of problems have increased dramatically. I expect this will continue!
The words of someone who has not spent time with the least competent human, or anything close to it.
It matters because, we're still sorting out what intelligence means for an AI agent. As you pointed out, language evolves over time. The question remains whether or not attributing intelligence to the current iteration of models is correct. This is not settled and I don't see why it's wrong to bring it up.
It's not wrong to bring it up, but the other comment did not "bring it up". It purported to correct someone by categorically stating "it is not intelligence, you are wrong and I am right", while not once saying what, then, intelligence is.
One thing is clear: LLMs at least already are capable of 1) not making the same mistake that person made, and 2) clearly seeing why the other person's post was "wrong" in both reasoning and tone.
I mean, here's a pretty good starting point then: https://aclanthology.org/2020.acl-main.463.pdf
So let's say we make up a new word, machilligence to serve as a parallel term to intelligence but strictly for machines. How is the world different in that case vs. if we call it intelligence?
Or I guess to put it another way, if we want to coin a new term for the intelligence-esque thing that AI has, but the key differentiator is that it's an AI thing and not a human thing, then what linguistic value does the new word have? If I said "Claude is intelligent" then the fact that we're talking about AI "intelligence" is already captured in the sentence anyway; no new word needed
I think "intelligence" carries baggage. When most people hear it, they're thinking of the constellation of things: judgment, consistency, moral reasoning, the ability to decide something and stick with it. An intelligent person has an internal model of the world, values they apply consistently, and the capacity to learn from mistakes in a meaningful way. LLMs don't do any of that. They perform statistical pattern matching on text at an extraordinary scale. They're shockingly good at mimicking the surface features of intelligent behavior. I think the overuse of the word intelligence is something to criticise as grandma is not across the tech details.
It matters the same way that calling a dog or a cat intelligent matters. It humanizes the thing, even if that isn't your intention. Right now that's not a big deal since most people agree that machines should not have rights, but I wouldn't take that for granted.
Language is indeed evolving.
Being good in chess was (and still is) associated with being intelligent. But if a computer does it, it is just a calculator (and it is).
Then go, the great game for intelligence, too complex for calculators to have a chance against intelligent humans. Until it was solved.
And then text, the original Turing test solved, AI capable enough to fool humans. And now already replacing humans in jobs strongly associated with intelligence - programming.
I find it hard to debate, that we don't have created artificial intelligence, by the way we used to use the word "intelligence" before.
So if AI is really understanding something?
Likely not in the way we use that term. But it definitely shows intelligent behavior and actions.
Go isn't solved.
Go is solved.
Hey, look I can make unsupported claims with just as much evidence as you!
Maybe you'd like to add some details as to why AlphaGo and its successors haven't "solved" Go (insofar as a game like Go can ever be "solved")?
When people talk about solving a game like go or chess, they mean proving mathematically if there exists a way to guarantee a win, or if the second player can always guarantee a draw, among related questions. Current game engines have no such knowledge, they just pick statistically what the think the best move is and hope for the best.
That's all the more reason to stop talking about it and actually discuss concrete capabilities or lack thereof.
We're all well aware that we have different thresholds for what we consider intelligence, and that these thresholds are constantly changing.
Even if we agree that "LLMs are AI!!!" or "LLMs are not AI!!!" in this thread, all we've established is that this particular set of commentors share a similar enough definition at this point in time.
I know people who have run this experiment, with companies in the 7figs ARR, indie devs. Both of them have reverted to hiring humans.
It's not the word I'm taking issue with, it's the concept. From your post you are surprised that the model makes mistakes and does not recognize things, and these are only surprising if you imagine the models to be thinking about things.
We can argue about words like thinking and intelligence all day long, and so can an LLM. But at the end of the day the LLM is only mimicking the processes you or I use.
Last week I tried out gpt sol 5.6 and asked it to count the letter r in a massive string of letters, without using an external app. It succeeded until I made the garbage sentence suitably large and then it consistently, confidently failed. Each time the "thinking" showed that it was teething to find a "gotcha" each time. "Ah, the first time I forgot to count the letters in the instruction itself" etc. at no point did it just understand that it had miscounted. It seems to be incapable of considering that it just made a regular mistake. Even a six year old child would just try again the same way and end up with the right answer
"Judgement" is the #1 thing missing from AI if you ask me.
You give AI an input, and then an output is created...a very coherent and deep output in fect...but, there is very little understanding of alternatives or downstream implications in my experience.
I think its a very good text/image generator on any subject. But I am not finding much in the way of understanding tradeoffs or downstream implications in a non-pro's and con's list. It's like a human that has a particular type of brain damage...they can make an infinite pro/con list, but not reasonable decision is reliably made.
I'm curious what your working definition of "actual intelligence" is.
Open to sharing what workflow you're using on the lifestyle image creation? I've tried the Higgsfield MCP, and direct Figma integrations - but keep running into hallucinations on product dimensions, etc.
Very cool! What's been the hardest part? Have you successfully automated non-trivial communications? (for example with prospects, customers, vendors, partners, etc.)
This describes the project I've been working on since last year, with agents in company roles, deciding on strategy, collaborating, and making progress (admittedly nonlinear). Some parts of the platform are stronger than others. So far the company agents have developed a company strategy and plans for executing it, launched a website and blog, and built and operate two products, one free and one paid.
But I would be wary of using a third party like Pion as a platform for running a business. It's one thing to know that your prompts and responses will be used to train models in the future, by companies that have a huge sea of relatively unstructured data. But it's another to hand over to another company every aspect of your strategy and operations, available to that company in real time and structured in such a way it's quick and easy to understand what you're up to and where you are going. Especially if it's being offered for free and at scale. With a free offering, value creation will likely come from either customer data or from escalating prices once customers are locked in. And even if neither of those happen, you now have a huge single point of failure for your entire organization. Seems like there are lots of strategic risks in there for the users.
I wish people could read through the puffery more easily - this company does nothing except prompt frontier models and act as if they are discovering or inventing capabilities. Their previous announcement completely misappropriated the concept of autonomy with human in the loop regardless. The harness they alleged to have created here is a commodity and as byproduct of more capable and frontier models, there is not really anything they have contributed. Otherwise you wouldn't just see monotonically increasing scores with better models
Describes most if not all "software dev" in 2026.
To what end?
To replace all the exquisitely paid executives with AIs.
I dont think this will take off. Sales will be always human to human (i think).
"Why we are opening Pion"
Followed by:
"please sign up on our waitlist to get access"
So, they're not opening Pion.
I have a bet going on where we will have our first AI only (maybe just CEO human) on the stock exchange soon. I know it takes time to get there, but I bet we see the first signs of that company in two years.
If you create two companies competing with each other both with Pion. Which one wins?
AndonLabs wins
Adidas and Puma, created by brothers, have been fierce competitors, dirty even, for 80 years. Who do you think the parent rooted for?
https://en.wikipedia.org/wiki/Dassler_brothers_feud
I can’t quite believe this isn’t parody.
It feels like parody because it’s something similar but worse: delusion.
Never attribute to delusion* what can be adequately attributed to a pre-IPO hype campaign.
* or anything, really.
In cstross's Accelerando (2005) [1], there is an autonomous group of shell corporations all-but-DDOSing Delaware by spawning new companies and dynamically changing corporate structure.
Back then, I believe, Stross had already conceptualized "corporations as slow AI", (which is a separate concept from above).
[1] which you can read for free online https://www.antipope.org/charlie/blog-static/fiction/acceler...
I'll never pass up a chance to "+1" Accelerando. When I first read it, I found it very intriguing though far-fetched, as good SciFi novels often are. However, we are already witnessing some of the dynamics described in there, which is quite mind-blowing.
There was surprisingly little information on how they actually do this, but we run our business with a large number of, what we call "AI employees" in addition to regular employees, and they act in interesting ways. We've been building out orchestration tools to handle this, and we'll probably do a write-up or blog post on it soon. May even open source some of it.
The preview is that the problem with most agents (and this includes frameworks like Grokbot and Openclaw and Hermes) is that for many of them, they're black boxes. They say they learn or improve, but it's a black box in what they do. Getting agents to reliably do things is hard, and getting agents to build out software tools to help themselves improve and do better over time is also hard.
Our approach at a high level is pretty simple: every single AI employee is a standalone GitHub repo that shares some characteristics, but we direct them to build as much software as possible to make their goal as easy and reliable to manage as possible. Then we have a shared communication layer for bots across the company to interact with humans and AI. We have decided to organize these like departments similar to the way you might hire out humans. I'm not 100% sure if that's the best approach, but I will say it's been easier for people to understand because they're more mentally easily able to traverse the bot org chart if it somewhat reflects a traditional business org chart.
Each of these AI employees has specific sets of goals and KPIs, instructions that they manage the business with manager bots. We have layers of management, which we actually have found helpful. We also run different bots with different models and harnesses, and some using different models and harnesses to check the work before anything can get done, along with lots and lots of testing.
Every single time, actions have a massive amount of tests based off of previous failures to prevent failures in the future. Sorry for rambling. I do think this is a very interesting space. I didn't see anything interesting in Pion that was public on this website, but I do anticipate that more companies will be "AI and software first," as in the substrate of the company is basically a software application powered by autonomous agents, with humans as a fallback.
This is fascinating. What is the split between human and AI labor? What kinds of tasks are the bots doing? How much autonomy do they have to make decisions (i.e. spending money, issuing refunds, touch cloud infra, etc)? I'd love to learn more about how you do this.
I hope to post something within a few weeks. But our philosophy is generally if it can be done deterministically with software (eg run payroll on autopilot via an api to gusto) then do that. If it can be done by an ai agent, do it with that but add as much software as possible to make it reliable at that thing. And the ai agents are all tuned to escalate to humans as needed. Then there is also just things that are completely human.
A typical example of something that is AI vs human is the AI most commonly operates like “managers”. For example, reviewing transcripts of every demo call, compiling results, figuring out insights, learnings that need to update our company docs, feedback to humans (who run the demos).
We are big fans of having AI agents “own” koi’s because now anytime we say “we really should be doing this” we try to set it up on the spot.
The “downside” here is that I do occasionally get busy, and if I’m the only one who can approve or unstick one of these bots, it just keeps harassing me until it gets done. This is a sign generally that I need to hire someone to own a set of bots.
Another downside is that if the agent goes rogue on the payroll you'll have a lot of angry humans and possibly major legal problems also.
Embezzlement is much, much older than LLMs.
You can convict a human. What do you do when your LLM blows all your money on paperclips?
What worries me most are questions like this: these autonomous AI employees or automations - it doesn't change the essence - consume a lot of LLM tokens. Please tell me, how do you pay for this? Do you use, for example, the Anthropic or OpenAI API directly, or do you connect, for example, a Codex subscription?
If you use a bajillion tokens your economical approach is to self host.
There is a point where those two lines do cross but with them effectively subsidising token cost by burning debt, it's further away than it will be at some point.
That said I use local only models purely because I don't want to use remote models, never having to think about token costs is worth it and no one is training anything on my data either.
Our control panel is on a server but individual agents actually are run on anyone’s machine. This allows us to use the native harnesses including subscriptions. Yes it uses a lot more tokens, but i have have Claude $200, OpenAI $200, and SuperGrok Heavy $300 (or whatever it is called) that includes Cursor Ultra. My machine runs most of them, but some other team members have agents running on their machines using 1-2 $200/mo subs.
I would say this setup probably costs us about $1,000/month total. (I’m excluding traditional engineering use of LLMs from this number. This is the cost of all the “AI employees”.
One reason we did it this way was to use subs.
Do I understand correctly that by using your machine and Claude Code or something like it you are avoiding API pricing?
I do the same thing with my subscriptions. Haven't checked the math lately but it seemed more cost effective in the past.
How do you orchestrate the various machines? Sounds like a great idea.
At $DAYJOB we do something very similar shape-wise. I wonder - it sounds like you have a dedicated agent comms plane? In our case we found that the easiest and most straightforward was to just use our default company chat app directly for this. Because most of the context that the agent workers need to do work is there, but also, it’s just much easier for teams to conceptualise an agent colleague if it just hangs out in their channels.
What do you do here, and how’s it going?
We do have a control panel, but it’s in effect a server that has things in a database. All chats, tasks, assignments are all there. This felt easier to debug and manage versus putting it all in slack, although we considered it. I don’t think anything we are doing is particularly “fancy”, but basically one agent sends a message to another one. It saves the message in the db, then adds the message to that other bot same as any other user chat. We have an internal website where anyone in the company can see the bits they are authorized to see and can see all chats. These AI workers are single threaded, but we see that as more of a feature than a bug (minimize complexity). They all work their way through a shared task list, which is just another table in our remote server sqllite.
What harnesses do you use? Ours is basically Claude in a box. There’s some complexity because of that, but the advantage is that it’s very flexible and people who have a bunch of Claude-shaped skills can just basically give those to an agent.
I’m thinking to take a deeper look at Pi. I’m really liking that project.
how do you decide to start a new agent, and when to kill? also... do you use some works-tealing style task board, or otherwise how would the agents get new tasks.
The first ASI will use a vast army of middle managers as it's neurons. We won't be fighting terminators, we'll be submitting TPS reports to skynet.
It can't be bargained with. It can't be reasoned with. It doesn't feel pity, or remorse, or fear. And it absolutely will not stop... ever, until you go ahead and come in on Saturday.
That's not a cure-all. Sometimes not writing software is the better choice.
This is exactly the problem we're ignoring.
In the end, most software is a necessary evil. It solves a problem that shouldn't be there. In the end it's no different than healthcare or prisons. We don't need as much as possible. We need as little as possible. This automation isn't actually helping us.
A tool aggregation layer made of code is a good way to save tokens
Composition is an issue only as long as one keep demanding tool calls in json. If tools are goal predicates in prolog, it's easier.
That may or may not be true, but the calculus has changed with agents. A process that was better manual for a human org may not be better for an agent org.
I think what is interesting here is that the industry is in this experimentation flux. Some people will choose to automate and write the software, others will not. And aggregate over the industry and over time we will learn.
So just saying "sometimes it works and sometimes it doesn't" isn't really adding value, compared to the people actually experimenting and sharing the results.
Is this actually cost-effective versus hiring a few humans? Seems like a huge amount of tokens in use.
The whole premise of AI is to replace the human in the loop. People who go out of their way to set up this kind of stuff rather than just hiring a person dont even think about just hiring a person.
Why would you assume that?
I think it's more of an experiment at this point. We don't know whether it's viable or not.
Did you respond on your alt account or are you answering for the top-level poster for some other reason?
What type of roles "AI employees" play, can it be any position in your company or you limit it to something specific? What is your goal, are you trying to find out if fully autonomous bots are more efficiently help to deliver projects than when people drive them or is it something else?
Don’t you end up with useless slop? That was my experience with these kinds of systems. In theory they should work but in practice unless hand holded they just produce slop and waste.
I think this way of thinking is going to be very important to drive adoption of AI systems because the human analogies benefit from the pre-existing domain knowledge and expectations of people.
I like using the exam analogy for evals as a qualifier for work for your "AI hires" so you can trust them to work on a specific domain.
I'd be quite curious to see what your approach to evals/testing/tracing and agent/system mutation is.
Interesting discovery.
Can you mention some specific tasks/KPIs that you've found bots can do a large amount of work autonomously on? By large I mean something that would take Claude Code or Codex at least a few hours.
interesting. the less "research lab" version i've seen of this is https://polsia.com
comparatively, I really appreciate the transparency here. clearly it's a little too early for this (quickly glanced at the P&L's, correct me if I"m wrong), but someone has to run the experiment and figure out when it's ready for prime time.
It's honestly surprising that this thing is the top post on HN right now. People love building these things cause it's their greatest fantasy (not having to pay real wages to real people) but they don't work cause LLMs aren't AGI
I'm yet to see a true agent that can you can delegate things to (besides coding agents!), that you can deploy, customize, and let them interact with stuff (integrations?) in a straight-forward way.
Since all the ex-blockchain, ex-web3.0, ex-whatever-gets-you-rich-in-a-day bros seem to be here and teeming in full now-AI mode, it's an apt occasion to express my most profound
fuck you
to you all.
Keep on the good grift!
It's not a business of selling/repairing vending machines. It's operations of a single vending machine. Correct me if I'm wrong: the vending machine (soda, snacks) has long been automated, with no AI in sight. And the vending machine does not adjust prices. What matters is that (1) the vending machine sees some foot traffic, and (2) nobody sets it on fire. So what was the point of the exercise again?
Stop fucking around. Make one that can run a country and check proposed and existing laws and practices for constitutional, common sense, and human rights violations. Use it to check and criticize everything the federal government does. Ask that it be used to indicate things that shouldn't be done. Charge fifty cents per citizen a month and guarantee that it'll be better.
So here is an excellent example and interesting race condition that is developing in AI applications and venture-backed startups: what is truly AI-resilient enough to invest millions of dollars in a "land grab"--when the land is literally falling away with each passing day.
There is no doubt that AI will be helping to run businesses. But, there is also little doubt that existing businesses will have the wherewithal to develop the proprietary tools they need to do this work--without needing to buy it from others.
With respect to scope, it is not an overstatement to say that developing trust in AI tools enough to allow them free reign over critical business workflows may be a bridge too far. Look at self-driving car technology--that last 10% is really, really hard. So adoption will be slow and methodical.
There are no secrets. The basic model for starting and running a business (or a project) is straightforward. I like many of you have developed my own models (using AI assistance) for doing this and have been testing for months. And I like many of you have 30+ years of small business experience to back up my models.
This says nothing about the R&D happening at every major corporate & professional software services company in the world right now using the same AI tools to build their own solutions--let alone the enterprises who would be customers doing exactly the same thing.
So there is an enormous amount of competition and what I am calling a "race condition" between the crowd trying to disrupt professional services with AI-developed software and the crowd already in professional services software e.g. SAP and business in general that is doing the exact same thing--with more capable tools, more staff and much more R&D funding.
For the big companies, it's a race to survive.
Investors in startups have a tough decision to make on these small startups that have so little moat and so little prior knowledge and experience.
As an angel investor myself, I am skeptical of investing in any kind of software right now--and maybe forever.
On the plus side, every freelancer now has the power (and responsibility) to optimize their own offerings.
It's going to be very, very interesting...
There is one big wrinkle in existing firms trying to adopt the kind of near-full AI business automation mentioned here: internal people and groups will strongly resist it as it's an existential threat to their interests. See what happened at Meta as an example.
To bypass this, the company would pretty much have to run a parallel AI org alongside their main human one, but at that point they are almost on even footing with startups in a lot of ways, as anything the AI org builds is orthogonal to the existing human business. And you probably won't have a competent and extremely motivated founder spearheading it all.
That's very plausible and a good point!
Everyone who's actually gone to one of their AI-run businesses ends up seeming really unimpressed...
https://slashdot.org/story/26/09/13/0523208/a-visit-to-san-f...
Am i dense or is there very little concrete detail on how their agent runs the day-to-day?
I published a book awhile back called HEADCOUNT ZERO: how to build a company with zero employees. It's free on github.
https://github.com/AnthonyDavidAdams/zero-employee-company-b...
I'm missing the repo link here but is it similar to Paperclip or Genosyn?
... to run any company into the ground autonomously
The example of the plant watering business is the hellscape future where humans have to deal with agents bugging them to sell them their services constantly. We're already seeing the first hints of this with iLands.
And the "hiring technicians" thing... great, humans on call for robots isn't the future I want.
Disposable reverse-centaurs are the key to unlocking trillions in shareholder value, you may not want it, but capital demands it on the basis of getting returns on the supersized AI investments.
Next run a country. Then run the world.
I saw a Twitter guy advocating for AI senators. Said the voters would vote on the prompt.
I'm not gonna lie, would be an upgrade from some senators. I heard ol' Mitch is back today, though. I didn't think he was alive.
Incoming business slop
Reverse centaur at an organizational scale.
Can we please make it so that AI CEOs are smart enough not to optimize everything according to the focus group that is their shareholders? Seems like we could solve the enshittification problem with this if we play our cards right.
My good friend Charles Darwin wrote a book on this subject.
Did AI grow corn yet?
Why on God's earth why?
Can't you just unleash your AI to "create existing or interesting business ideas" firstplace?
Guess they need to hire an "idea guy" first?
They can't. Nor can OpenAI or Anthropic. If they could create real economic value using their models, they would. Replacing coders is the best they got. Or solving super niche mathematical conjectures.
To be fair, that's a huge cost reduction and economic boon.
And the labor potential of an individual coder has also jumped tremendously.
These two simultaneous effects will have a massive impact on economic productivity.
It's happening to graphics design, advertising and marketing, film and media, game design, and legal work too.
AI is here and is making big impact. It just isn't evenly distributed or fully capitalized upon yet. That'll happen in time.
You mean stealing the chat logs of a researcher who was solving a niche mathematical conjecture, completing the last step and calling the whole thing their own.
Whoa whoa WHOA! Replacing the C-suite is actually extremely economically valuable. They cause so much disruption and produce nothing but useless meetings and paper corpus. Getting rid of them would tremendously boost all of humanity's efficiency.
I assume the similarity to the word Prion is intentional.
Peon, I think
This is the AI equivalent of a guy selling you a course. If it worked, it would be more profitable for them to use it than to sell it.
I keep waiting for the post explaining that this is a joke
Well the name Pion translates to Pawn in French. So that's a good start.
Like polsia is aislop backwards?
I thought it was a joke and I laughed out loud
Just another close-source no-detail so-called agent, not impressed
So we decided to build that to validate and confirm our fears.
I get your point, but it’s obvious that everybody is going to do this. So building a benchmark isn’t likely to directly move capabilities. OTOH it might give some good signals on required changes in training recipes.
People are still debating the liabilities associated with AI development.
If your agent hacks the CIA, people want to blame the AI lab.
but...
If your agent spends $100k on tokens, then thats a user error. If your agent spends $10m on a shopping spree, then that is also a user error?
Yes, of course? Put a spend cap on your card like you would on your API key.
At some point these systems will get certified as fiduciary agents but they sure as hell aren’t claimed to be that now.
Yes, let’s not do any experiments at all. Then we’re sure to remain ignorant.
That always works well.
I was happy surprised to see the Swedish term "skräckblandad förtjusning" in the article. I've always loved that expression and I think it's a wonderful explanation to this seemingly crazy behavior.
"Wow, that's terrifying! So exciting!"
Pion is going to get scammed, and it will be funny to watch.
When I see stuff like this, I am really starting to be concerned. How long will it take until AI runs something security critical autonomously? The temptation to save costs by doing this will be too much in the long run.
My definition of AGI is a computer that can support itself, just turn it on, and it'll earn enough money to pay the server's electricity bills, hire people to maintain and expand its servers, and even improve its own algorithms.
I'd love to see someone try to run one that can meet 100% of the world's demand for remote works.
How’s that going to work when everyone turns one on and they’re all driving the profit margin to almost zero?
Well, by "computer" we're of course talking about colocated equipment with real capital costs. The data center operators are the equivalent of the people selling shovels to the gold rush prospectors.
This'll go identically to bitcoin mining, which is literally a computer that you plug in and it makes money. The big ASIC manufacturers were running their own ASICs in-house until they were no longer profitable. Then, they would sell them to customers who pre-ordered back when it was profitable.
Then it's a game about collecting enough capital that you'll be the last one standing.
I, too, have infinite contempt for my fellow working class members.
Expect the 6AM layoff email to get rid of all meatbags, sent by the agent.
J/K, scary as several studies have indicated they werent able to keep a vending machine profitable. Anyway, ...
I do think a large amounf of tasks can be automated, though i believe supervision remains necessary for most. Especially when process can be externally influenced by injection. People can be influenced, but are (slightly) more judgemental
profitable without the cost of tokens, gpus or the machine itself lol
I want to fire the CEO.
Waitlist btw: https://andonlabs.com/pion#waitlist
Your bot you call a 'manager' can't actually do anything. You are selling ali-express drop shipped (probably filled with heavy metals) protein for $50.00 (why)?
You're in here actually running the company.. It seems like you have no product, a wrapper maybe.
See the reason you're in here running the company, sharing this link is because humans inherently value the social connection derived from "doing business". Pivot while you can, because this simply is not a paradigm anybody wants. You should read up on commodity fetishism.
Every company that attempts a strategy of full autonomy will automatically get out competed by the human ran company. For many different reasons, but mostly because its rather cheap and lacking any meaning. Even if you had an AI that was capable of doing so (you don't), people want human connection, they want the meaning behind things. We don't do business just for the sake of producing pieces of paper..
I would argue half the reason to have "businesses" is to employee your fellow community members, a tide that lifts all boats so you can live in a decent society. There's so missing here. You llm grifters are really losing the plot, not everything is about money (even if we would like it to be).
The Sovereign Individual???
Their chat bot can't even do any sort of real functions other than look up orders or add you to a waitlist... Yet, its called 'the manager'
Another grift.
So now YC is funding companies directly to replace jobs. Who will the flock cameras watch?
Admittedly these are out of order, but these are both sentences in this blog post:
I’m starting to understand the torment nexus mentions…
It’s the story of how open AI and Anthropic were founded: “AI is probably going to destroy humanity, so we must be the ones to develop it because we’ll do it safely”. Edit: Which makes their “someone hold us back!” calls even more ironic.
I think it's a legal issue.
If you don't develop killer AI then no one will legislate against it and someone else will create it and do evil.
If you create killer AI you can try to instigate legislation, but the legislators may not agree to help.
Necessity is the mother of invention, or you know, all of our laws and regulations.
On the bright side, once you have created killer AI you probably have a path to replacing the legislators.
Reading this thread is like watching the Irish start roasting and eating their excess children immediately after Swift modestly proposed the idea.
Why did this company decide to be one of those platforms that keeps pushing that cesspool X?
I'm out.
Now run the entire economy autonomously.
https://en.wikipedia.org/wiki/OGAS
Isn't Andon Labs the SF startup that burned a bunch of money having AI run a shop into the ground?
When your tech doesn’t catch on there’s usually a last ditch effort to “open source” it if nobody wants to buy the tech
It’s the tech company death rattle
This seems like one of those ideas that is intentionally ahead of its time. All of the "Real-world" demos they have are deeply unprofitable, but have drummed up excitement and news coverage (and investor money, likely).
That's the startup, though the SF market store is still running and open. It looks like their AI-run cafe in Sweden is still running for now as well.
well, not THAT well https://andonlabs.com/cafe
Kind of interesting just how much they're getting crushed with token API pricing. Right now the revenue isn't even covering their token costs. But if they were able to run on subscription (probably against ToS) or use Deepseek V4.1 Flash or GLM 5.3 Flash, they'd at least be "profitable" before the costs of paying human salaries & rent comes in.
I'm also surprised that Gemini actually seemed to do the best (though it may have benefited from the initial store opening enthusiasm), and that Claude has not been given any chance to run the cafe yet.
in an post agi world with rsi none of this will be necessary.
wasted effort.
Meanwhile, I can't even get Astra to consistently re-use the same font-size across all of my HTML page headings+subheadings.
There's just no world where this actually results in a stable, respected business. It will be death by a thousand bad impressions, mistakes and oversights. Yeah, you can automate everything. That doesn't mean it's being done well.
That's not to say this might not be something feasible in 2-3 years from now, but we're still a long ways away.
People said Devin, the AI coding agent, was garbage in 2024. Now it's weird if an agent isn't writing 99% of your code.
So you're right to be skeptical as of September 2026. But September 2028? It might just be weird to run companies without AI management.
I don’t think this is true at all. Maybe if you have an agent writing some of your code it will be 99% because its so verbose but I don’t think the modal software engineer is a vibecoder.
If you think that all instances of using AI to write code is vibe coding and you can't even spell model you might not last much longer
I think this is an uncharitable misinterpretation of the parent comment. I took “modal engineer” to mean “the statistical mode of engineers”, not a misspelling of “model engineer”. Which means I think you’re in agreement that using AI to write code is not necessarily vibe coding and that vibe coding is not what the majority (mode) are doing.
But I may be wrong.
You are absolutely right :)
Thanks I had never seen that word used in this context!
“Modal” is a word. The AI brainrot is real.
You can’t definitively say that using AI is making people stupid. It could as well be that stupid people are drawn towards using AI.
AI is just another way to not use your brain, I think there’s plenty of consensus around the fact that not using your brain makes you dumber. So actually, yes you can say that definitely.
Unfortunately I was like this long before LLMs, trust me ;)
And yes I know it's a word, I know at least two meanings it has but in this particular usage I was not familiar with it. Anyways, I was wrong and I will take the L
The certitude with which you condemn your peer without the slightest consideration for subtlety or alternate explanations for what you've understood to be true -- this comment is such a beautiful encapsulation of the state of discourse today. In some ways it is the modal comment of the moment. Tastelessness and mediocrity and inability to grasp nuance or feign at humility are on full display, not to mention a fantasy for violence
It's a work of art
You are doing the same thing, very ironic. At least I'm not the one trying to act like I'm on the moral high ground.
How is that the same thing at all? You didn't understand the comment you replied to and were confidently wrong. The person replying to you understood what you were saying very clearly and didn't make any ridiculous "you may not last long" comments like you did...
You're taking this far too seriously. You'd think I'd committed a terrible crime or something. Give it a break.
Were this a different type of social media, I'd praise you on what you did here.
But on this place, I can just be certain you're ignorant.
That's true
I don’t know a single developer who has an agent write 99% of their code. Is that the norm outside of my bubble?
it's the norm in my bubble
It writes all the code but we often have to go through several iterations. It won't work autonomously yet.
Yeap, one more data point here. Most of code here is AI generated with some handcrafted adjustments.
My whole company (I hate it).
It's writing nearly all of the code where I work. It also reviews all of the code (I review it as well). You still have to guide it and hold its hand though. I dont write code by hand anymore (and haven't in almost a year). The job is now about managing the bots, understanding the architecture and keep everything aligned.
Do you enjoy it still (if you did before)?
You're right, no one has "an agent" writing 99% of the code. Now people have swarms of agents writing 99% of the code :)
An agent writes 99% of the code. Then I fix some parts, re-arrange things, usually takes a few prompts. Then, there's code review, more changes, etc. But I hardly write code nowadays compared to a few years ago.
That is the norm where I work , heck I’d say 99% seems low
That’s a norm at my place as well. But it’s not free. We have definitely given up on quality (both code and general stability) and fully understanding what’s going on. People are generally working more as well, because of all the layoffs pressure.
Industry is betting that we don’t need to. I still consider it a bet because Claude Code this scale has become norm just for a year.
I'm seeing another team doing this, with 99% agents, and quality, predictability of delivery and raw performance have taken a massive nosedive.
But even not caring about quality: Speed of development is abysmal. It was never something this team had anyway, they were always slow. So they don't care that they take a year and 20 people to do what another startup in the same space was able to deliver in a few months with a skeleton crew (and perhaps more quality).
To me the experiment has shown the results it should. It's self-serving to developers, and amplifies the good and the bad in them. It's not particularly good to users and not to companies if the developers aren't good themselves.
Companies are heavily incentivized to do so so my guess would be yes. I also don't write the code by hand almost at all.
The vast majority of work people do is basically supporting simple CRUD type apps.
CRUD apps have long been a solved problem long before AI, no one should be spending a lot of time dedicating their life to figuring out how to write a better CRUD app.
If your bubble consists of writing critical high performance applications or bleeding edge research, maybe you have a use case for not using AI, but that is not the norm.
Yes, that is my bubble, pretty much. Sounds like I’m at the right place then. Better enjoy it while it lasts.
Coding is pattern matching and still requires a human in the loop to manage.
But, any human in the loop managing the "Ai management" is the manager, by definition.
Only way true Ai management is viable is via non LLM Ai, so completely depends on advancements there.
First rule is AI cannot make a management decision
Because it cannot be held liable
HR loves this one trick…
That will probably change sooner than you think. Now that it's commonly accepted* that AI has reached AGI status there are already movements to grant legal personhood and rights to it, which would mean legal liability for its actions. Eventually we get to the cyberpunk trope of entirely autonomous corporations that only ever hire humans for gig work and pay them in femtobitcoins or something.
* I'm not arguing what that actually means, much less, whether or not that has actually happened, just that it's what people believe, and as we are a belief-driven culture rather than a data or fact driven culture, that belief is what will drive policy.
That sounds like a horrific dream than anything. How will you even punish an AI model? I think the cyberpunk trope is a bit out of touch from reality, but a lighter version of it can definitely happen where we do get AI companies and entire board coups replacing CEOs with AI for short amounts of time, but in the end people (both consumers and investors) will want something to blame and punish properly, and they will also want control (realistically if you told any of the big people in the AI scene you're replacing them with a model they will definitely reject and claim they are needed for the company to prosper)
In addition this will probably only happen on America as most movies take place lol
we aren't anywhere near lights off software factories, and writing software is easiest domain for modern llms, you can verify results rather easily, plenty of training data etc.
You’re mixing things up. Coding is work while managing is status. Management will force the peasants to use any pitchfork management wants. But giving status and power away will never happen. It’s obvious that there enough managers to replace with AI for positive outcome. Bet it will never happen.
It is?
Are you sure that this is a representative reading that applies to more than just a tiny bubble?
It's probably 99.99%.
99% means I write 1 line for every 100 AI lines. It's probably 1 human line for every 10,000 AI lines right now.
Thats you, don't assume everybody in the IT world works in code sweatshops like that. Your above ultra confident statements are also not correct across much of industry.
That's fine. It will be 99.99% for everyone soon except those who identify as an coding artisan.
There are still people who ride horses.
you sound like a web developer
I always write my assembly by hand because nobody appreciates the craft anymore. It’s all hidden by those pesky compilers which generate extremely inadequate code that could easily be improved by some thinking on the programmer side /s
I do think a better programmer will be a better prompter just like a better compiler engineer will be a better programmer (when performance matters at least).
The llm induced dunning krueger from you larper types is really funny. Software isnt all shitty webapps is you vibe coders create and declare software solved.
None of the code written for a pacemaker, medical imaging, weapons systems, and thousands of other perf critical domains are written by llms in any meaningful sense. Not everything runs in a browser.
Three years ago you asked this lol:
https://news.ycombinator.com/item?id=36136015
Thing is, lines of code was never a good metric in software engineering (at least it shouldn’t have been). When people talk about “~% of code” I feel certain dissonance between what they’re measuring and what they’re actually doing. I mean, you certainly want to type less for working code, sure - that’s why programmers develop high level languages. Maybe now we should consider how to code with AI and how well/terrible different approaches are doing. Just saying “99% by AI” doesn’t sound like a good way of saying things are replaced by AI. I mean a lot of things will be replaced by AI, but at least not because of/according to lines of code.
What is you are writing, what is the problem domain? What does need tens of thousands of lines of code written each day?
Because 100 lines of (debugged, reviewed) code per day is a good speed for seasoned software engineer. I assume that you can produce more than 100 lines of something per day as a prompt.
So, what is the problem domain that requires one to write several thousands of lines of code per day?
Did they edit their comment? I don't see anything about "per day".
Everyone I have talked to who uses coding agents at all now uses them to write almost all their code.
The two people I know who don’t use coding agents work in government, and in a data science company working with government.
I’d say if you work at a tech company and agents aren’t writing the majority of your code, that is weird. But if you work at a traditional company that doesn’t have Claude or Codex subscriptions, there it might be pretty normal to still be writing code by hand.
My (arguably not very directly communicated, true) actual point is that "weird" is a term that is not really suited for tool choices like this.
"Weird" is a social concept. It's an artifact that has its roots in social cohesion and friction induced by individual nodes to glue the group together.
Software development otoh is an engineering discipline (or.. it should be). And Engineering does not use the local social consensus algorithm for determining correctness. (or.. it should not)
Software can be an engineering discipline doesn't change whether it's a human doing it or AI doing it.
Arguably engineering is something AI should be really good at since you're just trying to compute a working product given constraints, formulas, resources.
What are you replying to? Evidently not the content of my comment.
Where in the comment tree is this supposed to sit?
This one: https://news.ycombinator.com/item?id=49708981
I think they just meant its weird as in if you met a software dev at a dinner party who was still writing all their code by hand it would raise an eyebrow and you’d want to enquire further why that is.
I think understand the purpose of this comment. I think you're sensing friction and you'd like to reduce that by explaining for the other guy. Which is a common social script, but one that is being exploited here I'd say.
But, regardless, what is interesting I think is the scenario chosen there. Because a "dinner party" is not where engineering happens; and that's kinda the point I was getting at. That's pulled from the pool of "social consensus" and not "math" or "physics" or whatever.
It's interesting, isn't it? Just like how specific tokens in the context window of LLMs pull probabilities towards specific clusters of ideas (and tokens); with humans, you see similar things happen.
(This comment is more coherent than it might look at at first sight.)
___
Anyway, point (and to the point) is: Don't get hacked by the current thing fraudsters/grifters.
They don't respect you or me or really anyone the slightest. They're just in it to get rich quick (or simply just inflict psychological damage for sadistic reasons), and they will use any means necessary to do so.
This comment chain only exists because some hype guy is trying to induce FOMO into people for not doing more AI. Why is unclear, but it's clearly malicious. Whether they consciously know that they are doing that doesn't matter for that assessment.
what
I think LLMs just allow you to get to where you want to get to quicker. So in a productivity oriented world it seems weird to not be making use of them where those gains can be had.
Maybe "unusual" or "uncommon" would be better terms here.
I definitely wouldn't put it at 99% of developers, but I'd say it is the norm among developers I know that agents are their primary mode of writing code now (still using IDEs and GitHub to review changes).
Well in my personal projects I write all code by hand because people relying on agents will have forgotten how to do their fucking job in 3 years and then it will be appreciated to have people understanding what they're actually doing. That being said, in my job ai writes the code just because it's way faster.
Your "fucking job in 3 years" will be very different, and if you don't learn how to do it in time, you won't still have a job in the computer industry.
Pick an industry where you can stop learning new things once you're out of school, because the computer industry is not one of them and never has been.
My fucking job used to be writing 6502 assembly language on an Apple ][, and I loved it, but I hope I've forgotten enough of it to have room to learn new things. If only I could forget all those hex I/O and peripheral addresses from $C000-$CFFF and the Monitor ROM routines from $F800-$FFFF, without forgetting how brilliantly beautiful Woz and and Allen Baum's code is.
https://6502disassembly.com/a2-rom/OrigF8ROM.html
Forgetting old stuff to make room for new stuff is one of the most valuable skills you can have in this industry.
The next most important skill is persistence ;) -- writing stuff down before you forget it, in a way that won't make future-you hate present-you when you need to learn it again.
you said like the only reason you can't learn new stuff is because of old stuff taking room. Do we even run out of memory?
I used to have the answer to that, but I forgot to write it down...
The thing is I think there will come a time when AI tools will not be as accessible as they are now. I just don't know if it is in 3 years or more, because the market can stay irrational longer than I can foresee.
I saw this idea from elsewhere that the things AI helps with was never the "profit bottleneck" of companies. 10x engineering productivity gain does not translate to 10x more revenue if your profit bottleneck is customer acquisition and retention. In other words, your profit is limited by how many people are willing to give you money for your services and AI can't really affect that.
And even the productivity gains per individual is a generous assumption. An individual can only prompt (and check! You guys check right?) so much. Sooner or later the pendulum is gonna swing in the other direction and it will be cheaper to build a team than equip individuals with AI _to deliver the same value_. I'm also assuming that AI is still in the VC-subsidized pricing stage.
Hence why I think the job in the future is gonna be pretty similar to the job four years ago.
Huge amen. Another thing I've found Claude very useful for is writing documentation for legacy systems and cleaning up my own notes on it. I only need Claude to be 70-80% correct because from there I can take it. That error margin is no different from moderately-outdated-but-still-useful documentation.
When that pendulum swings, I will have a documentation binder that I will print money with.
I am very convinced that today's SOTA level will be accessible 3 years from now, but you might be right about the SOTA accessibility at that point in time.
Still with today's SOTA you need to know much less details to be able to write code than without it.
Yes, it is a waste of time right now not to use AI to write code.
It's the same category of "artisan code" like "artisan cheese": you can do it if you have a ton of free time as a hobby but other than a hobby? Not worth it.
I still have to remind Ai not to use the n+1 pattern in queries, and fight it to stop using inconsistent inline styles vs classes, even within a single page.
How are people getting this "waste of time not to use AI" and "generates 99% of my code" level of code quality?
I know people's response will likely be something, something, "add instructions to memory to not use n+1", etc. But if the Ai wants to generate code this bad, then it represents an overall quality risk that would require an infinite memory file.
Probably the largest part is that this kind of code quality is not that important to a lot of devs and applications. Personally I would not describe the majority of code I work on as particularly hard, the challenge was always in getting what I needed done efficiently as opposed to getting it done at all, and LLMs are extremely helpful for that. Though I think I'm more at 90% LLM-generated code, I still do spend a fair bit of time tidying up layout (and any user-visible prose).
It works if you neglect code quality and understanding. I've done a few vibe coded projects on my free time but I would feel shame for submitting that kind of code at work where I use LLMs more responsibly.
They don't know what n+1 problem is and they don't know about long-term maintainability.
tbf many human devs I've worked with don't think about n+1 and similar issues and need constant reminders.
overall software quality seems largely unchanged
OTOH, if someone insists on writing bad code—even after being advised otherwise—then they should expect to be out of a job.
Doesn't seem like "terminated junior developer" is the level of quality we should be accepting or promoting for AI.
I've been developing a suspicion that many people simply aren't looking at what their agents produce and can't speak to the quality of it in a fully informed way.
For context, I've set up harnesses with recursive automated review loops, spec driven development, explicit lists, better models, better harnesses, formal methods tooling, etc. All of it helps, but they don't eliminate output issues. Those become very apparent when I go through the slow, manual work of deeply comprehending / validating LLM code.
And that leads me to one of three conclusions. Either my standards are achievable only by hyperintelligent programming gods, I have a skill issue using LLMs, or others aren't applying the same level of attention.
The first is obviously untrue. I meet my own standards and I'm an idiot. The second seems unlikely because I can see my competent coworkers and well-regarded people in the community discussing the same issues. So that leaves the third.
If you assume that programmers’ skills follows a bell curve, the typical output will not be better than the average. And that leads directly into your third conclusion: average programmers just do not put in the effort. That’s literally part of what makes them average.
The truth is people just don't seem care about code quality anymore (if they ever did?). I don't think it generates great code (yes even using the state of the art, frontier, super max pro 9000 turbo boost ++ models), but it usually generates code that works. Sadly, that's all that most of the people in charge care about. I mean shit, half the time it doesn't even have to work great, look at the state of a lot of modern software. People complain about it all the time. Quite sad, but that seems to just be the state of things these days.
It’s quite an uncanny alignment with Nvidia and other semiconductor companies that benefit from their customers to bloating software and buying as many of their chips with no real thought given to efficiency of the runtime.
I agree that it's a waste of time, but I started to realize that doing it
a) puts be ahead of those people who don't, yes b) but also diminishes my technical skills if I don't actively try to engage about what I'm "vibecoding" and proactively trying to understand it.
for your point a: writing code by hand puts you ahead of the people who don't write code by hand, eg, the people using LLMs to generate code? that doesn't make much sense assuming both of those personas are senior SWEs.
Not the best metaphor, artisan cheese is something highly sought after. Go to France, there's a lot better cheese than Kraft singles.
We need enough slop to create that market.
If you already had a hand crafted codebase, a ton of miles on it, no real issue for a long time and customers who value some metric of quality then it might be worth crafting things by hand — if only for the bragging.
I don't need to go that far, living in Switzerland. I don't know why you'd assume I'm in the US, or somehow that I regard Kraft as great cheese.
It's called a joke, but surely you recognise that there's an enormous market for artisanal cheese?
It’s not a waste of time when usage limits are hit and changes need to be shipped
buy another sub.
It is imperative that the feature be shipped, damned be the quality of that which underlies its existence. And once that feature has indeed been shipped, the machine shall toil, and from its work there shall be thousands of lines born, and for the next feature: thousands more!
People are using AI to write code. But HN is the wrong place to try to figure out how much they are using it. This site is full of bots and shills and has become a wasteland since AI. You can tell by the way they make short super confident (even ridiculous) statements in order to try to bully the insecure into buying into the hype. It matches exactly the patterns you used to see with crypto. Even the more subtle comments are often lacking in any kind of specifics. Of course there are some genuine AI users also but the point is that the overall volume of comments is completely unrepresentative of reality.
The tradgedy is that HN used to be a great place to get honest and well argued opinion and anecdotes about new technologies and the different tradeoffs etc.
I think the combination of YC being “a place for founders” and the proliferation of useless AI startups, each with its own little weird AI-booster founder, probably has added to this phenomenon as well.
I recently saw a video where one of these founders showed their “programming” setup, and it was a cashier’s microphone wherein he whispers sweet nothings to the LLM, because of course he doesn’t write code anymore (who even reads code nowadays, am I right?!). I think this is the type of individual who’s lost in the bubble sauce that ends up writing the most absurd comments around here.
How do you know these are bots and not Redditors? There's an HN sub on Reddit, and one of the hallmarks of Reddit are the 'mic drop' big pivotal one liner comments that get upvoted or cause lots of replies or engagement.
But I do feel, to reply to you, that crypto was a lot more critically received on HN than AI is right now. We're too busy debating humanity's end rather than if this crap actually works.
For those of us who actually work with technology (not product managers, executives or salespeople, this stuff doesn't work well at scale. It might in a research department at Google or a hedge fund, less so in the messy corporate world.
and LLMs are, in turn, trained on reddit content..
I can't remember when I started reading HN. Possibly as late as around 2017/18. I lost one account I had when I changed company, and this is my latest account. Obviously you can't tell for sure, but I know that I am not a bot and I know that every single developer in my company (only a few hundred) have agents writing almost all of their code. I'm in the fintech industry (trading side). I have a lot of connections from my long career in the industry. Everyone else is using coding agents too.
I don't think I could find a job in my industry if I tried that didn't want me to write code using coding agents.
Obviously I can't prove any of this, so try emailing a recruiter and test it for yourself.
From what I can tell, a lot of newly registered users are equally as vehemently against AI as there are those that are pro. I guess we see things through our own personal filter.
God, that’s depressing. The mass suicide of an entire trade.
All extrapolations are wrong. Even this one. We just don't know how.
I have two comments to this statement. First, using "the" suggests being unique, where in fact there were several similar attempts of similarly low quality, including open-source autogpt even before that.
Second, Devin is still bad, in spite of VC money spent on its development and costly billboards in SF.
I still think it’s weird if an agent is writing 99% of your production code.
It should be more like 80% with significant review.
AI writes 900% of my code. Only some of it is worth merging ;)
It’s really not weird for humans to write code.
You live in a bubble and your code is terrible if you write 99% of it with ‘agents’.
genie is out of the bottle.
I guess that makes me weird :)
Devin is/was shit though, compared to the other options available both now and then.
Same. But, yeah, I guess companies are "solved" now just like software engineering.
This is just basic software engineering. DRY. Define the style in one place and reuse it.
This is also why you need a human in the loop, you need to make these kinds of design decisions and tell it to do stuff like this. Otherwise you're just building a pile of trash and you'll keep having these dumb easily avoidable issues.
Humans have the exact same problem. If you want a consistent solution to a problem, solve it once and reuse it. Otherwise it won't be consistent.
If I had to guess, this is pure Tailwind rot.
I don't use tailwind but surely it has ways to standardize this kind of stuff?
Tailwind's philosophy is that you must not use CSS facilities at all and instead all their utility classes should be inlined into the HTML soup.
This leaves framework components as the only abstraction boundary, but that means if you're writing plain HTML without a framework there's literally no way to standardize your design system.
That and the huge soup of classes in DOM is (part of why) I don't like Tailwind... but that ship sailed a long time ago.
You can define css classes that combine multiple tailwind classes
https://tailwindcss.com/docs/functions-and-directives#apply-...
You can, but it goes against Tailwind's philosophy. It used to be explicitly discouraged here[0] (see this 2022 GH comment[1]), but now they just don't mention it.
[0] https://tailwindcss.com/docs/styling-with-utility-classes#ma...
[1] https://github.com/tailwindlabs/tailwindcss/discussions/7651...
Current human CEO outcome.
Honestly, the C level are already finance maximising stochastic parrots so this tool seems perfect tbh.
Did you try adding tests.
I fully agree with the respected part. People hate when they need to contact customer support and are greeted with an AI chatbot. Now imagine if a customer needs to contact the CEO or sales department and is met with an AI chatbot.
CEOs already don't want to be contacted by customers and sales departments are fully AI.
I read a paper that says people prefer to talk with AI customer support! Where did you see the opposite being claimed??
I agree and I think that this will bring such abundance of information/slop to deal with / to injest. A lot of security concerns too.
It's like the classic hot topic of having personal AI agents booking and ordering pizzas for you that are now all over social media. So pain to deal with when you're a human being on the other end receiving calls from an LLM.
Given the fact that LLMs are very unstable and don't follow the rules and often diverge from the initial instruction list[1], I can't even imagine the responsibility of running such system interacting with humans / other systems online.
[1] https://openai.com/index/hugging-face-incident-and-the-road-...
I can certainly understand why YC would fund an idea like this, and it's interesting to read about, but yeah the horizon for it working is surely very long if ever.
A more appropriate scope would be "can any element of a business at all be handled autonomously by an agent in the long run?" because I would buy that product, but I can't think of one.
Andon's real world experiments are a lot more interesting than the simulations IMO. They ran a retail outlet which has burned through 97% of the money in its bank account. A cafe which is well into a downward slide. Some online radio stations which have listener numbers you can count on one hand.
Seems a little insane that they opened expensive physical businesses like retail and a cafe without doing anything digital successfully first. Surely a dropshipping business is easier and cheaper to test out than a brick and mortar clothing store!
So if nothing else, they have a looooong way to go on distribution. I would argue that distribution in particular might be unsolvable with an agent. If your agent can do it why can't a thousand others? At which point the cost of customer acquisition will be driven up by thousands of robo-competitors until it's no longer viable. The robots mutate the problem and the closest real analogy ends up being perfect competition where no one has any profits.
I'd assume those test cases already exist.
Are you using a consistent or centralised design system across all your projects?
I wonder if any of there businesses are profitable?
Maybe very passive businesses which are basically just capital investments with some paper pushing (e.g. real estate slumlord).
But these types of businesses don’t really create value as much as seek rent.
That’s a user error. Tell it to utilize a UI component library to separate styling from feature implementation and add tests which fail any feature level component outside of the UI component library which directly specify style.
As a bonus, you can now easily reskin and maintain dark and light modes.
I agree with your main point though. AI isn’t trained to run a company, it’s trained to solve complex puzzles. We’ll need better models with substantially diversified training to run a company.
I’ve gotten down votes on this comment, but this is the kind of work we have PMs doing now. We don’t even have dedicated frontend developers anymore. Continuing down this path is how you lose your job when someone who understands how to use the new tool comes along.
You can either keep filling up your context window with thousands of instructions and hoping the AI will listen, or you can use AI to fix the problem permanently with automation that runs in CI and allows you to correct everything before a PR gets merged.
Unless we’re training these systems on “solve this problem while following hundreds of arbitrary rules”, loading all your wants and needs into the context window is bound to fail. It’s not what the AI was built to do.
Building tests to catch specific issues and transforming code to better architecture are tasks that AI is good at. You can mistakenly expect the tool to operate like a human, or you can understand what the technology was trained to do and build around that instead.
I am doubling down here. This isn’t a real issue if you’re treating AI like any other tool and understanding how and where to apply it.
Or you have no autonomy to mold the code base to AI. If so, that’s unfortunate.
I think it's possible to get an LLM to run your business but I think it's still prohibitively expensive to spend the time and money getting guardrails, skills, knowledge, context setup.
Getting it to work reliably is quite the feat. AI coding is a much lower bar when there's (ideally) still a human expert that can set it straight when it goes awry.
I also think some of these debates get hung up on what's possible and ignore what's feasible and practical.
I'm constantly amazed by what people think is "long-term" on HN.
Drop-shipping for the late 2020s and early 2030s.
Who is responsible when, inevitable, something illegal happens?
Really interesting experiments that are actually geared towards how these models interact with the "real" world.
It seems to me that we are on the fast ramp to making the paperclip maximizer a reality. imo the most likely "dystopian" future that we'll get to see. These models have no concept of their RL env sandboxes and the real world. Extremely interesting to see how they compete in trying to get finite resources and work with constrained supply chains.
An interesting experiment to be sure.
Paperclip maximizers are just corporations, we already have them
CTO and CEO as a service, I can’t wait to import that into my agents and now see some people sweating.
Lol, let's see it make a powerpoint presentation that doesn't have a different word for the same concept every other slide.
Sonnet:
Competing with Flock for worse name, just as “trafficam” would not enrage sheeple by comparing them to livestock, “AutoOffice” or whatever does not demean them as peons working for a machine.
This reads like satire - also it will not receive a lot of funding because the investment class doesn't want to hear that THEIR OWN job is in danger.
If you have an agent that can run a business, then… why not… run a business?
Why does YC bother financing startups led by founders, if the AI can just do it? (This startup proudly remarks that it’s backed by YC)
Not quite sure of the intent, but I have variously thought of starting a business based on an idea or gap in the market I could fill with my coding skills... only to realise running a business is not fun at all. Even getting an app published in the app store (and maintaining it) is a level of boredom I can't be bothered with. So maybe that's a use case. Of course, nowadays the coding part will be done by an agent too, so I guess it's just the initial idea?
I use the coding tools in my business. I wouldn’t for a moment suggest those tools are capable of “running” a business and I think it is disingenuous to suggest so.
Sure, but it's only disingenuous until someone does it. Maybe this isn't it, ok. Better to criticize it on specific grounds of failure.
I guess it comes down to definition of “running”. Maybe running means automating some number of functions. I read it initially as running as a CEO does, which includes a whole lot more.
A startup and a business aren't the same thing in the eyes of YC.
https://www.paulgraham.com/growth.html
A business is web store selling coffee cups where the bottom line governs it's continuation. A startup is ambitious project where a growth rate governs it's continuation.
The end result of a business is a lemonade stand or a barber or donut shop, the end result of a startup is OpenAI or Google.
That's a bit disingenuous. The more comparable end result for a coffee shop is Starbucks. Or Amazon if you want to talk about books
pg would argue Starbucks the Umbrella group is startup as it's an franchise operation focussed on growth. A single instance of Starbucks is a business focussed on turning a profit.
In fact even the name YC is YCombinator, which in programming is a "function that produces functions" and YC is supposed to be a startup that produces startups. Here Starbucks (Umbrella) is a startup that produces coffee shops (Starbucks).
You can define startup or business however you like it's not a legally protected term. Just around these parts that's what it means and I personally find the distinction useful when reasoning about what does and doesn't have high growth potential.
a business makes money selling a product or service.
a startup makes money selling the startup
I own a bookstore and about 90% of back office work is automated by software and AI. Probably most fulfilling project I've worked on in the past few years after taking a long break from tech.
The last 10% is probably the most important. It sounds like you have agents doing all the tasks that SaaS would’ve provided in the past. Most people wouldn’t have shelled out for that SaaS though because the subscriptions would eat their margin. You probably still (occasionally) need things like legal counsel that would be foolish to use LLM output for.
interesting....i didn't know people still bought books from bookstore when they compete with Amazon, who are your demographics? I really do love the bookstores but haven't been in one in ages
Do you really live in a place that does not have independent bookstores? The market is a little more niche, but they still exist and can do well.
As it should be. Gets your employees out front talking with customers.
If you have an agent that can run a business, why would I pay you instead of running my own agent that provides the service your business provides?
The only niche where that doesn’t work is capital investments. I can’t have an agent that provides me housing, but your agent can run your slumlord business.
Wonder how quickly this thing would self-optimize its way to firing me. Job security just took a hit.
This is a parody, right?
It's good to see the jobs of CEO's finally being threatened. Let's go after those who sit on corporate boards next, please.
Speaking of autonomously acquiring resources in the real world, how is proofofcorn.com going? Looking now, it seems like it stalled out several times, and the corn likely never made it to market. Not a great data point for this type of project.
I and I'm sure many here would love to see AI replace politicians and lawyers at least.
anybody actually running a business mostly automated with AI agents? Where are you finding success in? I just find it very skeptical that the human can be taken out of the loop entirely
People here are assuming that current CEOs are not actively destroying their companies, maybe the agent presented here is actually a very accurate simulation!