- 183comments
- 286comments
- 145comments
- 29comments
- 121comments
- 11comments
- 92comments
- 66comments
- 44comments
- 17comments
- 3comments
- 6comments
- 15comments
- 24comments
- 31comments
- 21comments
- 1comments
- 245comments
- 43comments
- 3comments
- 105comments
- 52comments
- 44comments
- 21comments
- 32comments
- 22comments
- 160comments
- 29comments
- 160comments
- 6comments
Why would they use Notion in a system like this?
Seems like setting yourself up to have a massive pile of janky cruft.
managers that love notion are the cruft you're speaking of
My team is in the agentic orchestrator phase. I like this software factory pattern in concept, but our biggest challenges in development are acceptance testing of anything UI-related. Mobile app testing in particular is still a huge bottleneck that requires a human. AI models really suck at identifying poor usability and jank, particularly because they only typically process snapshots of the app from an instance in time.
I know there are traditional testing frameworks that can detect jitter and frame drop to a certain level. We could potentially start having agents build that in.
If we had concrete designs and specs on every project, that would also be helpful, but in a fast-moving startup, that gets delegated to the builders. That puts a human back in the loop every time.
Curious to hear what anyone else does to fully adopt a software factory pattern.
If your UI is created for humans to use, you should still have humans involved in testing it.
I totally agree with this. Agents seem tremendously bad at UI to me. Maybe it's just because I am a back end guy.
Right now I'm working on a declarative UI framework which can help me along here. My thought is that if I sacrifice a little control for sane primitives, that will make that spec /build loop easier.
I think ClayUI is a really interesting "reduced instruction set" for UI. I don't know that immediate mode UI is the right call for anything web related (that's how you get React lol) but his reduced primitive layer is very interesting to me
After your first para, I was about to suggest exactly that (well, maybe not writing your own). I find that frameworks (both front-end and back-end) constrain the LLM's choices and result in both sensible defaults and improved consistency.
Of course, you will immediately hit the problem all frameworks have: customer requirements that the framework components don't quite meet.
Still, great for RAD.
Can you elaborate on what frameworks you've used? and how they've helped
Missed the part in July when the token bill starts rolling in
May I ask what you’re making that somehow is improved by turning everyone into chat bot controllers?
There was/is a believe that more folks can unblock themselves with access to AI. There is some truth to this. Our product managers and sales folks are achieving more with Claude and MCPs pointed at tools they use.
The downside, however, is they are also unknowingly digging themselves into holes. For example, we have AI-generated skills that are thousands of lines long and include Python scripts with hundreds of lines of tests. Some of these Python functions are literally just emitting MCP tool names.
Are there known examples of software where the "software factory pattern" (an incorrect term, given that it's not a pattern but a workflow. Or rather an idea. A hope. A wish.) proved to work long-term?
Preferably ones that I could validate myself instead of just having to take someone's word for it.
It’s all experiments at this point
Have you produced anything shippable AND maintainable?
and economically viable
I don't believe the "software factory pattern" can even be established as a pattern. I wonder if it's considered normal these days to call a Silicon Valley management style a "pattern."
Fundamentally, a pattern is a reusable solution to a recurring problem. But what the author is describing here is simply an operational loop and a management strategy.
I think current agent methodologies are practically indistinguishable from human developer management theories. Isn't this just a feedback process? I would consider this an agentic workflow.
Also, people casually use the metaphor of a "factory" when churning things out, but a factory fundamentally operates on "orders"—it has a specific objective and a target production volume. The entire approach of just trying to dynamically respond to everything under this label feels somewhat contrived.
Furthermore, I always have this underlying question: why are we building software factories in the first place? Where are we manufacturing the people who will actually buy all this?
I've heard stories of people finding success by running "app factories" in the early 2010s when apps were scarce. But in the agent era, I believe we are already drowning in AI slop. We have to remember that back then, producer friction was high, and simply submitting an app was a difficult hurdle.
Now, AI handles most of the basics by default. Things that were once highly praised are now just the baseline. To create something actually worth selling today, you have to break the existing grammar entirely, build much more complex architectures, and offer deeper features. Given this new standard, I seriously question whether mass production is the right answer.
You say this but then
To do something much more complicated, you need to rework the fundamental ways of working which is exactly what the author is proposing. I think you are already sold that this is the base line (but it doesn't seem like it based on your post).
Will has trouble with this in a relatively nimble org, where he has good control of what is adopted. Imagine how much fun this is in more ossified enterprise environments, where getting work done efficiently means bypassing process with dev lead and line management blessings. The differences in performance among individuals has never been wider, and not just gains in productivity, but performance losses as some people who have iffy judgement cause trouble a lot faster.
Maintaining quality when an organization just lacks the muscle to make changes and nobody has the mandate to try to make sure uptime is in good shape is a challenge. A lot of things break, precisely because there's just as much change as Will sees, but it's very unevenly distriuted.
Has anyone demonstrably gained market share over a competitor who is vocally _not_ adopting agentic software development flows? This article outlines a lot of process churn without a clear through line to how it is impacting feature development or revenue.
I suspect that the vendors out there who have not adopted AI development and are losing market share to competitor who has are not vocal about not being an adopter, they are just complacent.
Writing about these things in public and putting yourself out there is greatly appreciated. Hats off for that!
However, as far as what's being pursued, it seems more like wantonly trying to ride a hype cycle without strongly questioning the end-to-end value of new software development approaches or vetting their immediate suitability.
Personally, I think it would be more sensible to take a few individuals or a smaller team(s) and do more isolated/skunkworks experimentation and adopt as justified based on what those people report/experience. The smaller group can adjust faster and iterate/advise the larger dev org about the good approaches/techniques/strategies, and avoid more broad damage/chaos for things that aren't that well thought out.
For more conservative AI use cases like adding to code review, writing low stakes PR summaries, or beefing up security checking, a more global, but still not off-the-rails, approach would be the kinds of things that would make more sense to push more broadly.
I find it a bit suspicious that they don't talk about the amount of human intervention required (apart from writing the RFC, which itself can be done by the agent).
I think this can work if you've set up your codebase with proper AGENTS.md with all the guarantees and low level design you expect. I think some human intervention is required here so that you keep the codebase maintainable - you as the human know the domain and future plan well so your addition is valuable.
I also think its necessary to have some garbage collection - scheduled jobs that look at the codebase and think of ways to tighten it and come up with better design.
Lastly, even if one doesn't take away the full factory, I still think that deployments should be fully automated. In the companies I have seen, deployments are still a cognitive burden - one must look at 100 different dashboards and test in staging, look at logs and so on. This is something the agent can do very well and is best automated. Very few orgs have done this!
So I ended up doing something similar:
- Create a Git Hub project board for issues
- Connect Grok to the above
- Use Grok voice mode to take ideas, have Grok refine them with me and then save them as issues
- Created slash commands in OpenCode like /ni (new issue), /do (do an issue), /curr (what is the current issue), /done (self explanatory)
- I generally tell the OpenCode instance to /do <number> and then off it goes
This give me several benefits:
- I can use Grok Voice while walking or driving to develop and test ideas
- I can lose my entire local OpenCode setup but still have relevant data in the issues
- multiple machines can read from GitHub
- I could go even further and have separate user accounts for each of my bots.
Having been both a PM, dev, SRE and manager, this really does feel like managing a team of devs.
That’s very close to my setup. Agree it’s a massive shift that doesn’t even resemble how I used to create products before.
https://jaisenmathai.com/articles/sojourn-for-ios-was-45-one...
and do you ever look at the code afterwards? The 'software factory' gets exponentially nastier the longer it is kept running without someone with actual experience 'shoving it' back in shape ever so often ... and even then.
Sounds like token burn maxing to me. Ie end up with a lot of “not quite” prototypes soooo try again?
don't work while driving, you're putting other people at risk by being distracted
Bro this board is called Hacker News. Let us shove the needle in our prefrontal cortex and use dolphin white matter to hack into SpaceX's network to issue commands to Grok using some ancient sumerian dialect and keep this shit signaling for your Washington Post interview.
Do all that, when you arent operating a vehicle that can kill me and my family
Musk would be so proud. Do you by any chance drive a cybertruck too?
Seriously, don't do this while you're driving. Or even while your car is driving. Its dangerous to be distracted in that situation.
Isn’t this the same as just telling it the goal directly but with more steps?
Yeah, I strongly feel this kind of stuff will just fall away as models get smarter and have a longer and longer viable time horizon per task.
Prompt Engineering, OpenClaw, Ralphing (remember that?) have all fallen already.
There is no problem in software development processes that can't be solved with an additional layer of process.