Crafting a good prompt is what differentiates an expert's output from a beginner's. Of course you need good constraints too. But those good constraints are created precisely through good prompts.
This is an amazing article. The problems it describes are exactly what we found when building a production system that used LLMs to (most of the time) produce reliable results. Extensive tests are necessary, and stakeholders have no idea how their suggestions fail in production. They just see the handful of times they tried something and had it work, not the long tail of cursed results. (“How hard can it be? Why don’t you just…”)
We didn’t get to the point of self-built prompts, as the article suggests, but it’s an intriguing idea.
I hope the future of AI isn't this sort, where companies provide the user/customer with an interface to an AI that can do things for the user... I would much prefer that companies instead provide an interface FOR an AI, and the user brings their own AI which connects to that interface.
In other words, provide my AI with tools, instead of providing me an AI that uses your tools.
That way, my AI can bring all the context it needs, and I can bring all of the settings and knowledge about what I want with me. I don't want a fractured world of tons of AIs i interact with where I have to explain all the fundamental information about what I want and how I work every time.
This also has the benefit of sidestepping the issue the essay is talking about. You provide a consistent tool, and the AI weirdness is not your issue anymore. You don't have to worry about solving for all the weird ways people prompt the AI, or the ways they break.
I don’t think that will happen. It sounds good, and I would like it if things operated that way - but from the company perspective how do they, for example, have a tool call that gives the customer a discount, without it getting used when it shouldn’t?
Companies want ai to replace human customer service decision making, which means it can’t just be an api that an external agent can interact with, because it needs private knowledge of company processes and access to capabilities that are abusable.
But we’re already at the point where if you manage to talk to a human, mostly you end up speaking to someone with no actual power to resolve your issue - so i think basically the future is just going to suck
have a tool call that gives the customer a discount, without it getting used when it shouldn’t?
without comment on the rest of your post, a lot of discounts are just dark patterns, buy 2 for $10 or one for $6 is one of the oldest dark patterns around, and a lot of people fall for it, when they really didn't need that much whatever it was they were buying. In fact there are laws outlawing these practices coming into force for various types of products (alcohol, sugar) .
Other types of discounts like loyalty are more vague personally I think are a minor dark pattern. Negotiated discount on bulk purchases b2b are a different class, but it could be argued that you should only get the discount that the seller saves on logistics.
I'm just pointing this out because its very normalized, 'signup and get 3 months free', 'buy 2 get one free' etc, people dont bat an eyelid at, the same as advertising , but they are all just the earliest versions of dopamine hacking and dark patterns that consensus now is coming around to say probably isn't the best thing when scaled beyond a simple one to one interaction. Why are snapchat dopamine hacking with streaks but your local coffee shop stamping your card for your 10th coffee free different?
from the company perspective how do they, for example, have a tool call that gives the customer a discount, without it getting used when it shouldn’t?
I might be misunderstanding your point but I think you're in agreement with the gp, as in they are arguing for a locked down interface behind which the AI sits, narrowed to use within whatever particular operations the user can access, as opposed to a wide open AI interface with access to any operation (in principle).
So in your example there's no way to ask the AI for any old discount because the interface for doing so is locked down to just the discounts you can access, although the AI may be able to apply a discount you didn't ask for dynamically, if allowed, to give the customer a better experience.
Of course both versions can be implemented securely, it's just probably in general a smaller attack surface if you provide a tighter interface first.
how do they, for example, have a tool call that gives the customer a discount, without it getting used when it shouldn’t?
One possibility is that the company and the user both have AI agents, and that these are able to negotiate with each other. Then I talk to my local customized agent, which knows my preferences, and it explains the required context if the company’s agent does something that won’t make sense to me. But the company’s agent is still privy to the details required to offer discounts, say.
There's no interop between video conferences services, no BYO-video-conferencing app, why would this be any different? It took government regulation to allow telephones other than those provided by the phone company to be plugged into the phone network. Can't control you or usage if they just provide dumb pipes/interfaces that anyone's tool and interact with.
provide the user/customer with an interface to an AI that can do things for the user
This is exactly what MCP is. But the reality is it will likely be about as popular as browser extensions and most normies will avoid.
Simplifying UX at the cost of personal control is inevitable because not everyone wants to think about tool selection and coordination. But maybe we can angle the future toward "tool bundles" that interoperate well or (if we are dreaming) mandate models remain accessible by any harness, which none of the players want but would be best for users and ecosystem development.
I think this is basically the motivation behind MCP servers, so you're in good company.
But I've spoken with many people at companies who've decided to add an in-UI agent to their apps, and I don't think this trend will persist. Absolutely AI/agents will increasingly become a major component of interfaces, but the current common incarnation of "Clippy for X" feels like a kludged bridge between an app that was not designed from first principles for agents (and often was poorly designed for humans) and the impulse to be "AI native".
From my experience, some of the most successful niches for this sort of UI so far have been in apps that were already well suited to it. I'm thinking specifically about analytics/dashboarding/"this is a portal for you to query things easily" software. They already started with a lot of the elements you want: visible provenance of the agents actions via the queries it writes, a malleable interface that you already expect to be customizable and ephemeral, and most importantly, "navigation" that is genuinely difficult for many users (in the sense that "navigating" can mean "querying specific data"). The agent provides a ton of value to users and its actions are intuitive and legible.
But the agents you see right now in a lot of apps that do things like navigate you to the right page by... sending you a link, which you could have clicked from the navbar? Or worse, which hijack your navigation and throw you on a page you're unfamiliar with, and where you have no sense of place or how to make your way to/from? I don't think they're particularly long for this world.
So I actually struggled with this recently (2 months ago). My "vision" was to have customers BYO agent and provide a secure MCP with an in-app (forced) approval gate for sensitive operations. But everytime I described to one of my clients, how they could bring their own ChatGPT, Claude, Gemini, w/e, I got blank stares.
So now I'm building (already have done so mostly) cheap ZDR models into the application. I'm trying to build them in a way that I would want to use them, so a less "I'll be your AI today!" help bubble, but who's kidding... At least they aren't just customer pacifiers that use the help system, but actually can act on behalf of the user.
After which I'll end up building the MCP, but for power users, hopefully leveraging much of the same work.
Also I had to build a human UI, and a mobile UI, and SEO, and... etc.
I think we might be in this interim state that requires us to build ALL of the UIs, to address ALL of the users, and it's a bit painful. If I was "the user" myself, then I would just want a CLI w/ OAUTH with clear documentation, that an agent could use (how is this worse than MCP?). But I'm not my own customer. That being said... I probably WILL build this version to scratch my own itch and bet on the future.
the textual nature of prompts leads us to take the intentional stance towards systems which aren’t conscious, and thus miss the essential nature of their non-meaning
I see LLMs as being capable of making useful distinctions and having a rich action space. They are widely used because their operation is useful, and that can only happen when semantics work well in practice. But useful things that pay for themselves don't need our "essential nature" blessing, they already have persistence by mutual entanglement with us.
No idea how effective this is, but it sure looks a lot more like engineering than most of the 'prompt engineering' things I've seen in the past years. Kudos to the author for writing this concisely without aggrandizing his work.
One issue with this that we ran into is that it costs actual countable money to run the test suite, which is distinct from anything else I’m used to, so the notion that we’d do enough testing to generate a statistically significant gauge of performance - man, I know it’s correct, but I’m not sure my company will survive the process.
Crafting a good prompt is what differentiates an expert's output from a beginner's. Of course you need good constraints too. But those good constraints are created precisely through good prompts.
This is an amazing article. The problems it describes are exactly what we found when building a production system that used LLMs to (most of the time) produce reliable results. Extensive tests are necessary, and stakeholders have no idea how their suggestions fail in production. They just see the handful of times they tried something and had it work, not the long tail of cursed results. (“How hard can it be? Why don’t you just…”)
We didn’t get to the point of self-built prompts, as the article suggests, but it’s an intriguing idea.
I hope the future of AI isn't this sort, where companies provide the user/customer with an interface to an AI that can do things for the user... I would much prefer that companies instead provide an interface FOR an AI, and the user brings their own AI which connects to that interface.
In other words, provide my AI with tools, instead of providing me an AI that uses your tools.
That way, my AI can bring all the context it needs, and I can bring all of the settings and knowledge about what I want with me. I don't want a fractured world of tons of AIs i interact with where I have to explain all the fundamental information about what I want and how I work every time.
This also has the benefit of sidestepping the issue the essay is talking about. You provide a consistent tool, and the AI weirdness is not your issue anymore. You don't have to worry about solving for all the weird ways people prompt the AI, or the ways they break.
I don’t think that will happen. It sounds good, and I would like it if things operated that way - but from the company perspective how do they, for example, have a tool call that gives the customer a discount, without it getting used when it shouldn’t?
Companies want ai to replace human customer service decision making, which means it can’t just be an api that an external agent can interact with, because it needs private knowledge of company processes and access to capabilities that are abusable.
But we’re already at the point where if you manage to talk to a human, mostly you end up speaking to someone with no actual power to resolve your issue - so i think basically the future is just going to suck
They can want anything they like, if customers want to use agents and they don't provide APIs they will lose out.
You prevent abuse of the API by models the same way you prevent abuse by humans: you have server side checks.
without comment on the rest of your post, a lot of discounts are just dark patterns, buy 2 for $10 or one for $6 is one of the oldest dark patterns around, and a lot of people fall for it, when they really didn't need that much whatever it was they were buying. In fact there are laws outlawing these practices coming into force for various types of products (alcohol, sugar) .
Other types of discounts like loyalty are more vague personally I think are a minor dark pattern. Negotiated discount on bulk purchases b2b are a different class, but it could be argued that you should only get the discount that the seller saves on logistics.
I'm just pointing this out because its very normalized, 'signup and get 3 months free', 'buy 2 get one free' etc, people dont bat an eyelid at, the same as advertising , but they are all just the earliest versions of dopamine hacking and dark patterns that consensus now is coming around to say probably isn't the best thing when scaled beyond a simple one to one interaction. Why are snapchat dopamine hacking with streaks but your local coffee shop stamping your card for your 10th coffee free different?
I might be misunderstanding your point but I think you're in agreement with the gp, as in they are arguing for a locked down interface behind which the AI sits, narrowed to use within whatever particular operations the user can access, as opposed to a wide open AI interface with access to any operation (in principle).
So in your example there's no way to ask the AI for any old discount because the interface for doing so is locked down to just the discounts you can access, although the AI may be able to apply a discount you didn't ask for dynamically, if allowed, to give the customer a better experience.
Of course both versions can be implemented securely, it's just probably in general a smaller attack surface if you provide a tighter interface first.
One possibility is that the company and the user both have AI agents, and that these are able to negotiate with each other. Then I talk to my local customized agent, which knows my preferences, and it explains the required context if the company’s agent does something that won’t make sense to me. But the company’s agent is still privy to the details required to offer discounts, say.
Companies will do whatever creates the most lock-in for the user and generates the most profit for themselves.
Ask yourself what's in their best interest as a business? That's probably what they'll do.
There's no interop between video conferences services, no BYO-video-conferencing app, why would this be any different? It took government regulation to allow telephones other than those provided by the phone company to be plugged into the phone network. Can't control you or usage if they just provide dumb pipes/interfaces that anyone's tool and interact with.
This is exactly what MCP is. But the reality is it will likely be about as popular as browser extensions and most normies will avoid.
Simplifying UX at the cost of personal control is inevitable because not everyone wants to think about tool selection and coordination. But maybe we can angle the future toward "tool bundles" that interoperate well or (if we are dreaming) mandate models remain accessible by any harness, which none of the players want but would be best for users and ecosystem development.
I think this is basically the motivation behind MCP servers, so you're in good company.
But I've spoken with many people at companies who've decided to add an in-UI agent to their apps, and I don't think this trend will persist. Absolutely AI/agents will increasingly become a major component of interfaces, but the current common incarnation of "Clippy for X" feels like a kludged bridge between an app that was not designed from first principles for agents (and often was poorly designed for humans) and the impulse to be "AI native".
From my experience, some of the most successful niches for this sort of UI so far have been in apps that were already well suited to it. I'm thinking specifically about analytics/dashboarding/"this is a portal for you to query things easily" software. They already started with a lot of the elements you want: visible provenance of the agents actions via the queries it writes, a malleable interface that you already expect to be customizable and ephemeral, and most importantly, "navigation" that is genuinely difficult for many users (in the sense that "navigating" can mean "querying specific data"). The agent provides a ton of value to users and its actions are intuitive and legible.
But the agents you see right now in a lot of apps that do things like navigate you to the right page by... sending you a link, which you could have clicked from the navbar? Or worse, which hijack your navigation and throw you on a page you're unfamiliar with, and where you have no sense of place or how to make your way to/from? I don't think they're particularly long for this world.
So I actually struggled with this recently (2 months ago). My "vision" was to have customers BYO agent and provide a secure MCP with an in-app (forced) approval gate for sensitive operations. But everytime I described to one of my clients, how they could bring their own ChatGPT, Claude, Gemini, w/e, I got blank stares.
So now I'm building (already have done so mostly) cheap ZDR models into the application. I'm trying to build them in a way that I would want to use them, so a less "I'll be your AI today!" help bubble, but who's kidding... At least they aren't just customer pacifiers that use the help system, but actually can act on behalf of the user.
After which I'll end up building the MCP, but for power users, hopefully leveraging much of the same work.
Also I had to build a human UI, and a mobile UI, and SEO, and... etc.
I think we might be in this interim state that requires us to build ALL of the UIs, to address ALL of the users, and it's a bit painful. If I was "the user" myself, then I would just want a CLI w/ OAUTH with clear documentation, that an agent could use (how is this worse than MCP?). But I'm not my own customer. That being said... I probably WILL build this version to scratch my own itch and bet on the future.
TL;DR: you need to do... essentially supervised learning on your prompts? I mean, if I wanted to do ML, I'd already have been doing it ten years ago.
I see LLMs as being capable of making useful distinctions and having a rich action space. They are widely used because their operation is useful, and that can only happen when semantics work well in practice. But useful things that pay for themselves don't need our "essential nature" blessing, they already have persistence by mutual entanglement with us.
No idea how effective this is, but it sure looks a lot more like engineering than most of the 'prompt engineering' things I've seen in the past years. Kudos to the author for writing this concisely without aggrandizing his work.
One issue with this that we ran into is that it costs actual countable money to run the test suite, which is distinct from anything else I’m used to, so the notion that we’d do enough testing to generate a statistically significant gauge of performance - man, I know it’s correct, but I’m not sure my company will survive the process.
The format of this article makes it almost impossible to read.
I gave up after about 10 "pages".
It’s a talk, not an article.