Let's say it kind of does what you want, but not exactly. I now have a problem. How do I fix what I told it to do so it does exactly what I want? I do not know. Do I need to be more persuasive? Should I be argumentative? Do I need to be rude? Was I ambiguous in a way that I did not understand as I wrote it?
Yes! This is one of the major challenges I still have. It's easy to say "fix model errors in AGENTS.md / the system prompt." But finding a way to prove the effect of that change is the hard part.
It's so frustrating when a model says "no, your guidance was fine, I just didn't follow it, I'll do better next time..." It just makes me want to yell "No! You won't do better next time! This is the problem!"
In that respect, working with an LLM is much more like working with another person than working with traditional software.
Yes, but if you work with another person, at least they'll have a chance of remembering suggestions you make.
The simple answer is "Have the AI do a refinement -> eval -> refinement loop." That pushes the problem somewhere else, into the "what goes in the eval" question ("and make sure you don't overfit" -- something AI-driven prompt refinements are not good at).
Yes! This is one of the major challenges I still have. It's easy to say "fix model errors in AGENTS.md / the system prompt." But finding a way to prove the effect of that change is the hard part.
It's so frustrating when a model says "no, your guidance was fine, I just didn't follow it, I'll do better next time..." It just makes me want to yell "No! You won't do better next time! This is the problem!"
Yes, but if you work with another person, at least they'll have a chance of remembering suggestions you make.
The simple answer is "Have the AI do a refinement -> eval -> refinement loop." That pushes the problem somewhere else, into the "what goes in the eval" question ("and make sure you don't overfit" -- something AI-driven prompt refinements are not good at).