photo credit: metamorworks/Getty Images

I have worked in business process automation for a third of a century and as such I have had a front-row seat to every new wave of software technology that has come along since Lotus Notes, The Mac and Windows. With each wave I was immersed in, if not drowned by, the enormous hype surrounding “transformational capabilities,” “breakthrough operational efficiency,” and “hockey-stick ROIs.” The accelerating pace of these claims is only overmatched by their ever-increasing bombast.

Still, nothing in my experience comes close to the ongoing claims of a business revolution, nucleated around generative AI. Businesses are making a multi-trillion-dollar bet on what is essentially an advanced gambling platform; applied statistics with a friendly user interface.

 

Band Aids for Arterial Bleeds

I recently came across some posts here on LinkedIn gushing with praise over Google’s new Generative AI training course, featuring a whitepaper on prompt engineering.  This whitepaper, authored by Lee Boonstra and some of his fellow Googlers, lays out all of the detailed steps for how you write “good” generative AI prompts, and outlines, if unintentionally, some of the fundamental flaws of LLMs.  Most mentions of this paper seem to miss the several such Google publications on this topic over the last few years, but nothing sells quite so effectively as the appearance of new-ness.

 

While it is clear that the authors intended for this paper to benefit users who are trying to get positive outcomes from their LLM investments, it also makes clear the need to deeply question whether or not we should be taking this technology seriously in the first place. Prompt engineering is itself an admission that these platforms should not be taken too seriously as business tool.  As Boonstra states in one of his recent posts one of his recent posts,

“Given how much prompt outputs can change across different models, sampling settings, and even different versions of the same model, it’s super important to document everything. You might get a response with slightly different wording or formatting, even with the exact same prompt, so keeping good records is key for future work.”

The last time I checked in with my friends in Compliance or Legal, the idea that a business system should be expected to provide different answers to the same question, without a clear understanding as to why there was a difference, was fairly disqualifying. And while the suggestion that users ‘keep better records’ is a tip of the hat to the requirement of repeatability and traceability in business processes, it sure sounds like the amount of increased effort, and records produced, may call into question any notions of a positive return on investment.

 

Your Results May Vary

Again, Boonstra, et. al., implicitly acknowledge this challenge in a further observation,

“Remember, prompt engineering is all about continuous improvement. You’ll need to create and test different prompts, analyze and document the results, tweak your prompts based on how the model performs, and keep experimenting until you get the results you want. If you change the model or its configuration, go back and test your old prompts again. [emphasis added] This iterative process is key to refining and optimizing your prompts for the best possible performance.”

Now, having to re-test old code or old business rules is nothing new; this is and always will be a fundamental part of designing, building and maintaining business information systems.  However, what prompt engineering reveals quite clearly is that the behavior of these systems is inherently unpredictable.  Ask an LLM a certain question one day, then ask another day or in a slightly different manner, and you may expect to get a wildly-different response.  This isn’t inherently bad.  Different questions often demand different answers. 

What is a challenge is this: there’s no real way to understand WHY the responses are different. What was the source of the difference in responses? Better prompting? New ‘learning’ by the LLM? Sun spots? Slight fluctuations in the gravitational pull of Jupiter? In deterministic systems, like rules-based software, we test all such possible variations to create predictable systems with predictable outcomes. This is otherwise known as governance.

LLMs are inherently probabilistic; they’re guessing engines. This can be valuable when one is trying to come up with guesses to answers, but it is a poor choice when one is trying to come up with answers to guesses. Guessing at guesses is the antithesis of governance, and for a broad and deep array of business processes, such compounding guess work is simply unacceptable.   With all business automation tools, use case selection is critical for operational and financial success.  With generative AI, use case selection is the difference between positive ROI and a class action lawsuit.

 

Utensils and Utility

Much of this discussion around prompt engineering seems like a crutch. Reading a 68-page guide on how to prop up an inherently-unpredictable business systems feels a lot like having a 68-page guide on how to eat broth with a fork. It might be possible, and your results may vary based upon your techniques.  But one must wonder: wouldn’t we be better served by simply using a spoon instead? A fork is a wonderful utensil with very high utility, but only when eating solids.  Using a fork for eating liquids sounds an awful lot like the buffoonery we used to watch on after-school-TV as kids; or like the buffoonery that “adults” now watch on platforms like TikTok. It does not belong in, and we cannot afford it to take hold in, serious places like the workplace.

Generative AI can and will provide business value. But, as guidance from experts such as Boonstra clearly shows, picking the right tool for the job has never been more critical or determinative of success.  Please choose wisely.


About the Author: IRPA AI Senior Analyst, Chris Surdak


Chris Surdak is a Senior IRPA AI Advisor and was formerly White House Chief Transformation officer, Automation & AI Practice Lead at EY & Executive Partner for Digital Transformation at Gartner. He’s an engineer, futurist, transformation executive and best-selling author, with over 30 years’ experience in technology development and deployment, digital transformation, blockchain, data and analytics and AI & intelligent automation.


CLICK HERE TO SCHEDULE AN ANALYST CHAT

Links:


Originally posted on 2025-04-11 in the IRPA AI Network — Announcements & Updates