Is software engineering the lone unicorn of AI use cases?
To transfer the success in agentic programming to other disciplines, we need to understand how automated verification systems and the use of an intermediary formal language makes this possible in the first place.
Software engineering is the poster child of agentic AI automation - other disciplines are way behind. But is it even possible to close this gap, or are we hitting capability limits of today's AI paradigm? Learn how verification systems enable the successful adoption in software engineering, what the constraints are, and how it can be transferred to other fields.
The productivity of software engineers seems to have skyrocketed, with sites like Github.com reporting almost exponential increases in the amount of code submitted and Google claiming that 75% of their new code is being written by AI.
Recent studies on broader AI adoption, such as the 2026 State of AI report from McKinsey, make it clear that the gap with other disciplines remains large. In the screenshot from the McKinsey website below, a clear sweet spot for software engineering in technology companies can be identified in deep blue color:

While this much lower degree of adoption may be attributed to many different causes - leadership, lack of skills, etc. - there are also problems inherent in the technology itself that can stop adoption cold. LLMs don't provide deterministic output, we don't understand how they arrive at their results, they are vulnerable to bias and exploits, and so on. In two words: unreliable, and unsafe.
However, despite these issues, software engineering does not seem to suffer as much from these problems as do traditional business functions such as Finance or Operations.
In order to be able to transfer the success of AI in software engineering to other disciplines, we need to look at what it is that makes software such a distinctly good field for making broad use of AI, and how exactly it mitigates the shortcomings of the technology.
Semi-autonomous software development - the Bun example
A typical programming project still requires substantial amounts of human guidance and oversight, but it is interesting to look at an extreme example, published by the team at Bun.
Bun allows code to run on a server - do things like show you a webpage, store and retrieve data, call other web services, and so on. Bun also manages dependencies on other software, and helps run tests. Most software projects are like a stacked house of cards, building on pre-existing functionality provided by others. Already a popular tool on its own, it has been acquired by Anthropic, and popularity has since skyrocketed to 22M monthly downloads because it became part of their Claude Code product.
Bun itself was originally written in a programming language called Zig, but at some point started running into more and more issues connected with how Zig manages the memory available to Bun. The team at Bun thus decided to do something that basically no one ever does (and smart people - before the AI era - have argued why): rewrite the software from scratch in a different programming language.
But of course they didn't do it manually: this is where the AI came in.
Using state of the art AI models, the rewrite was reportedly completed in an astonishing 11 days, burning about 165k USD in tokens (which were free for the team, given they work at Anthropic) - both a timeline und budget that would be absolutely unthinkable for a human team (the estimate runs at 3 persons for 1 year).
Verification systems create agent autonomy
While the quality of the result may be a contested topic among some project insiders, and of course the majority of agent supported coding is by far not this autonomous, it's undeniably a huge feat to get a rewrite done in such a short time that actually works in production and improves on the previous version!
The key learning from their report is that the team had created a thorough and automated verification system for the program code. The system was able to automatically test the functionality, identify memory and security issues, and automatically test integration behaviour with other software. And this is not a unique pattern - if you talk to engineers working deeply with AI based coding, you will find that indeed building such systems to verify the output of an AI process is what they spend most of their time on, in an effort to create as much autonomy for their AI coding agents as possible with minimal human oversight.
Now, you can of course say that it's cheating when you already have a working product that you can let the AI compare against. But software engineering has a deep tradition of building verification. On most projects, we usually start with the test cases, and only then proceed to create the actual code. We test units of code, integrations, we load test, we security test, and most importantly, all in an automated way by writing (even more) code.
Now, given a sufficiently elaborate verification system, any change can be automatically tested against this system to see if it still works as intended or whether new functionality has been correctly implemented - no matter whether a human, monkey, or AI agent created the code.
It is the existence of these verification systems is what allows humans to step back and delegate lists of changes to autonomous coding agents, which will design, coordinate and implement the changes while testing against it, because coding agents can then use error messages in log files, screen shots, UI automation, knowledge bases and other tools to try to self-heal and recover from error, proposing and trying new solutions until the task has been completed successfully.
The asymmetry in AI
In a way, automated verification is the key to benefit from what I like to call the asymmetry of AI impact. LLM-based AI is really bad for tasks where you need 100% accuracy so the automation pays off. When the cost (or risk) of one or a few failures outweighs the benefit of many good runs, and you can't detect the failure without a human, you will be struggling to fix the shortcomings with workarounds and fixes, things people call guardrails. Think invoices, flight safety, etc.
On the other hand, when you have a scenario where you are fine with many failed runs as long as you get at least one good run, and your current limit is the cost of performing a single run, you have a great scenario. With a solid test suite, it doesn't matter (much) whether the LLM needs just one pass or ten passes to arrive at a working code example.
Results from program code, not LLM answers
Everything we've discussed up to now builds on an important, but often overlooked fact when we compare AI use in coding to that in other disciplines:
We're not building stuff where AI is handling the actual end user process.
Instead, we use AI to generate output in a formal, controlled vocabulary - program code - that can be audited and run in a reproducible way.
This is a crucial difference to applications where we rely on directly providing the output of an LLM at runtime - such as an agent employed to create content in a powerpoint, do a financial reconciliation, or filter applicants in a recruiting process.
The availability of high-quality training data
Software also differs in another aspect. Training sets that go into LLM pre-training include huge amounts of text, Reddit posts, etc. - but LLMs are also trained on huge amounts of programming code available on the web, directly ingesting good and working examples of both the expected output and the test suites. Because code repositories also include the change history of the code base, we can even train on how a software project typically evolves, and based on what human inputs (documented in the tickets, comments and check-in messages). This is made possible by the many public repositories, especially open source software, that contain both code and test cases, and sites such as Stackoverflow.com.
This amount of structured public data currently simply isn't available for any other discipline. Also, these often do not use the same class of structured intermediary language, nor computer readable verification systems.
Making the transfer? Adopting software engineering AI practices in other disciplines
In order to learn and transfer from software engineering, we can take a step back and look at our work from a different perspective. Where can we move away from direct LLM output to an intermediary language? Where can we build verification systems? And how can we generate a similar abundance of training data, possibly private to us as an advantage?
Some areas have already adopted this approach. The way Pharma companies set up their agentic automation for research in processes like protein folding is by creating a series of small program libraries, written in Python or similar, that allow an agent to start simulations, observe outputs, and perform verifications against certain key measures as well as failure modes. These program libraries are given as tools to long running agentic processes, which can start experiments and use the tools available to verify results and test different recovery strategies in case of failure. So instead of having an agent directly manipulate outputs (such as protein structures), this is done by deterministic code that an agent can invoke. Instead of relying on human feedback on each run, a verification system is used to check whether simulations run correctly and their output is useful.
How can we transfer to other areas? For Finance workflows, it may help to focus on having an LLM create the required formulae and automations in Excel, or maybe even as a Python Notebook, as opposed to letting the LLM manipulate the data or work in the actual spreadsheets. Building up clearly labeled test sets and constraints for different scenarios with expected inputs and outputs helps for automated verification of changes, especially to avoid regressions, i.e. breaking of previously working functionality.
Classification and routing tasks performed by an LLM e.g. on support tickets or incoming documents, may in many cases be delegated to a deterministic algorithm created with the help of an LLM, such as a decision tree. This improves auditing and explainability, and reduces the direct use of LLM classification to edge cases.
In engineering, we're often unable to directly generate the actual target data with an LLM, because it may be binary - such as the files for a Blender 3D project. But what currently may seem like a "computer use" scenario - i.e. the agent operating the computer and programs like Blender just like a human via keyboard and mouse, as seen e.g. in the launch video for OpenAI Astra - has already been transformed into a coding task: the Blender application exposes a programming interface that uses Python for all its functions, and LLMs can be trained on pairs of instructions with resulting Python scripts.
What we're missing here, though, is the long historic baseline of examples that we have for traditional software engineering, as users are just starting to transition from manually using such software to using it via code-based automations and tuning the results.
While the first learning is thus that we can try to apply a coding approach - with an intermediary, deterministic language as the target of LLM output - to many more disciplines, the follow-up learning is to vastly improve at generating training data to teach AI systems the expected inputs and outputs in our fields.
This can already be helpful without adopting the coding approach: generating detailed logs of how business processes are performed, from input to output, across different IT systems and even manual steps, provides an invaluable resource to anyone looking to automate these processes or onboard new team members. Solutions for this exist; they can be a side effects of observability tools, or dedicated projects in process mining. For LLM interactions, dedicated tools exist to record the interaction traces.
This data, along with screen recordings and explanations, can be both used by LLMs to create process automations, as well as by data scientists to better train LLMs for the tasks themselves. Incidentally, such training data can also be a core competitive asset, if it allows an organisation to automate steps that others cannot.
Re-thinking how we train our teams on AI
What we require to be able to execute on this idea is a new perspective on what it means to educate our teams for adopting AI. It feels like past initiatives have mainly focused on three things: adhering to regulatory requirements, and learning to interact with an LLM through prompts, and acquiring knowledge on existing off the shelf tools.
The result is superficial, ritualistic knowledge on how to improve the chances of the LLM tool producing the right output - be it text, Excel, PowerPoint, or image - with the newest prompt tricks or downloaded agent skill.
This builds on a tradition of "computer literacy" education that prioritized how to operate different business software applications (Word, Excel) over an actual understanding of the computing system and how to instruct it to process data.
The uncomfortable truth is that we urgently need to step out of this paradigm and embrace a new one. In a world where we automate complex workflows, it becomes an essential skill to know how to break it down into smaller steps, model the expected inputs and outputs of a process step, and understand the failure modes and process flow patterns. Grasping whether a bigger chunk of processing is done by an LLM - burning tokens - or by code generated by an LLM - slightly heating your laptop - can make a big difference. An understanding of how a piece of computing technology - program code, an LLM, an agent harness - operates, what its constraints are - become basic building blocks of knowledge for those driving change.
With coding agents taking away the grunt work of programming, creating code based automations becomes accessible to many. For this to scale, we need basic education of computing systems, data, and the required infrastructure. The future may belong to those acquiring this knowledge and applying it to transform their discipline.