Why We Measured AI Instead of Trusting It

Generating test cases with AI was easy. Working out whether they added value was much harder.

This blog series grew out of a conference presentation exploring how AI-generated UAT test cases could be evaluated through evidence rather than assumptions.

Where we started

AI is everywhere. We’re told AI will revolutionise everything bringing unparalleled efficiency and savings! Happy face!  

Or possibly that it will bring about mass professional unemployment, neo-feudalism and a sea of AI Slop. Sad face!

Either way, my organisation, like may others, is exploring how we can use AI in pretty much every context, including in software testing of enterprise IT systems which is my area.

And we are far from alone in this; looking at the testing conference circuit this year AI was clearly the biggest topic. I should know I was one of the people giving a talk on it.

In our current paradigm shift there is a lot of pressure on people to just “use AI”… Use AI to do what exactly, to solve what problem? Who cares, just use it!

And AI is deceptively easy to start using, be it with specific testing tools sold by specialist vendors or just generic LLMs.  Churning out tests is easy with AI but how useful is it really?

I could easily have just set up an AI pipeline to see how it operates and how I feel about it then told my managers “Yes, we’re doing AI now!” But that’s not exactly a measured or well thought out approach.

I’m a tester, I’m interested in getting solid metrics on how well systems work, applying judgement. That is fundamentally my job and if you’re a test professional it should be yours as well.

Testers shouldn’t run on vibes, they should run on evidence.

I decided to do what I do best and gather evidence to see how effective AI test creation was for me.

What began as a comparison of human and AI-created tests would eventually raise wider questions about quality, human judgement and whether a project is ready to benefit from AI in the first place.

Setting up a study

I run Business Acceptance Testing and User Acceptance Testing on Enterprise IT systems at the BBC. My job involves checking that the various outsourced systems we procure work as intended and do what the end user wants. This frames where I came from in setting up my study; the type of systems I test, the developmental stage they’re at, my resources and so on. 

What works for me might not be what works for you because you’re testing different systems, using different tooling, with different stakeholders and different priorities. My results in terms of value delivered will be different to your results.

Your inputs are different from mine, your results may be different from mine, but the method for evaluating whether AI assisted test creation will work in your environment is transferable.

Setting up a study

The Proof of Concept already existed: AI could generate test cases. This has been demonstrated many times by many people.

But that wasn’t the question I wanted to answer. I wanted to know what value those tests added in the environment where my team worked. What I needed was a Proof of Value study.

The difficult part was not persuading AI to write test cases. It was working out how to recognise whether those test cases had value.

What followed challenged some of my assumptions, not only about what AI could produce, but about requirements, source material, human judgement and what “good” testing means.

Designing the study

With that in mind, how do you design a study that works in your own circumstances? I recommend the following process:

  1. Understand your starting point
  2. Decide what value means for you and your ogranisation
  3. Select key metrics based on those values
  4. Set up a fair system of comparison
  5. Stabilise your AI approach before you begin measurement
  6. Collect and evaluate the evidence from human and AI generated tests

Understanding your starting point

To start with I did some basic analysis of where my team and I were. I ran a SWOT analysis, where I analysed what our Strengths, Weaknesses, Opportunities and Threats were when it came to AI.

I then ran a resource audit where I looked at what was available to work with and what wasn’t available

  • team composition and skills
  • capacity and workload
  • available tools and technologies
  • training needs
  • constraints and opportunities

Deciding what we valued

Different testing teams are trying to solve different problems. A team under intense delivery pressure might care most about reducing test design effort. A highly regulated environment might prioritise traceability and auditability. Another team might be struggling with test coverage and want to know whether AI helps identify missed scenarios.

Looking at what we valued for us it came down to three broad questions:

  • Does AI help us create tests more efficiently?
  • Does it create tests of sufficient quality?
  • Can it realistically fit into the way we actually work?

Those questions became the foundation for the metrics we selected

Defining key metrics

Once we had defined our values, we could consider the measures that might represent them.

Working with stakeholders and the wider BBC testing community, we generated more than 20 potential metrics. Collecting data for every one of them would have created considerable overhead, so we needed to identify the measures that best represented what we valued.

Setting up a fair system for comparison

Having established what we valued and wanted to measure, we needed a fair way to compare human-created and AI-created tests. That was harder than it sounds.

Once a tester knows that a defect exists, they cannot simply unknow it and independently rediscover it. A perfectly controlled comparison would therefore require different testers working in parallel.

I run on a shoestring, so spare people available to repeat the same work was not going to be easy to find. Our study design had to be credible, but it also had to reflect the resources available in the real world.

Stabilising the AI approach

There was another fairness issue. An experienced tester brought years of context, judgement and learned technique. Our newly established AI-assisted process did not.

We therefore needed to give the AI approach an opportunity to stabilise before formal measurement began. Otherwise, we would have been measuring the immaturity of our first attempt rather than the potential value of the approach.

Collecting and analysing the evidence

Using several measures allowed room for nuance. The outcome would not be reduced to a simple claim that either humans or AI were “better”.

This was not simply a comparison between two sets of test cases. It also raised questions about what experienced testers contribute that is difficult to encode in a prompt or provide through a source document.

What happened next

We had established the question, assessed our starting position and identified the measures that mattered. We had also found a real project that could give us sufficient scale to compare human-created and AI-generated tests.

The eventual results would be specific to our environment. The method for reaching them is something other teams can reuse.

The next step was turning an abstract claim such as “AI adds value” into something we could observe and measure.


Next in the series

In the next post, I’ll explain how we defined value, reduced more than 20 possible measures to a practical set, and discovered why the most obvious measures were not always the most useful.

SNAFU, FUBAR and Clusterfudge

Software projects don’t always deliver on time or deliver what’s wanted; according to the Standish Group only 31% of software projects are successful, and of these successful projects only 46% return high value…

Fundamentally testing as an activity is about finding and preventing problems that will affect the quality of what you’re delivering. This isn’t always about plain bug hunting or automated test suites, it’s about smoothly managing a project because when projects slip, as they do, quality suffers.

Generally we use RAG statuses, Red, Amber, Green, to measure how a project is running. But what do these mean? Red, it’s late, Amber it might be late, Green it’s going OK? Or something like that; to be honest I’m not entirely sure because everyone’s take is different. Of course, RAG statuses should be clearly defined at project inception however this isn’t always the case. As I frequently say, if everything that should happen did happen I wouldn’t have had a twenty five year career in testing.

So completely tongue in cheek I’d like to suggest a different taxonomy on how we look at project status based on a hierarchy of military terms. I’ll also suggest what test management should be doing for SNAFU and FUBAR projects. When it comes to a clusterfudge you’re on your own!

SNAFU (Situation Normal All Fouled Up)

This gentleman is Private SNAFU. He featured in a series of cartoons made from 1943 to 1945 by the US Army Airforce First Motion Picture Unit. The cartoons showed what happened when service personnel didn’t follow regulations .The majority of the cartoons were written by one Theodor Geisel who you might of heard of by his more well known writing name, Dr Suess.

Organisations are made of people, and because of this they can be surprisingly functionally messy. When I was young I used to love blooper programmes on TV as they sowed that behind the slick and polished exterior of the TV screen things went wrong on a regular basis.

This the base state of projects. Things don’t run smoothly and there are a series of low level issues to overcome in order to deliver on time.

What should test management be doing?

This is where test management should be making sure that potential project issues have been anticipated and mitigation measures are ready. For me a key element in tackling SNAFU is planning, particularly in keeping up to date RAID logs.

However these need to revisited regularly; as the old adage goes “No plan survives contact with the enemy”. Any plan, including your RAID logs, that is created in a changing environment must routinely adjusted and assessed.

FUBAR (Fouled Up Beyond All Recognition)

It appears this phrase came from the USMC in World War 2 and then was popularised, if that is the correct word, during the US Vietnam conflict. Which when all is said done was pretty FUBARed from the American perspective.

In my taxonomy FUBAR is for projects that are going off the rails. In conventional RAG status these would either be in the Amber or Red.

What should test management be doing?

When trying to turn around a project that is FUBAR realism is key. Testing is unlikely to be the reason it’s shifting right testing but it will be a key part of getting it back on track. As a tester or test manager you need to provide accurate estimates, not just what people what to hear. It’s a regular occurrence on FUBAR projects to be asked to trim testing estimates for activity occurring at the end, such as UAT or SIT. Fine, you could do that but if you do you’ll just be revisiting timelines later.

Be honest with your stakeholders and don’t be afraid to ask for help. Ultimately our job should be to ensure our organisations survive and thrive so don’t be afraid of talking to project sponsors, keeping the sunk cost fallacy in mind.

Of course, it may be that your project sponsors are so deeply committed to the project that they don’t want to hear what you have to say. That brings us to the next status.

The Clusterfudge; Red and beyond

The “clusterfudge” (please note I’ve sanitised the wording) is where things are now seriously wrong. The photo above is by Hubert van Es and shows the last US helicopter to leave Vietnam in 1975, truly a clusterfudge of a situation from the US perspective.

The etymology of the word is somewhat obscure, but it came about during the 1960s and was popularised during the US Vietnam conflict. One attributed is to beat poet Ed Sanders, however my favoured origin and the one that fits with this talks is that is stems from the Oak Leaf Clusters worn by Majors and Lieutenant Colonels in the US forces. A cluster “fudge” came about when these officers got involved in a FUBAR situation and made it worse.

Stanford Prof Bob Sutton describes a clusterfudge as being made up of a mix of illusion, impatience and incompetence. Illusion, when a decision maker believes a goal is easier than it really is. A current high profile example might be Putin believing that Russian troops could take Kyiv in three days. Impatience, trying to release in a hurry or within false time frames. Incompetence, where leaders don’t have technical competence and aren’t willing to listen to experts.

Hopefully your projects will be more SNAFU than FUBAR; chaos and churn are part of life, it’s our job as test professionals to look for problems before they happen and have contingencies in place to deal with them.

Shifting Quality Left with Fagan Inspections

At recent test conferences it seems one of the phrases of the moment is “Shifting Quality Left”. If you test as early as possible bugs are found when they are easier and cheaper to fix. Ideally catch them those you can before coding even starts!

Possibly because the industry has pivoted towards Agile one method that seems to have fallen somewhat out of favour in recent years is Fagan Inspections. If you’ve not come across the term before I’m going to describe it here, it may be something you want to try. A Fagan Inspection is a document review process to make sure that the objects you’re working with are defect free. I say document review, it’s also applicable to code although personally I’ve not used it in that context. This makes it a type of static analysis, testing without actually running the software being created. It’s called a Fagan Inspection as it was devised by Michel Fagan at IBM in 1976. Personally I came across this while working at Sony Ericsson where known as a Granska Bra (“well review”). Sony Ericsson had taken the concept from SAAB and applied it to mobile phone development.

A Fagan Inspection is a formal review process where objects (documents or code) are examined team of people to check to see if they’re ready for use. When I was using this process at Sony Ericsson we estimated that we had a ten fold return on the time we invested in it. So for every hour each person spent reviewing and preparing for reviews, we would save ten hours of time down the line.

How does it work? Well, when the object owner thinks what they’re working on is ready they take it to a moderator and the two of them have a pre-inspection run through just to make sure things really are ready to review. The moderator then sets up a panel of relevant experts to examine the object. For example, when using this at Sony Ericsson we would generally have a developer, a UI designer, a tester and an architect. These reviewers should run through the object from their domain perspective before the meeting.

During the meeting a presenter presents the object. This could be the object owner or it could be the moderator. The reviewers then step through the object and comment where they feel the need to.

There are two dangers at this point and it’s the moderators job to tackle them. Firstly, the moderator has to ensure that the review is a fair review and that the owner of the thing being inspected doesn’t feel like they’re being picked on. Secondly, it’s easy for people to drift onto finding solutions at this point “Ah, we could try doing this instead!” .That’s not the point of the meeting and if you let it start happening the meeting will soon be derailed. A two minute rule works well for this, rigorously enforced by the moderator. People can arrange follow up meetings to work on solutions if they need to.

At the end of the meeting the an exit decision is taken. Does the object need reworking and a subsequent re-review or can it be passed (perhaps with minor modifications with that won’t need re-review).

Adding this process can be seen as a drag, an extra step for people who just want to get on and code. However, ensuring our inputs are actually fit for purpose and clearing up misunderstandings before work is created based on those misunderstandings pays massive dividends.

Why we blog

The aim of this blog is to talk about software testing. It’s a broad area so there’s a lot to talk about from tools and techniques, to methodology and into the philosophy of test,

The inspiration for this blog came from a talk at the UK National Software Testing Conference 2021. As someone with a fancy job title at a well known tech leader I occasionally get invitations to do things like talk at a conference. As it happens the invite for the NSTC talk dropped into my inbox soon after I’d read an interesting article about levels of project failure so I said “Yes, why not?”

After that a blog seemed to be the next thing…