Skip to main content

If it's not good testing, it's not good regression testing either.

Pick a coin from your pocket, and hold it at arms length. Take a good look. Now take another one, of the same denomination and hold it out at arms length as before. Based on your observations alone - can you say they are the identical?

Lets go a step further. If someone had given you one coin to look at, then exchanged it for another, could you have determined whether they are the same or different coins? Maybe, yes? If the differences had been large enough e.g. one coin was heavily tarnished or scratched, then the different coins would be identifiable. Or if you'd been given the opportunity to examine the coin using magnifying equipment, you probably could of found differences.

But lets assume our only test was a standard set of checks i.e.: viewing at arms length and comparing what we see with our notes/records. It's better than nothing, I would see some differences, some might be important ones. For example if my next coin was blank: I might have suspected an issue with my coin supply, and investigated.

What about my next coin... it is blank on one side. Unfortunately it's not the side I check when I hold it at arms length. So as far as my checks are concerned there has been no regression in the quality of the coins being produced by my pocket. So until I go 'live' and try and spend my coins out in the real world of shopkeepers, I'm none the wiser.

Do you see the flaw in our logic here? If we noticed a degradation in coin quality the testing is good. If the testing does not find an issue, it still must be good, because previously those checks found a different issue. Because I was only performing one test or one set of tests I was blind to issues that I can't see with that one test.

If we'd been testing the coins independently, we probably would of been more critical. We might of thought: sure it looks good in the arm length test, what about the weight: maybe thats wrong. We'd try a number of different tests trying to find an issue. We'd ask other people about coins, learn about their two sided nature and perform tests for it.

But as soon as we enter 'regression testing' mode, we often start to disregard this behaviour and start to mindlessly run the same tests. We avoid exploration, sometimes without noticing. Sometimes people actively avoid exploration during regression testing thinking it's inappropriate. This approach would assume that the test you have been running is some kind of super-observer, capable of helping you to see all problems.

If the system has changed significantly, with the addition or removal of complex behaviours, surely the tests might not also need to adapt? The assumption that the same test will somehow catch a change in functionality, reliability etc is based on the premise that our super-test was testing everything -before- and still is. As testers we know it didn't, doesn't and never will be that super-test. We need to adapt to each new release in an attempt to find new issues. If our tests aren't finding an issue, it's just as possible that the tests are ineffective as it is that the system isn't defective.

Comments

Popular posts from this blog

Can Gen-AI understand Payments?

When it comes to rolling out updates to large complex banking systems, things can get messy quickly. Of course, the holy grail is to have each subsystem work well independently and to do some form of Pact or contract testing – reducing the complex and painful integration work. But nonetheless – at some point you are going to need to see if the dog and the pony can do their show together – and its generally better to do that in a way that doesn’t make millions of pounds of transactions fail – in a highly public manner, in production.  (This post is based on my recent lightning talk at  PyData London ) For the last few years, I’ve worked in the world of high value, real time and cross border payments, And one of the sticking points in bank [software] integration is message generation. A lot of time is spent dreaming up and creating those messages, then maintaining what you have just built. The world of payments runs on messages, these days they are often XML messages – and they ...

What possible use could Gen AI be to me? (Part 1)

There’s a great scene in the Simpsons where the Monorail salesman comes to town and everyone (except Lisa of course) is quickly entranced by Monorail fever… He has an answer for every question and guess what? The Monorail will solve all the problems… somehow. The hype around Generative AI can seem a bit like that, and like Monorail-guy the sales-guy’s assure you Gen AI will solve all your problems - but can be pretty vague on the “how” part of the answer. So I’m going to provide a few short guides into how Generative (& other forms of AI) Artificial Intelligence can help you and your team. I’ll pitch the technical level differently for each one, and we’ll start with something fairly not technical: Custom Chatbots. ChatBots these days have evolved from the crude web sales tools of ten years ago, designed to hoover up leads for the sales team. They can now provide informative answers to questions based on documents or websites. If we take the most famous: Chat GPT 4. If we ignore the...

Text to SWIFT - making data from prose (What possible use could Gen AI be to me? - Part 2)

 As I write this, my dog is grumpily moving around the room pausing intermittently to give me disappointed looks - looks that only my elderly mother could compete with. She (my dog) is annoyed by the robot vacuum cleaner. Its not been run for a while in that room - and its making a noisy foray into dark corners in a valiant effort to cleanse the mess. Its grinding gears and the cloud of dust in its wake is not helping to ease the dogs nerves. The dog's pleading puppy dog eyes & emotions have of course been anthropomorphised - at least a bit - by me (My dog is 7 years old and weighs over 20kg - so has little to fear). That is - I've taken human feelings and mapped them onto my dog. I know she has emotions - but she lacks language - or at least a language that (1) we humans understand, (2) maps to the same phrases or concepts I'm using. But I'm human, That's how I think and how I interact with people and sometimes - machines. Deciphering the problem and representi...