Validation
Premise. Validation asks what evidence is sufficient to deliver a software system into the world.
Software's updateability changes the economics of validation. Engineers can sometimes deliver under uncertainty, observe what happens, and repair what they discover. But an update repairs the software, not necessarily the consequences of its previous behavior. It cannot recover money already lost, make disclosed information private again, or reverse a physical injury.
Validation therefore does not seek certainty before every delivery. It asks what evidence is sufficient for this delivery. That requires several related judgments: what is at stake, which claims require confidence, where those claims exist, what evidence can bear on them, how strong that evidence is, and how much uncertainty engineers can responsibly carry into the world.
What is at stake?¶
Consequence establishes an evidentiary burden. A graphical defect in a game, a TODO application that loses information, and a medical device that delivers an incorrect dose do not demand the same assurance because being wrong does not have the same consequences.
Engineers must consider both the severity of a possible consequence and how directly the software can cause it. Almost any defect can be connected to serious harm through a sufficiently long chain of events, but that does not make every defect safety-critical. An incorrect radiation dose can directly injure a patient; a graphical glitch does not acquire the same significance merely because an annoyed user might subsequently act badly. More severe and more directly caused consequences generally demand stronger evidence.
Reversibility matters as well. Some failures can be repaired cheaply after delivery; others leave consequences that an update cannot undo. Demanding more assurance than the stakes warrant wastes resources and delays useful software; demanding too little transfers unjustified risk into the world.
The question is therefore not can we prove the system correct? It is what uncertainty can we responsibly carry through this delivery?
Where must we establish confidence?¶
Validation begins with claims, not techniques. Specification states what the machine must do under assumptions about its environment; requirements identify outcomes promised in the world. Validation asks what evidence bears on those claims and where that evidence must attach.
Evidence can exist at several scopes. For a payment system required to charge a submitted payment at most once, engineers might obtain local evidence that an idempotency component rejects duplicate identifiers, compositional evidence about interacting retry mechanisms, system evidence at the assembled service boundary, and world evidence about the payment provider's actual semantics.
Smaller scopes make failures cheaper to reproduce, localize, and diagnose. But some properties exist only in composition. Components can each satisfy local latency budgets while their end-to-end path exceeds its system budget; individually correct mechanisms can interact incorrectly; machine-side evidence cannot establish that an environmental assumption actually holds.
Validate a property at the smallest scope capable of establishing it — but no smaller.
The familiar terms unit, integration, system, and acceptance describe roughly these scopes. They tell us where evidence attaches, not how that evidence was obtained.
How can we obtain evidence?¶
Scope and evidence mechanism are independent choices. Engineers can obtain evidence through several mechanisms:
- Dynamic testing observes selected executions. Its limitation is sampling: executions not performed remain unobserved.
- Static analysis reasons about possible behavior without executing it. Its conclusions depend on the abstraction and guarantees of the analysis.
- Human review and inspection contribute semantic knowledge and judgment that may not be encoded mechanically, but human attention is finite.
- Formal methods reason mechanically over explicit models and properties. Their guarantees extend only as far as the model, property, assumptions, and bounds.
- Measurement and experimentation obtain quantitative evidence under stated conditions.
- Operational observation obtains evidence from the deployed system in its actual environment, after exposure to consequences has begun.
No mechanism simply establishes that the software is correct. Each observes or reasons about different aspects of the system and leaves different uncertainty behind.
Dynamic testing introduces another choice: the validation strategy. Executing a program is usually cheap; deciding whether the result is correct can be difficult. The available oracle determines which executions can practically be searched. Example-based testing supplies known expected answers. Property-based testing states what should hold across generated executions; sometimes the useful property relates several executions, such as how an image operation should behave after its input is rotated. Differential testing obtains an oracle by comparing independently developed implementations. Fuzzing uses weak failure signals such as crashes, hangs, and assertion failures to search enormous spaces of unusual inputs.
Choose the strategy for the uncertainty you need to reduce.
How strong is the evidence?¶
Neither the name of a technique nor the quantity of evidence establishes its strength. Engineers must ask what the evidence actually supports.
Important dimensions include coverage — how much relevant behavior was examined; detection power — whether the evidence would expose the failure of interest; representativeness — whether the conditions resemble delivery; scope — whether the evidence attaches where the property exists; independence — whether several pieces of evidence share the same potentially mistaken assumption; and residual uncertainty — what consequential possibilities remain unresolved.
These distinctions matter increasingly as evidence becomes cheap to generate. Ten thousand tests derived from one mistaken interpretation are not ten thousand independent reasons to trust that interpretation. They are one reason, repeated.
Measurement for decision-making¶
Measurement requires the same care. A useful chain is:
Property → Metric → Measurement → Evidence → Judgment
A metric specifies how observations will bear on some aspect of a property. A measurement applies that metric under particular conditions. The resulting value becomes evidence only through the property, metric, and conditions that give it meaning. A measured p99 latency of 420 milliseconds matters because, for example, an architectural model bounded that path at 500 milliseconds under a stated workload.
A disagreement between model and measurement is itself information. The implementation may be defective, an assumption may be false, the model may be incomplete, or the metric may not represent the property we thought it did.
When is the evidence enough?¶
Evidence must ultimately support an action. For quantitative properties, engineers can reason about margin: the separation between expected or observed behavior and an unacceptable boundary. A system measuring 420 milliseconds against a 500-millisecond limit is in a different position from one measuring 499 milliseconds, even though both currently satisfy the requirement. The defensible margin must also account for relevant uncertainty in workloads, measurements, models, and operating conditions.
Not every software failure has a useful numerical margin. Software behavior is often discrete: one unexpected transition, malformed message, or missing authorization check can move execution onto a qualitatively different path. Engineers therefore also build containment into systems. Interfaces, isolation, permissions, transactions, resource limits, timeouts, staged rollout, and rollback do not prove that enclosed software is correct. They limit what an unanticipated failure can affect.
For quantitative uncertainty, ask how much margin remains. For discrete and unanticipated behavior, ask how far failure is permitted to propagate.
The final judgment considers the consequential claims, the strength of the evidence supporting them, remaining margin, containment, consequence, and reversibility. Several actions may follow:
- Deliver when the evidence justifies accepting the remaining uncertainty.
- Gather more evidence when additional information could materially change the decision.
- Change the system when reducing the risk is preferable to gathering more evidence about it.
- Revisit an upstream decision when the evidence exposes a problem with a requirement, specification, architecture, or design.
- Refuse when no available course makes delivery professionally defensible.
A failed validation therefore need not mean test more. Evidence can tell engineers that the implementation should change, that an architectural assumption was wrong, or that a commitment itself should be reconsidered.
Delivery does not end the argument. Operation produces new measurements, incidents, and observations, and changes to the software can invalidate evidence obtained for an earlier realization. Continuing to deliver is therefore another engineering decision under the evidence now available.
Validation asks what must be true, where evidence for that claim must attach, what evidence we have obtained, how strong it is, and whether it is sufficient to act.
Read the expanded treatment: The Software Engineering Handbook, "Validation" →
Materials¶
- Lecture slides — Validation — source (PPTX)
Readings¶
Review and testing in engineering practice
- Winters et al., Software Engineering at Google, Chapter 9, "Code Review," and Chapters 11–14 on testing. These chapters provide the practical foundation for the unit. Chapter 9 treats review as a way to bring another engineer's knowledge and judgment to a change. Chapters 11–14 move from testing strategy through unit tests and test doubles to larger-scale testing. Read them as an account of how engineers obtain evidence at different scopes: what can each practice tell us, what assumptions does that evidence depend on, and what failures remain outside its reach? Full citation: Titus Winters et al., Software Engineering at Google: Lessons Learned from Programming over Time (O'Reilly Media, 2020), chaps. 9 and 11–14.
Human attention is a finite validation mechanism
- Warm, Parasuraman, and Matthews, "Vigilance Requires Hard Mental Work and Is Stressful" (2008). Why human monitoring is itself demanding work, and why assurance cannot scale simply by putting more machine output in front of human reviewers. The authors gather evidence that sustained watching consumes attentional resources, imposes measurable workload, and produces stress, rather than being the passive activity it appears to be. The paper is not about software, but code review is this unit's own vigilance task: a reviewer scanning a diff for defects that are rare, scattered, and usually absent is doing exactly the work these experiments measure. Read it against your own code reviewing, and against the growing volume of machine-generated code arriving for review: which defects can you still rely on a reviewer to catch in the tenth diff of the day, and where are we asking reviewers to serve as indefinitely scalable detectors of rare defects instead of as the source of judgment that makes them valuable? Full citation: Joel S. Warm et al., “Vigilance Requires Hard Mental Work and Is Stressful,” Human Factors 50, no. 3 (2008): 433–41, https://doi.org/10.1518/001872008X312152.
Fuzz testing: let unexpected inputs find the failures
- Miller, Zhang, and Heymann, "The Relevance of Classic Fuzz Testing: Have We Solved This One?" More than thirty years after Miller's original fuzz-testing experiments, the authors return to a remarkably simple question: if we feed programs unexpected inputs, do they still fail? They do. Read this paper both as an introduction to the basic idea of fuzzing and as an unusually long-running engineering experiment. Why can such a simple validation strategy continue to find defects after decades of improvements in languages, tools, and development practices? Full citation: Barton P. Miller et al., “The Relevance of Classic Fuzz Testing: Have We Solved This One?,” IEEE Transactions on Software Engineering 48, no. 6 (2022): 2028–39, https://doi.org/10.1109/TSE.2020.3047766.
Property-based testing: state the property, search for counterexamples
- Ordinary example-based tests choose particular inputs and determine what should happen for each one. Property-based testing asks the engineer to state properties that should hold across a class of inputs, then generates inputs in search of counterexamples. Cockx develops this idea through QuickCheck, including generated inputs, shrinking failures to simpler counterexamples, conditional properties, and different ways of discovering useful properties. Read for the underlying validation strategy rather than the Haskell syntax: what does the engineer still need to specify when test generation becomes cheap? Full citation: Jesper Cockx, “An Introduction to Property-Based Testing with Quickcheck,” 2020, https://jesper.sikanda.be/posts/quickcheck-intro.html.
- The canonical paper introducing QuickCheck. Read it for where the technique came from. Full citation: Koen Claessen and John Hughes, “QuickCheck: A Lightweight Tool for Random Testing of Haskell Programs,” in “Proceedings of the Fifth ACM SIGPLAN International Conference on Functional Programming (ICFP 2000),” special issue, Proceedings of the Fifth ACM SIGPLAN International Conference on Functional Programming (ICFP 2000) (New York, NY), 2000, 268–79.
- A contemporary application: inexpensive generated implementation makes the separation between producing code and stating useful properties about that code especially visible. Full citation: Anthropic, “Finding Bugs across the Python Ecosystem with Claude and Property-Based Testing,” Anthropic, 2026, https://www.anthropic.com/research/property-based-testing.
Only the Cockx introduction is required. Claessen and Hughes and the Anthropic write-up are further reading — neither is necessary for understanding Cockx, but together they show where the technique came from and why it remains relevant in an agentic setting.
Differential testing: create an oracle by comparison
- McKeeman, "Differential Testing for Software" (1998). Testing is harder when engineers can generate inputs but do not know the correct output for each one. McKeeman shows how independently developed implementations can serve as partial oracles for one another: run the same input through each and investigate disagreements. Read this as a general strategy for obtaining evidence when specifying expected answers individually would be prohibitively expensive, not merely as a historical testing technique. Full citation: William M. McKeeman, “Differential Testing for Software,” Digital Technical Journal 10, no. 1 (1998): 100–107.
Model checking: search the behaviors of a model
- Palshikar, "An Introduction to Model Checking" (2004). Model checking makes the relationship among models, properties, and evidence unusually explicit. Engineers describe relevant system behavior in a formal model, state a property the model should satisfy, and use a model checker to explore the modeled behaviors. If the property fails, the checker can produce a counterexample showing how. Read this for both the strength and the limitation of the evidence: exhaustive analysis of a model can provide much stronger evidence than trying selected executions, but it establishes properties of the model engineers actually wrote, under the assumptions that model contains. Full citation: Girish Keshav Palshikar, “An Introduction to Model Checking,” Embedded Systems Programming, 2004.