Black Box Testing: Five Techniques That Cut Case Counts
Black box testing derives cases from the specification. Equivalence partitioning, boundary value analysis, decision tables, state transition testing, and error guessing, with worked examples.
Yuvan Sundrani · 12 min read
autosana.ai

TL;DR
- Black box testing derives cases from the specification and observable behavior, with the implementation treated as opaque.
- Five techniques do most of the work: equivalence partitioning, boundary value analysis, decision tables, state transition testing, and error guessing.
- Applied properly these techniques reduce case counts while increasing defect detection, because they replace arbitrary inputs with representative ones.
- Autosana executes black box flows described in plain language on real iOS, Android, and web devices, which is the same vantage point a user occupies.
A field accepts ages between 18 and 65. A tester enters 25, 40, and 50, runs three cases, and reports the field working correctly.
Those three values sit inside one equivalence class, so the three cases tested one thing. The values that find defects are 17, 18, 65, and 66, because defects cluster at boundaries where a developer chose between less-than and less-than-or-equal.
Black box testing is the discipline of picking those four values instead of the other three. This guide covers the five black box testing techniques that do it systematically.
What Black Box Testing Is
Black box testing derives test cases from requirements and observable behavior, with no visibility into source code. You supply inputs, observe outputs, and compare against what the specification promised.
The ISTQB Foundation Level syllabus classifies these as specification-based techniques. The ISO 29119 standard covers the same family under test design techniques.
Two consequences follow from that opacity, and both shape how black box testing gets used. Cases stay valid through refactoring, since they depend on behavior rather than structure. And coverage becomes hard to measure directly, because you have no view of which code paths your inputs reached.
Security work uses the same vantage point. NIST SP 800-115 describes black box penetration testing as assessment performed with the tester holding no internal knowledge, which mirrors an external attacker.
Black Box and White Box in One Table
| Dimension | Black Box | White Box |
|---|---|---|
| Visibility | Specification and behavior | Source code |
| Cases derived from | Requirements | Control flow and structure |
| Survives refactoring | Yes | Frequently rewritten |
| Coverage measurable | Indirectly | Directly, as a percentage |
| Finds | Requirement gaps, wrong outputs | Unreachable paths, uncovered branches |
The two answer different questions and both are needed. What follows covers black box testing in depth.
The Five Techniques
Equivalence partitioning
Divide the input space into groups whose members the system should treat identically, then test one value from each group.
For the age field, three partitions exist: below 18, between 18 and 65, above 65. One value from each gives three cases covering the entire input range for that dimension.
The economy is substantial, and it is the main reason black box testing scales. A field accepting any integer has billions of possible inputs and three meaningful ones.
Partition invalid inputs too. Non-numeric entries, empty submissions, and negative numbers each form their own class, and each deserves a case.
Boundary value analysis
Defects cluster at the edges of partitions, because that is where comparison operators get chosen.
For each boundary, test the value just below, the value itself, and the value just above. Age gives you 17, 18, 19 at the lower edge and 64, 65, 66 at the upper edge.
A two-value variant tests only the boundary and the value immediately outside it, which halves the case count while catching most operator errors.
This technique finds more defects per case than any other in black box testing, and it applies anywhere a range exists: string lengths, file sizes, dates, quantities, and monetary thresholds.
Decision table testing
When output depends on several conditions combining, a table makes the combinations explicit and exposes the ones left unspecified.
Consider a discount rule with three conditions: the customer holds a membership, the order exceeds 5000 rupees, and a promotional code applies.
| Membership | Over 5000 | Promo Code | Expected Discount |
|---|---|---|---|
| Yes | Yes | Yes | 25 percent |
| Yes | Yes | No | 15 percent |
| Yes | No | Yes | 10 percent |
| Yes | No | No | 5 percent |
| No | Yes | Yes | 20 percent |
| No | Yes | No | 10 percent |
| No | No | Yes | 10 percent |
| No | No | No | 0 percent |
Three binary conditions produce eight rows. Building the table usually surfaces a row the requirements left undefined, and finding that gap before implementation is where the technique pays.
State transition testing
Use this where the system holds state and valid actions depend on which state it occupies.
An order moves through created, paid, shipped, delivered, and cancelled. Legal transitions run created to paid, paid to shipped, shipped to delivered. Cancellation is legal from created and paid.
Cases cover every legal transition, then attempt illegal ones: cancelling a delivered order, shipping an unpaid order, paying twice. The illegal attempts are where the defects live, since developers implement the happy path first.
Mobile adds transitions the specification rarely mentions: backgrounding mid-flow, termination under memory pressure, and returning after a token expired.
Error guessing
Experience-based case design, drawing on defects seen previously in similar systems.
The recurring candidates: empty submissions, leading and trailing whitespace, maximum-length strings, zero, negative numbers, duplicate submissions, unicode and emoji in name fields, timezone boundaries at midnight, and leap day.
This resists formalization, which is exactly why it complements the other four. Keep a team list of guesses that found something, and it becomes an institutional asset rather than an individual skill.
Applying These to Mobile
Black box testing techniques carry over to mobile, with three additions.
Interrupt states extend state transition testing. An incoming call, a backgrounded app, a low-memory termination, and connectivity loss are all state changes with legal and illegal follow-on actions.
Permission responses form equivalence classes. Granted, denied, and denied-permanently each represent a class, and the third is the one usually left uncovered.
Device and OS combinations multiply every case. That makes partitioning discipline more valuable on mobile than on web, since an undisciplined case count gets multiplied by the matrix. Testing a mobile app economically depends on it.
Mobile UI testing benefits particularly from boundary analysis applied to screen dimensions, text scale settings, and content length.
Where Black Box Testing Reaches Its Limits
Coverage stays invisible. You can run a thorough black box testing suite and leave whole code paths untouched, with no signal telling you which ones.
Defects the specification omits stay invisible too. A requirement that fails to mention concurrency produces cases that fail to test it.
Diagnosis is slow. A failing case tells you the output was wrong while saying little about where, which lengthens investigation compared with a failing unit test.
The resolution is combination rather than replacement. Pair black box testing with structural coverage measurement, and use API testing at the contract layer to narrow where failures originate. Regression suites built from black box cases stay valid across refactors, which is their durable advantage.
How Autosana Fits
We built Autosana around the vantage point black box testing occupies, which is the outside of the application.
Our flows describe what a user does in plain language, so a case designed through partitioning or boundary analysis translates directly into an executable flow. Our guide on writing effective flow instructions covers phrasing that runs reliably.
We execute on real devices across iOS, Android, and web, which covers the interrupt and permission classes that emulator runs exclude.
We organize cases into suites and trigger them through automations on your schedule or from pipeline events.
Because flows anchor to intent, a refactor or redesign leaves them valid. That preserves the main advantage of specification-based cases, which selector-bound end-to-end suites usually forfeit.
Conclusion
Take one input field in your product and write down its partitions and boundaries. A quantity field, a date range, a monetary threshold.
Then compare that list against the black box testing cases currently covering it. The usual finding is several cases sitting inside one partition and zero cases sitting on the boundary, which means the existing coverage costs more and catches less than four well-chosen values would.
Start with boundary value analysis on your highest-value numeric inputs. Among black box testing techniques it returns the most defects per case.
FAQ
What is black box testing?
Black box testing derives cases from requirements and observable behavior, with the implementation treated as opaque. You supply inputs, observe outputs, and compare against the specification.
What are the main black box testing techniques?
Equivalence partitioning, boundary value analysis, decision table testing, state transition testing, and error guessing. The first two handle input ranges, the next two handle combinations and stateful behavior.
What is equivalence partitioning?
Dividing the input space into groups the system should treat identically, then testing one representative value per group. A field accepting ages 18 to 65 has three partitions covering billions of possible inputs.
What is boundary value analysis?
Testing values at the edges of each partition, plus the values immediately outside them. For a range of 18 to 65, that gives 17, 18, 65, and 66, which is where comparison operator errors surface.
When should you use a decision table?
When output depends on several conditions combining. Three binary conditions produce eight rows, and building the table usually exposes a combination the requirements left undefined.
How does black box testing differ from white box testing?
Black box derives cases from the specification with no code visibility, so cases survive refactoring. White box derives them from control flow and structure, which makes coverage directly measurable.
What are the limits of black box testing?
Code coverage stays invisible, defects the specification omits go untested, and a failing case indicates the wrong output while saying little about where the fault originated.
How do these techniques apply to mobile apps?
Interrupt states extend state transition testing, permission responses form equivalence classes, and device combinations multiply every case, which makes partitioning discipline more valuable than on web.
.png)