White Box Testing: What Coverage Percentages Actually Tell You
White box testing measures coverage directly. Statement, branch, condition, and path coverage explained, why 100 percent statement coverage misleads, and where the useful floor sits.
Yuvan Sundrani · 11 min read
autosana.ai
.png)
TL;DR
- White box testing derives cases from source code structure, so coverage becomes directly measurable as a percentage.
- A function with one conditional reaches 100 percent statement coverage from a single test while leaving half its branches unexecuted.
- Branch coverage is the useful floor to enforce. Path coverage grows exponentially with complexity and stops being achievable past a handful of conditionals.
- Autosana covers the layer white box testing structurally excludes, running intent-based flows on real iOS, Android, and web devices from your pipeline.
A function contains one if statement. A single test executing the true branch runs every line in that function, so statement coverage reports 100 percent.
The false branch has yet to execute. That branch is usually the one returning the error, and the report showing 100 percent is what stops anyone from looking.
White box testing is the practice of deriving cases from code structure rather than from the specification, and the value it delivers depends entirely on which coverage metric you read. This guide covers what each white box testing metric measures and where the useful floor sits.
What White Box Testing Is
White box testing designs cases with full visibility of source code. You read the control flow, identify the paths through it, and write cases that exercise them.
The ISO 29119 standard covers them alongside specification-based ones as complementary families.
The defining advantage of white box testing is measurability. Where specification-based work leaves you guessing which code your cases reached, white box testing reports it as a number.
The defining cost is fragility. Cases bound to control flow get rewritten when that flow changes, so a refactor that preserves behavior can still invalidate a suite.
The Four Coverage Metrics
Consider a discount function: if the customer holds a membership and the order exceeds 5000, apply 20 percent; otherwise apply 5 percent.
| Metric | What It Requires | Cases Needed Here |
|---|---|---|
| Statement | Every line executes at least once | 2 |
| Branch | Every decision takes both outcomes | 2 |
| Condition | Every individual condition takes both values | 3 |
| Path | Every route through the function executes | 4 |
Statement coverage
Every executable line runs at least once. This is the weakest useful metric, and the one most commonly reported.
For a function with a single if and no else, one test through the true branch achieves 100 percent while leaving the implicit false path unexercised.
Branch coverage
Every decision point takes both outcomes. This catches the gap statement coverage misses, and it is the metric worth enforcing as a floor.
For the discount function, one test with membership and a 6000 order plus one test with no membership and a 3000 order covers both branches.
Condition coverage
Each individual condition inside a compound expression takes both true and false. The discount function has two conditions joined by and, so covering each independently requires a third case.
This matters where short-circuit evaluation hides defects. An expression where the second condition is defective passes whenever the first condition is false, because evaluation stops before reaching it.
Path coverage
Every distinct route through the function executes. Two independent conditions produce four paths, and path count doubles with each additional independent condition.
Ten conditionals produce 1024 paths. This is why path coverage stays a theoretical target outside safety-critical work, where regulation requires it.
Reading Coverage Percentages Honestly
A number without a metric name carries little information. Eighty percent statement coverage and eighty percent branch coverage describe very different suites.
Three distortions recur.
Getters and setters inflate the figure. Trivial accessors execute during almost any test, so a codebase heavy in boilerplate reports high coverage with thin real verification.
Executed differs from asserted. A line running during a test that asserts something unrelated counts toward coverage while verifying nothing. Coverage measures execution rather than checking.
Aggregate figures hide distribution. Eighty percent overall can mean ninety-five percent on utility code and forty percent on the payment module, and the aggregate gives no signal about which.
A more useful practice: enforce branch coverage as a floor on new and modified code rather than chasing a global number, and read coverage per module against that module's business value.
Complexity Sets the Case Count
Cyclomatic complexity counts the independent paths through a function, and it equals the minimum number of cases needed for branch coverage.
A function with complexity 1 has no branching and needs one case. Complexity 5 needs five. Complexity 20 needs twenty, and a function that complex is usually better split than tested exhaustively.
Use the metric as a design signal rather than only a testing input. Where complexity rises past 10, the case count required is telling you the function carries too many responsibilities.
Martin Fowler's test pyramid reasoning applies directly. White box testing belongs at the base where functions are small and fast to exercise, and it becomes impractical higher up.
White Box Testing on Mobile
The technique splits cleanly by layer.
Business logic, view models, data mapping, and networking code all suit it. These are ordinary functions with measurable control flow, and coverage tooling reports on them normally.
UI code resists it. View hierarchies, gesture handling, and platform lifecycle callbacks carry control flow that lives partly inside the framework, so coverage figures there describe your code while the failure modes live in the interaction.
Platform behavior sits outside it entirely. Vendor OS forks, permission dialogs, and thermal throttling have no representation in your control flow graph, which means white box testing reports full coverage on code that fails on a Galaxy.
Android test automation and mobile test automation cover that outer layer, and the two approaches measure different risks.
Instrumentation adds a wrinkle. Coverage tooling on Android instruments the build, which changes timing and can mask race conditions that appear in release builds.
White Box Testing in CI/CD
Run white box testing coverage measurement on every pull request and report the delta rather than the absolute figure. A change dropping branch coverage on modified files is actionable, while a global percentage moving from 76.2 to 76.1 is noise.
Gate on new code rather than the whole codebase. Requiring 80 percent branch coverage on changed lines is enforceable. Requiring it retroactively across a legacy codebase produces a large ticket that stays open.
Keep the measurement fast. Coverage instrumentation slows execution, so run it on the unit tier rather than across regression suites where the cost outweighs the signal.
Pair the coverage gate with performance monitoring in production, since structural coverage says little about behavior under real load. Teams running CI/CD integration can surface both signals on the same pull request.
Combining the Two Approaches
Structure-based and specification-based work find different defects, and a suite using only one carries a predictable blind spot.
White box testing finds unreachable code, uncovered branches, and logic errors inside a function. It stays blind to requirements the code fails to implement, because a missing feature has zero control flow to cover.
Specification-based work finds wrong outputs and requirement gaps while leaving coverage unmeasured. The two together give you both the map and the territory.
A workable division: structure-based cases at the unit tier where they are cheap and measurable, specification-based flows at the system tier where they survive refactoring. End-to-end suites built the other way round decay quickly.
How Autosana Fits
We built Autosana for the layer white box testing structurally excludes, which is behavior on real hardware.
Our flows describe user intent rather than control flow, so they stay valid through the refactors that invalidate structure-based cases. For teams that prefer keeping tests alongside source, code-managed flows live in your repository and version with the code.
We run on real iOS and Android devices, covering vendor OS behavior and thermal effects that have no representation in a coverage report.
We organize flows into suites triggered through CI/CD integration, so structural coverage and behavioral verification report on the same pull request.
Our performance monitoring tracks startup time and frame rates, which complements coverage by measuring what performance testing covers and structure leaves out.
Conclusion
Check which metric your coverage report displays. If it says statement coverage, the number is higher than your actual verification.
Switch the gate to branch coverage on changed lines, then read the per-module figures against what each module is worth. A payment service at 45 percent branch coverage beside a string utility at 98 percent is the specific finding worth acting on, and an aggregate figure will hide it.
Structure tells you which code ran. Behavior tells you whether the product works. Measure both, and treat each as a complement to the other.
FAQ
What is white box testing?
White box testing designs cases from source code structure, using visibility of control flow to exercise specific paths. Coverage becomes directly measurable, unlike specification-based approaches.
What are the main white box testing techniques?
Statement coverage runs every line, branch coverage takes every decision both ways, condition coverage exercises each condition in a compound expression, and path coverage runs every route through a function.
Why is 100 percent statement coverage misleading?
A function with one if and no else reaches full statement coverage from a single test through the true branch. The false path stays unexecuted, and it is usually the error-handling path.
Which coverage metric should teams enforce?
Branch coverage, applied to new and modified code rather than the whole codebase. It catches the gap statement coverage misses and stays achievable, unlike path coverage.
What is cyclomatic complexity?
A count of independent paths through a function, equal to the minimum number of cases needed for branch coverage. Complexity above 10 usually signals a function that should be split.
Why does path coverage become impractical?
Path count doubles with each independent condition. Ten conditionals produce 1024 paths, which keeps full path coverage confined to safety-critical work where regulation requires it.
How does white box testing apply to mobile apps?
It suits business logic, view models, and networking code. UI behavior, vendor OS forks, and thermal effects have no representation in your control flow, so coverage reports stay silent about them.
How do white box and black box testing work together?
Structure-based cases find uncovered branches and logic errors at the unit tier. Specification-based cases find wrong outputs and requirement gaps at the system tier, and survive refactoring.
.png)