
A clean sample can be a genuine observation and still leave defective items in the batch. There is no contradiction. The inspection describes the items selected, while the receiving decision concerns a larger set. Connecting those two statements requires a selection rule and a clearly specified question. A calculation about the chance of finding a defect is not automatically a calculation about the quality of everything that remains.
On 13 August 2025, Interfax reported discussion of a honey-market charter in Russia. Roskachestvo deputy head Elena Sarattseva linked the initiative to laboratory research. The fictional sampling exercise below does not describe that programme's inspection rules.
The business question is deliberately narrower than whether a charter can restore confidence. Suppose a receiving team inspects only part of a delivery. What does the inspection design imply about its ability to encounter a specified problem? An anonymous numerical example can answer that question without making claims about any named producer, recommending a food-testing procedure or assigning an actual defect rate to a market.
Begin with a fixed batch, not a changing market
Imagine a batch of twenty distinct sealed items. Exactly two fail an invented binary specification; the other eighteen meet it. The items do not deteriorate, improve, enter or leave while the sample is selected. Nobody replaces an inconvenient item. This is a fixed population, and its two failures are stipulated by the analyst. They are not a finding obtained from the sample or a belief inferred by the receiving manager.
The distinction between the analyst's premise and the manager's knowledge is essential. We are asking what would happen across possible samples if the batch contained exactly two failures. We are not giving the manager a magical view inside every package. Nor are we assuming that real deliveries contain two failures in twenty. The numbers make a conditional question small enough to examine completely.
Assume that a perfect test identifies whether each selected item meets the fictional specification. It never misses a failure in an item actually tested and never labels a conforming item as defective. This isolates selection coverage from test accuracy. If the sample contains a failure, detection is certain under this assumption. The only way the inspection misses the problem is to select entirely from the eighteen conforming items.
Select a fixed number of distinct items uniformly at random, without replacement. Every subset of that size is equally likely. An item already selected cannot fill another place in the sample. This is not the same as repeatedly examining the nearest package or allowing a supplier to present whichever items are convenient. The calculation depends on the stated selection rule, not merely on the number of completed tests.
Two checks leave many ways to miss
Start with a sample of two items. The first selected item is conforming with probability eighteen out of twenty. Conditional on that first clean selection, seventeen of the remaining nineteen items are conforming. The probability that both selections are clean is therefore eighteen twentieths multiplied by seventeen nineteenths. The result is 306 divided by 380, approximately 80.5 percent.
That is a high probability of seeing no failure even though two failures are definitely present in the fictional batch. A clean pair is therefore not a surprising result under the stipulated condition. The result does not mean the test is inaccurate: both tested items really meet the specification. The problem lies outside the selected pair, where the perfect test has not been applied.
The chance of detecting at least one failure is the complement, approximately 19.5 percent. These two outcomes exhaust the possibilities for the question being asked. A sample either contains at least one of the two failures or contains neither. We do not need to distinguish one detected failure from two to calculate the probability of detecting any failure at all.
Count possible samples, not inspection anecdotes
There are 190 different unordered pairs that can be selected from twenty distinct items. Of those pairs, 153 contain only conforming items because they can be formed entirely from the eighteen conforming items. The fraction 153 out of 190 is the same as 306 out of 380. Counting pairs provides an independent check on the sequential multiplication without treating the order of inspection as an additional sample.
A single completed inspection produces one pair, not the entire set of possible pairs. Probability describes the selection design across those possibilities. Once the two particular packages have been tested, their results are observed facts. Confusing these levels can lead a report to describe the realised clean sample as though it had resolved every alternative sample that could have been selected.
Increase coverage while keeping the batch unchanged
Now let the sample contain five distinct items, leaving everything else unchanged. The chance of selecting only conforming items is the number of five-item subsets available from eighteen conforming items divided by the number available from all twenty. In standard counting notation, that is C(18,5) divided by C(20,5). It equals 210 divided by 380, approximately 55.3 percent.
The same reasoning gives a compact expression for a sample of n items: C(18,n) divided by C(20,n). For n from zero through eighteen, it simplifies to (20 minus n) multiplied by (19 minus n), divided by 380. The denominator stays fixed because the batch size and the number of failures stay fixed. Changing the sample does not improve the batch; it changes the chance of encountering its existing failures.
- Testing two distinct items leaves an approximately 80.5 percent chance of detecting no failure.
- Testing five leaves an approximately 55.3 percent chance.
- Testing ten leaves an approximately 23.7 percent chance.
- Testing fifteen leaves an approximately 5.3 percent chance.
- Testing sixteen leaves an approximately 3.2 percent chance.
Every percentage is conditional on exactly two failures in this twenty-item batch and on the uniform, without-replacement selection rule. None is a measured rate for a company or a product category. The list compares inspection designs within one invented situation. It does not establish how many items a real business should inspect, what a regulator requires or which level of residual risk is acceptable.
For example, an arbitrary comparison with five percent separates the samples of fifteen and sixteen: the former remains slightly above that line, while the latter falls below it. This is only arithmetic within the exercise. Five percent has not been supplied by a contract, a safety assessment or a legal standard. Giving the line a familiar appearance would not give it an external authority.

The boundary cases explain the mechanism
With no items tested, the probability of detecting no failure is one. Nothing has been examined, so detection cannot occur. With eighteen items tested, the probability of no detection is two divided by 380, approximately 0.53 percent. It is small but not zero: the sample might contain precisely the eighteen conforming items and leave both failures outside.
With nineteen distinct items tested, missing both failures becomes impossible under the assumptions. Only one item remains untested, and that single place cannot contain both failures. Testing all twenty also guarantees detection. These endpoint results follow from the stipulated count and the perfect test. They are not assurances about a real inspection in which the number of failures or the test's accuracy is unknown.
There is another way to see the general expression. A clean sample requires both failed items to be among the twenty minus n unselected positions. Counting the possible positions for that pair produces the same ratio. This view shifts attention from the packages examined to the room left outside the sample. The inspection can miss only if both failures fit there.
The reasoning also shows why distinct items matter. Testing the same conforming package repeatedly does not reduce the number of unexamined packages. It may answer a different question about the repeatability of a test, but it does not reproduce the coverage gained by selecting additional items. A dashboard counting test runs alone could obscure this difference between repeated measurement and wider selection.
Do not reverse the conditional probability
After a clean sample, a manager might want to know the probability that the entire batch is clean. Our calculation has not answered that question. It answers the probability of a clean sample given a batch with two failures. Reversing the order of those conditions produces a different quantity. The numerical result cannot simply be relabelled as the probability that an untested item or the whole batch is acceptable.
Within the fictional premise, the batch contains two failures by definition. A clean sample does not erase them. To infer an unknown batch condition from an observed sample, another model would be needed: one describing the possible batch conditions and the information available before the inspection. This article does not supply such a model or choose prior probabilities on a receiving team's behalf.
The distinction prevents two opposite mistakes. One is to claim certainty of cleanliness because the sample was clean. The other is to announce that the batch has an 80.5 percent probability of being defective because that number appeared in the two-item calculation. Neither statement follows. The number describes a clean-sample outcome under a particular already-defective condition, not a posterior judgement about an unknown delivery.
A precise record can instead say what was tested and what was found. Separately, a design calculation can state its conditional coverage for an assumed configuration. Keeping those statements separate does not make the inspection useless. It identifies the information actually obtained and the additional assumptions that would be required for a broader inference.
Convenient selection changes the question
The formula gives equal probability to every subset of the chosen size. If the selection procedure favours the front of a pallet, the earliest arriving container or a specially prepared set of packages, that condition may not hold. The count of tests could remain unchanged while their relationship to the batch changes. In that case, quoting the uniform-sample probability would describe a design that was not used.
This is a logical boundary rather than a claim that any particular supplier behaves improperly. Convenience can arise without deception. A receiving team may simply choose items that are easiest to reach. But ease of access and equal selection probability are different properties. The fictional calculation cannot establish that they coincide in a real storage arrangement.
The population boundary matters too. If the twenty listed items belong to one delivery while the sample is taken from another, accurate testing does not repair the mismatch. Similarly, adding later arrivals to the batch after selecting the sample changes the set to which the result is being attached. A calculation for one fixed population cannot silently become evidence for a larger or different population.
Coverage and test accuracy belong on separate lines
Even complete coverage would not solve every possible testing problem outside the model. The perfect-test assumption deliberately removes false results and incomplete specification coverage. A test that addresses one property does not automatically establish every property of a product. An inspection count and a description of what each inspection can detect therefore answer different questions.
Our example does not quantify imperfect tests, correlated errors or multiple kinds of failure. Adding those features would require new assumptions, not a casual adjustment to the displayed percentages. The useful contribution of the simple case is to make one source of uncertainty visible before mixing it with others: a failure can remain undiscovered solely because the item containing it was never selected.
Nor does the example optimise inspection expenditure. Larger samples may require more work, but no cost schedule or consequence valuation has been specified. Choosing a sample size would require the real decision's objectives and constraints. The probability table alone cannot identify a commercially optimal or otherwise sufficient procedure. It is an explanation of conditional detection coverage, not a purchasing recommendation.
A clean result deserves an accurate description
The fictional receiving record has three separate elements: the fixed population to which the sample belongs, the method used to select distinct items, and the results obtained on those items. A design calculation can then add a fourth element, clearly labelled as conditional on an assumed batch composition and test accuracy. Collapsing these elements into a single pass label loses the distinction the exercise was built to reveal.
In the twenty-item example, two perfectly accurate clean tests leave a substantial chance of having missed the stipulated failures. Fifteen or sixteen distinct tests provide different coverage, and the endpoint cases explain why. None of these calculations changes the contents of the batch. They change what the selection procedure is capable of encountering. That is the difference between an honest clean-sample result and an unsupported claim about every item in a delivery.