
On 8 July 2025, Interfax reported that an artificial-intelligence service for identifying suspicious calls and warning users was among measures being discussed in Russia. That description concerned a proposed service, not a published measurement of its accuracy.
An alert raises a different question from the one answered by a detection rate. A detector might identify a large proportion of the events it is designed to find, yet produce a set of alerts containing many ordinary events. Those statements can both be correct. To understand why, a reader needs to know not only how the detector treats each category but also how common the target category is in the population being examined.
The following example isolates that relationship. It does not evaluate the proposed service, describe actual telephone traffic or suggest ignoring any real warning. The word useful has a narrow meaning here: the proportion of alerts that correspond to the stipulated target category. It is not a complete judgement about the commercial or protective value of a system. Consequences, response costs and missed events would all matter to a broader assessment, but no values for those consequences are supplied.
Start with events whose categories are already known
Imagine two separate collections of 10,000 events. Call them population A and population B. An event either belongs to a target category or is ordinary. For the arithmetic, the true category of every event is stipulated in advance and does not change. The detector produces either an alert or no alert. An alert is an output of the detector; it is not the definition of the true category. Keeping those two labels separate makes it possible to describe both correct and incorrect outputs.
In population A, 100 events belong to the target category and 9,900 are ordinary. The target share is therefore one percent. We stipulate that the detector alerts on 90 percent of target events and on one percent of ordinary events. These are exact proportions in our constructed example. They are not estimates obtained from a test of a real product, nor a promise that a random future collection would reproduce the same integer counts without variation.
The first rule produces 90 alerts from the 100 target events and leaves 10 target events without an alert. The second produces 99 alerts from the 9,900 ordinary events and leaves 9,801 ordinary events unflagged. All four groups are needed. Looking only at the 90 successful detections would conceal the ordinary events that also entered the alert set. Looking only at the 99 false alerts would conceal the much larger ordinary population from which they came.
- 90 target events generate alerts.
- 10 target events generate no alert.
- 99 ordinary events generate alerts.
- 9,801 ordinary events generate no alert.
The denominator changes the meaning of the percentage
The detector finds 90 out of 100 target events: a 90 percent detection rate, also called sensitivity. It flags 99 out of 9,900 ordinary events: a one percent false-positive rate. But there are 189 alerts altogether, formed by combining 90 target events with 99 ordinary ones. Only 90 out of 189 alerts correspond to targets, approximately 47.6 percent. That last proportion, often called precision, answers a question about the alert set rather than either original category.
There is no conflict between 90 percent sensitivity and approximately 47.6 percent precision. Sensitivity begins with target events and asks how many are flagged. Precision begins with flagged events and asks how many are targets. The conditioning group has changed. Replacing one percentage with the other would not simplify the report; it would change its subject. The numerical difference comes from the large ordinary population contributing false alerts, even though only a small fraction of that population is flagged.
Similarly, the one percent false-positive rate does not mean that one percent of alerts are false. In population A, the false fraction among alerts is 99 divided by 189, approximately 52.4 percent. The one percent figure uses all ordinary events as its denominator. The 52.4 percent figure uses all alerts. Both fractions describe the same 99 events, but they compare those events with different totals. A percentage without a stated comparison group leaves this distinction invisible.
Change the population, not the detector's conditional rates
Now consider population B. It contains 1,000 target events and 9,000 ordinary ones, still adding to 10,000. The target share is ten percent rather than one percent. Keep both detector rates exactly as before: it flags 90 percent of target events and one percent of ordinary events. Nothing in this comparison changes the decision threshold, introduces new information about individual events or makes the detector better at distinguishing categories. Only the relative sizes of those categories change.
There are now 900 target alerts and 100 missed targets. The ordinary category contributes 90 false alerts and 8,910 unflagged events. The total alert count is 990. Of those alerts, 900 correspond to the target category, giving precision of approximately 90.9 percent. The detector still misses one tenth of targets and still flags one hundredth of ordinary events. A much greater share of its alerts is now genuine because the population supplies many more targets relative to ordinary events.
Notice that the false-alert count falls slightly, from 99 to 90, despite the sharp rise in total alerts from 189 to 990. The ordinary category became smaller, so applying the same one percent rate produced fewer false alerts. Meanwhile, the target category became ten times larger, producing ten times as many true alerts. This is a composition effect. It does not require a change in the detector's conditional behaviour and does not establish that a new software version has improved anything.
Nor should the two populations be read as a recommended way to improve a real system. We have not described how events would be selected into either group, what information selection would require or whether it would change the conditional rates. The comparison is deliberately controlled on paper. It demonstrates that different alert precision can arise with identical category-specific detection rates. Applying that explanation to an actual system would require evidence about its population and outputs.
Overall accuracy can move in the opposite direction
Another familiar measure counts every correctly classified event, whether it generated an alert or not. In population A, the correct outputs are 90 target alerts plus 9,801 ordinary non-alerts, totalling 9,891. Overall accuracy is therefore 98.91 percent. In population B, there are 900 target alerts plus 8,910 ordinary non-alerts, totalling 9,810. Overall accuracy is 98.10 percent. By this measure, A scores higher even though its alert precision is far lower.
The reversal follows from the detector being correct on 99 percent of ordinary events but only 90 percent of target events. Population A contains more of the category on which the stipulated rule is more often correct. Population B contains more of the category on which it is less often correct. Changing the weights changes the combined accuracy. At the same time, the larger target category in B makes its alert set contain a higher proportion of true targets.
Calling one population's result better without specifying the measure would therefore be incomplete. A has the higher proportion of correct outputs across all events. B has the higher proportion of true targets among alerts. Sensitivity and the false-positive rate are identical. None of those statements alone identifies the preferred operating arrangement, because the example assigns no cost to a missed target, an unnecessary alert or a subsequent response. Different measures can legitimately rank the same constructed results differently.
A compact expression for the composition effect
Let p denote the fraction of all events belonging to the target category. In our example, true alerts occupy a fraction 0.9p of the whole population. False alerts occupy a fraction 0.01(1 − p), because ordinary events make up the remaining share. Dividing true alerts by all alerts gives the expression 0.9p / [0.9p + 0.01(1 − p)]. This is simply the same counting argument written without specifying a total population of 10,000.
Substituting p = 0.01 reproduces population A's approximately 47.6 percent precision. Substituting p = 0.10 reproduces population B's approximately 90.9 percent. The formula also shows why knowing sensitivity and the false-positive rate is insufficient: p still has to be supplied. A report that provides the two conditional rates but omits the population composition has not provided every input required to calculate the true share among alerts. The missing quantity cannot be recovered merely by restating the other two more precisely.
For an additional arithmetic check, half the alerts are targets when true and false alerts have equal counts. That requires 0.9p = 0.01(1 − p), giving p = 1/91, or approximately 1.099 percent. This is a dividing point for the population's composition under the stated rates. It is not a setting for a detector, a recommended operating threshold or an acceptable level of harm. No practical authority attaches to one half merely because it is an easy fraction to recognise.

An alert count cannot supply its own ground truth
Suppose a dashboard shows only that 189 alerts occurred. Without the stipulated category labels, that total does not reveal that 90 are true alerts and 99 are false alerts. Many different splits add to 189. The count is evidence of output volume, not independent evidence of output correctness. In the example we know the split because the construction supplies it. A real evaluation would need a basis for determining categories that is separate from the mere presence of a detector's flag.
The same limitation applies to events that generate no alert. Reviewing only flagged events may describe the reviewed alert set, but it does not directly establish how many targets were missed outside that set. Our sensitivity calculation uses both the 90 detected and the 10 missed targets in A. If the latter category were absent from the evidence, reporting 90 percent sensitivity would require assumptions or additional observations. A measurement cannot demonstrate coverage of a group it has never accounted for.
This article does not prescribe an evaluation study or assume that perfect real-world labels are easy to obtain. Delayed information, unresolved cases and changing definitions could complicate that task. The narrower point is logical: labels derived solely from the detector's own decision cannot independently establish that the decision was correct. Treating every alert as a confirmed target would make false positives disappear by definition, rather than demonstrating that the detector does not produce them.
Keep the unit of observation stable
The populations contain events, not necessarily distinct people, devices or organisations. Several events might involve the same participant. Counting ten flagged calls would not, by itself, establish ten different callers or ten different recipients. Our arithmetic uses event counts consistently from the initial population through the alert set. Switching to a count of people in the middle would change the denominator and break the reconciliation, even if every underlying event had been recorded correctly.
Likewise, the categories must refer to the same period and population as the alerts being evaluated. A target share from one group cannot automatically be used to interpret alerts in another. The controlled example intentionally changes that share and obtains a different result. A population boundary is therefore part of the meaning of a precision figure, not merely background information. Extending a percentage beyond its original population requires a reason to believe the relevant inputs still apply.
Repeated events also warn against treating the illustrative percentages as guarantees for arbitrary collections. The construction imposes exact counts so that the denominator changes can be inspected directly. It does not assert independent random events or quantify sampling variation. Those would be additional modelling choices. A real observed proportion could differ because of statistical variation, different event dependence or different conditional behaviour. None of those possibilities needs to be introduced to establish the composition effect demonstrated here.
More alerts and more work are not the same measurement
Population B generates 990 alerts rather than A's 189, even though a larger proportion of B's alerts corresponds to targets. That count could matter to a hypothetical review process, but it does not specify a staffing requirement. The example gives no handling times, no automation rules and no requirement that every alert be reviewed in the same way. Multiplying alert volume by an invented universal workload would introduce a separate operational model, not complete the existing calculation.
Nor does higher precision mean that missed targets have disappeared. Population B contains 100 missed targets compared with A's 10 because its target population is larger. The missed fraction remains ten percent in both cases. Thus a report could truthfully describe a higher true share among alerts and a higher absolute count of missed targets at the same time. Comparing only one of those statements would conceal the other, rather than settle the overall value of the detector.
Separate a changed population from changed behaviour
In this construction, the detector's conditional rates do not deteriorate. If an observer saw precision fall when moving from B to A, the entire change would be explained by the target share falling from ten percent to one percent. It would be incorrect to attribute that constructed decline to a software defect. In a real setting, however, neither stable rates nor changing composition should be assumed without evidence. The example supplies one sufficient explanation, not a diagnosis for every decline in a reported metric.
The distinction also limits comparisons between suppliers or departments. Two published precision figures can differ because their populations differ, because their detectors differ, or because both differ. Our example isolates the first possibility. It does not establish a universal adjustment procedure or make unlike datasets comparable by assertion. A meaningful comparison would at least identify the populations, the category definitions and the measure being reported before treating a numerical gap as a performance advantage.
Describe what a signal actually establishes
The two populations produce a clear result: identical sensitivity and false-positive rates need not produce identical alert precision. In A, 90 true alerts sit alongside 99 false ones. In B, 900 sit alongside 90. The difference arises because the starting populations contain different proportions of targets. Meanwhile, overall accuracy ranks A above B, showing why an unspecified claim of accuracy can conceal more than it explains.
A careful account therefore keeps three questions separate: how many target events were detected, how many ordinary events were flagged, and how many alerts were true targets. Each uses a different denominator. Stating those denominators and the population behind them does not tell a person how to respond to any particular warning. It makes the measurement readable, prevents one conditional percentage from being relabelled as another, and keeps a large alert count from becoming an unsupported claim of success.