When the system fails in both directions
A signal-based duty fails in two opposite ways, and a desk that only describes one is selling something. It fails when a rule reads an ordinary week as a pattern, which costs the player access they did not need to lose - 1,600 of 4,000 reviews in the sample. And it fails when a pattern that should have been caught is not, which costs far more - 360 of the 36,000 accounts that met no review at all.
- no-action reviews
- 1,600
- false-positive share of reviews
- 40.0%
- caused by a one-off
- 1,280
- caused by a data error
- 224
- accounts missed
- 360
- missed share
- 1.0%
Read the signals
Five comparisons the account already supports, run continuously rather than after a complaint. 6,000 flags in the sample, 96.0% of them relative to the account own history.
Choose a response
Nothing, a message, a limit or a break, or a restriction - one of four, graded by how much it takes from the player. 960 of 4,000 reviews in the sample changed access.
Write it down
Signals, crossing dates, rule version, path, response and outcome - one record per review, including the 1,600 that ended with nothing done.
Take the appeal
A route outside the desk that decided. 480 appeals in the sample, 176 of them upheld, which is 36.7% - high enough that the route does real work.
A protection duty fails in two directions. A rule can read an ordinary week as a pattern - 1,600 of 4,000 reviews in the sample ended with no action, 1,280 of them explained by a one-off event and 224 by a data error. And a pattern can be missed: 360 of the 36,000 unreviewed accounts should have met a trigger, 288 because the window had closed and 72 because the pattern was split across two accounts.
Failure one: the rule was too eager
A threshold that never fires costs nothing and catches nothing; a threshold set to fire early catches more and costs the players who did nothing wrong. The sampled policy sits deliberately on the early side, and the bill is 1,600 reviews out of 4,000 that ended with no action - 40.0% of every review in the quarter. Of those, 1,280 (80.0%) were explained by a one-off event and 224 (14.0%) by a data error, and 96 (6.0%) were never explained at all.
The 224 data errors are the ones that should be easiest to eliminate and are most often left: a shared device counted as two different players, a currency conversion figure read as a spend, a deposit recorded twice. Each one is a rule acting on a number that was wrong, and none of them is a judgement call.
Failure two: the rule was too quiet
The other direction is worse, because nobody is contacted about it. A retrospective audit of the 36,000 accounts that met no review found 360 that should have met a trigger - 1.0% of the unreviewed population. Of those, 288 (80.0%) were missed because the flag was raised after the review window had closed, and 72 (20.0%) because the pattern was split across two accounts and neither reached the rule alone.
| Too eager | Too quiet | |
|---|---|---|
| Where it shows up | reviews that end with no action | an audit of unreviewed accounts |
| Scale in the sample | 1,600 of 4,000 reviews (40.0%) | 360 of 36,000 accounts (1.0%) |
| Main cause | a one-off event, 1,280 of 1,600 | a window that closed first, 288 of 360 |
| Who pays | the player, in access and in a contact | the player, invisibly |
| How it is found | the review itself | only by going back and looking |
The right-hand column is why this desk keeps returning to the record and the appeal: the eager failure is visible, answerable and in 480 cases in the sample actually answered. The quiet failure has no route at all except a retrospective audit, and an operator that never runs one cannot know its own rate.
The trade is real, and it is a choice
Worked example / moving the threshold
- the sample trigger: 4.00x the declared band, which produced 480 financial reviews
- at 3.00x, more accounts clear it, so the missed financial cases fall and the no-action reviews rise
- at 6.00x, fewer accounts clear it, so the contacts fall and the missed cases rise
- in the sample, every 0.5x moved the financial review count by roughly 60 accounts in either direction
- there is no setting at which both counts fall at once, because both are produced by the same comparison
- 1,280 of 1,600 false positives were explained by a one-off event, so a rule that could read context would remove the largest single cause.
- 224 were data errors, which are removable rather than tunable.
- 96 were never explained, and the account was left unchanged anyway - the honest residue.
- 288 of 360 missed cases were a window that had closed, which is a design choice rather than a bug.