Demotion destroys its own evidence

AI governance · Perspective · August 2026

A marketplace that demotes a seller manufactures the evidence that will later justify the demotion. A return-rate spike trips the model, rank drops or a warning badge appears, and the rank cut pulls traffic down; fewer sessions mean slower orders and weaker conversion. Review the case a month later and every number has deteriorated, which reads as proof that the system caught something real. It proved no such thing. The metrics did not decay because the seller got worse. They decayed because the marketplace changed the conditions it was trying to measure.

We see this pattern across marketplace engagements, and the original trigger is usually ambiguous. A genuinely defective batch produces a return spike. So does a courier partner crushing boxes on one lane, a season of buyers ordering three sizes and keeping one, or a small cluster of abusive claims. The score that fired the demotion could not separate those causes at decision time, and that part is defensible; scores act on the information available. The harder part to defend comes after the action, when the marketplace stops collecting the information that would settle the question.

The evidence leaves with the traffic

Once visibility is cut, the marketplace no longer observes how the seller would have performed at normal exposure. An appeal can be filed, but an appeal argues over the old data. The data that would actually decide the case, defect behavior under ordinary traffic, was never generated. The loop closes on itself: the action degraded the numbers, and the degraded numbers defend the action.

Consumer lenders met this problem decades ago and gave it a name: reject inference. A declined applicant never produces repayment history, so a lender that studies only its approved book will keep concluding that its declines were sound. Credit-risk teams built a whole corrective methodology around that blind spot. Marketplaces run the same loop at far higher speed, with ranking suppression standing in for rejection, and mostly without the corrective discipline.

The cost of a wrong demotion rarely appears on any internal report. A good seller punished for a courier's failures tends to skip the fight; they move inventory to a rival marketplace and take their pricing and assortment depth with them. The enforcement ledger records only that the seller's numbers collapsed after the action, which is true of every demoted seller, including the ones who deserved reinstatement.

The experiment, designed before the action

The fix is to treat demotion as an intervention that must carry its own test. Keep a small, controlled amount of exposure alive after the action: the affected listings stay visible to a matched slice of buyers, sized so that customer harm stays bounded while the comparison stays valid. This is a holdout in the ordinary experimental sense, applied to enforcement.

Then measure the thing enforcement claims to reduce. Genuine defects per delivered order, controlled for category, geography, shipping legs and buyer behavior, is a defensible quality measure. The denominator matters as much as the numerator: delivered orders, so that courier failures stop double-counting against the seller. Raw return rate bundles size sampling, transit damage and abusive claims together with true product failure, and a seller can drown in sizing returns while shipping flawless goods.

Write the reinstatement rules before the first data point arrives. Thresholds drafted after the results are in will inherit every incentive to confirm the original call, and the experiment quietly becomes one more appeal. Set in advance, the read-out is clean. If measured quality improves only because exposure disappeared, the demotion proved nothing and the seller comes back. If the holdout keeps producing seller-caused defects at normal exposure, the demotion stands, and now it stands on evidence.

This widens what model governance has to cover. The standard review asks whether the score was reasonable on the inputs it had, and that review still matters. But a score can be audited against history because a prediction leaves the world unchanged. An enforcement action reshapes the very data any later audit will read. Governance that stops at score quality will certify systems that reduce orders and book the reduction as fewer defects. The audit has to follow the action, and the action has to preserve enough counterfactual exposure for the audit to mean anything.

The governance principle underneath all of this fits in one sentence: any decision that changes the data it will later be judged by must bring its own experiment, designed before the decision is taken.

Designing tests like this for enforcement and scoring systems is part of our custom engagements.