Enforcement Triage as Constrained MPC: From FIFO to Foresight
DOI:
https://doi.org/10.31224/8126Keywords:
Model Predictive Control, queueing model, stochastic modelingAbstract
In a trust-and-safety review backlog, harm accrues while genuine cases wait to be actioned, cases carry consumer-protection deadlines, review capacity is scarce and below the mean arrival rate, and the magnitude of harm caused by a case is observable only through a noisy classifier score until it is eventually reviewed by a human. Production systems prioritize this backlog in a myopic fashion, using a first-in-first-out approach or by choosing those cases with the highest estimated harm based on a machine-learning classifier. These approaches ignore the forecastable quantities of future load and the rate at which pending harm accrues. The problem can be posed as a constrained model predictive control (MPC) over a marked-arrival backlog. A seeded, config-driven simulator anchored to public data provides the simulated environment to test MPC's efficacy.
Evaluating against a series of baseline approaches (the c-mu rule, a Whittle-inspired index, a full-information greedy benchmark, and a PPO reinforcement-learning agent), a fluid linear-program MPC controller is designed and tested on paired seeds. Three findings emerge and form the primary contribution of this work: Firstly, the cost of partial observability is large and dominant under a realistic signal. A factorial attribution assigns roughly 20% of it to unobserved harm magnitude, which improving a label-only classifier cannot address under our score-magnitude independence assumption but which can be learned online, and 80% to label uncertainty. Secondly, the forecast-aware controller beats the strong myopic rules only conditionally, namely when the deadline constraint carries weight and load is near-critical, and can be significantly worse than the c-mu rule when its internal model is misleading. Finally, both failure modes share one cause and can be addressed by applying a dual-mode controller which executes the myopic rule when the plan's anticipatory value is found to be negligible. It reveals no statistically significant disadvantage against the better of the pure c-mu and MPC policies in any of the tested regimes. In fact, it is significantly better than both pure rules in several regimes, and it retains MPC's advantages where foresight is particularly valuable.
Downloads
Downloads
Posted
Versions
- 2026-09-02 (2)
- 2026-08-31 (1)
License
Copyright (c) 2026 Chinmay Koimuttum, Rohan Shekhar

This work is licensed under a Creative Commons Attribution 4.0 International License.