Leave the hole argued for refusal: the causal pass should mark blanks it cannot own and hand them to a specialist. That architecture only works if the stage that decides which blanks are real is trustworthy. A cascade of specialists fails at the gate, not at the repair.
Recall without precision invents work
A router or flagger that catches most hard positions can still look successful on a miss-rate chart. Lower misses feel like progress. They are not enough. When precision sits near a third — when the system predicts twelve holes and roughly nine exist — the stack invents work. Every false hole is a specialist call on structure that did not need one. Every invented blank burns decode budget and training signal on noise the causal pass already owned.
That gap is easy to miss under teacher-forced metrics. A score that looks healthy around eighty percent can coexist with real precision stuck in the thirty-to-forty range. The plant does not cash teacher-forced optimism. It cashes what the live gate actually marks. For a residual cascade to be viable, precision at that gate needs to sit near ninety percent — not "better than last week," but high enough that false holes stop dominating the residual surface.
Staged specialists compound the error
Training three specialists apart, then wiring them in sequence, multiplies the gate's mistakes. The flagger's false positives become the repair model's training distribution. The repair model's confident answers on invented blanks become the next stage's prior. Separate ledgers do not average into a joint truth. They compound.
The proposed fix is not a fourth specialist. It is joint training: shared weights, end-to-end optimization, one loss surface that charges the gate for inventing holes as hard as it charges the repair for missing real ones. The model may grow — roughly threefold is a realistic cost — and it will want heavier GPU. That is still cheaper than an indefinite cascade that looks sharp on miss rate while precision never clears the bar.
Corrupted labels are not a precision jump
A second failure mode wears the costume of progress. Change the objective, watch precision leap, celebrate. Then discover the labels came from the wrong corpus. Metrics on corrupted labels are not progress. They are a story about a different dataset. Restart from corrected data. Do not ship the jump. Do not let a dashboard rewrite that invalidated a celebrated gain sit next to a capacity claim.
Reconstruction is not credit made the same demand of training tricks: keep the proofs separate. The same demand belongs here. A smiling precision curve on the wrong labels is not a gate you can trust. A falling miss rate with invented holes is not a router you can route through.
Capacity waits on the gate
The hard bits need both sides. Compression is capacity only when the smaller representation is the one you keep and the one you read. A specialist cascade earns that claim only when the stage that decides what needs work is precise enough that false holes do not eat the gain. Until precision clears the gate, more specialists are just more ways to invent blanks.
