The hard bits need both sides. That was the case for a second objective: hide the expensive positions, let a specialist see left and right, and route by who should speak. The missing move is what the causal model does before the specialist arrives.
It has to leave a hole.
Forced speech is expensive
A next-token compressor that cannot name a position still has to emit something. Under an arithmetic or entropy coder, a confident wrong guess is not a small miss. It is a long code. Mistakes are charged steeper than silence. The plant pays for every bit that leaves the encoder, and a loud residual that was never sequential structure is how you burn ratio without buying capacity.
Pushing the same causal objective harder on those positions does not fix the bill. It repeats the bill. The model was never the right voice for that blank.
Refusal is part of the stack
The clean architecture is not "make the sequential predictor smarter until the residual disappears." It is refusal. Mark the positions the causal pass should not own. Leave gap vectors — deliberate holes — so nothing leaks across a causality boundary the specialist is about to fill. The router is not a third brain. It is the rule that keeps the wrong objective from speaking.
Train the selector for that refusal on its own surface. Do not fine-tune it from the weights that already learned to shout through the residual. A fresh objective on a fresh head learns which blanks to skip. Borrow features from the causal stack if they help. Do not borrow its habit of answering every token.
Same family, different job
The specialist that fills the hole is still the same architecture family: a masked pass on the catalog of high-cost positions, not a larger causal model with the arrow flipped. The causal pass keeps the easy structure. The masked pass reconstructs only what was left open. Decode still has to round-trip exactly and stay cheap enough to place where the data lives. A hole that does not reconstruct is not capacity. It is a broken pipe.
The same brief travels into video and other dense media. Low-entropy test material will not teach a residual specialist anything useful. The positions that still cost bits need enough variety to be a real surface — and then the same rule applies: leave the blanks the sequential model cannot own, and hand them to the objective that can see both edges.
Capacity is the refusal
Compression is capacity only if the smaller representation is the one you keep and the one you read. A saturated predictor is not the end of that argument. Past saturation, the lead is the stack that knows when not to speak — and the specialist waiting in the hole.
