Compression is capacity only when the smaller representation is the one you keep and the one you read. That argument needs honest metrics. It also needs honesty about which training tricks actually move credit where it belongs.
A falling bit rate is a real signal
In a neural compressor, average bits-per-byte is not a vibe. It is the plant's bill. When that number drops in sizable, repeatable chunks, something in the model is naming more of the file cheaply. That is the clean improvement signal. You do not need a story about architecture fashion to trust a sustained compression gain under the same decode constraints.
Treat those drops as capacity candidates. Log them. Reproduce them. Ask whether the smaller codes still round-trip and whether decode stayed cheap enough to place where the data lives. A bit-rate curve that keeps falling under those proofs is how a stack earns its next deployment, not how a lab earns a slide.
One metric can smile while another slips
The trap is reading a single objective as the whole training surface. Classifier and training quality are often uneven across passes. Precision can climb, then degrade, while recall still looks healthy. The loss that dominates the dashboard can keep improving even as a secondary surface starts to wobble.
That is not a reason to discard the compression signal. It is a reason to keep more than one score in view. A simpler-looking objective does not automatically mean fewer passes. It can mean the opposite: the easy structure is gone, the residual is thin, and each additional pass has to prove it still buys bits without buying a false confidence story. When prediction saturates, more of the same on one loss is exactly how you burn compute. The same discipline applies to the metrics you celebrate mid-run.
Reconstruction is a local story
The same skepticism belongs on architectural shortcuts that promise to rewrite credit assignment. Recent proposals for causal or auxiliary networks that functionally decompose backpropagation — predicting activation gradients, parallelizing updates, leaning into forward-oriented or test-time style learning — are motivated by a real pain. Full backprop is expensive. Forward-only training would change the plant if it worked at global scale.
Motivation is not proof. Many of these designs lean on reconstruction losses to stand in for the credit that real backprop carries across the whole graph. A local reconstruction can look tidy. It can even track useful features. That does not mean it preserves the global credit-assignment properties that make reverse-mode differentiation effective. Floating-point error budgets and parallel update schedules make interesting papers. Without a convincing end-to-end case under the constraints that matter — exact reconstruction for a compressor, portable cheap decode, sustained bits-per-byte that survive more than a toy surface — the result stays almost useful.
Almost useful is a research status. It is not capacity.
Keep the proofs separate
Capacity work should keep two ledgers. On one side: compression metrics that the plant can cash — bits-per-byte, round-trip, decode cost — watched across enough passes that a smiling precision curve cannot hide a later slip. On the other: architectural claims about credit assignment, judged by whether they still produce those cashable metrics when the shortcut replaces the global path, not by whether an auxiliary network reconstructs activations in isolation.
Leave the hole was about refusing to force the wrong objective to speak. The same refusal applies here. Do not let a tidy local loss speak for global credit. Do not let a falling bit rate excuse a training surface you stopped measuring. The lead still goes to the stack that can show both: smaller codes you can keep, and credit that actually traveled to make them.
