28B 上 xHC 只比 vanilla 多 3.0% 训练 FLOPs,却在多项推理和中文 benchmark上拉开明显差距,说明“扩残差记忆”这条路在大模型里确实有真实增益。 Figure 5 说明 t…...