DTR 动态张量重计算被曝对内存预算极度敏感:0.10% 差异导致 7.3 倍开销
Deterministic Regime Switching and Feasibility Inversion in Dynamic Tensor Rematerialization
一个 arXiv 预印本,用 DTR 官方模拟器实测发现显存预算只差 0.10%,训练开销能差 7.3 倍,还复现了奇怪的可逆 OOM 区间,做省显存训练的可以看看细节。
研究在 DTR 参考模拟器 simrd 上用公开执行轨迹发现确定性不稳定:LSTM 轨迹中内存预算相差 0.10% 的两次运行,开销差距最高达 7.3 倍,原因是对同一批存储的反复逐出(每存储逐出次数从 1.33 升到 8.27)。在 ResNet-32 轨迹上出现确定性可行性反转:预算比例 0.101 时可行,0.102-0.106 区间 OOM,0.107 起又恢复可行。OOM 的直接原因是被完全钉住的递归重计算边界超出预算。消融实验将 LSTM 不稳定归因于 size-staleness 联合评分项。
Deterministic Regime Switching and Feasibility Inversion in Dynamic Tensor Rematerialization
We report fine-grained, deterministic instability in Dynamic Tensor Rematerialization (DTR), an online eviction policy for memory-constrained DNN training, measured on the reference DTR simulator (simrd) using public execution traces. On an LSTM trace, memory budgets differing by 0.10% of unconstrained peak memory select fast and slow execution regimes whose overheads differ by as much as 7.3x; the slow regime is driven by broadly repeated re-eviction of the same storages (evictions per storage rise from 1.33 to 8.27 while the set of distinct evicted storages is essentially unchanged: 5,233 vs 5,236, with the two sets overlapping at Jaccard 0.999). On a ResNet-32 trace, a fine budget sweep reveals a deterministic feasibility inversion: the run is feasible at ratio 0.101, infeasible (OOM) across 0.102-0.106, and feasible again from 0.107. We trace the immediate cause of the OOM to a fully pinned recursive rematerialization frontier that exceeds the budget after every evictable tensor has been evicted. Ablations using the DTR authors' own variants implicate the joint size-staleness scoring term in the observed LSTM instability. We argue these are at least two distinct budget-sensitive pathologies rather than one mechanism, and we separate what is demonstrated from what remains hypothesised. All results concern the reference simulator; reproduction in a production runtime is future work. Code, instrumentation, and raw results accompany this preprint.