论文

InferOpt:把 LLM 推理配置搜索化为约束多目标优化

InferOpt: Constrained Multi-Objective Search for LLM Inference Configurations

精选理由

一篇教你怎么调推理配置的论文,Qwen 和 DeepSeek 两个空间上省缓存、降延迟都有具体数字,做推理优化的可以看看。

arXiv 论文提出 InferOpt,将推理配置调优重构为约束多目标黑盒优化,只需变量边界、确定性资源成本和评估钩子即可搜索。该框架在冻结的采样代理集上搜索,在调用模型前拒绝超预算候选,并在全量数据集上复验 Pareto 代表解。在 Qwen2.5-7B 的 28 维连续 KV 空间上,prefill 后剪枝将 16K 缓存削减 64.4%,TPOT 降低 22.9%–48.5%,搜索出的分层预算比匹配的均匀预算高 7.3% 和 14.0%。在 DeepSeek-V2-Lite 的 26 维离散 MoE 空间上,搜索出的 top-k 调度移除 43.0% 的 token–expert 路由对,性能与默认配置差距在 0.59 分以内。对比 Random Search、NSGA-II 和 MOTPE,InferOpt 在两个空间均领先。

原文 · arXiv: DeepSeek

InferOpt: Constrained Multi-Objective Search for LLM Inference Configurations

Serving an LLM means setting dozens of inference-time knobs, from per-layer KV retention to per-layer expert counts. Practice sets them with mechanism-specific heuristics that return a single operating point and do not scale to layer-wise search spaces. We recast inference configuration as constrained multi-objective black-box optimization and build InferOpt, a reusable search framework that requires only variable bounds, a deterministic resource cost, and an evaluation hook. InferOpt searches on a frozen sampled proxy set, rejects over-budget candidates before any model call, tightens the budget adaptively, and re-validates Pareto representatives on full-scale dataset. One pipeline covers a 28-dimensional continuous KV space (Qwen2.5-7B) and a 26-dimensional discrete MoE space (DeepSeek-V2-Lite). On KV, post-prefill pruning cuts the 16K cache by 64.4% and TPOT by 22.9--48.5%, and the searched layer-wise budget by InferOpt beats a matched uniform budget by 7.3% and 14.0% of the Full KV reference points. On MoE, a searched top-k schedule by InferOpt removes 43.0% of routed token--expert pairs while staying within 0.59 points of the default, closer than the matched-budget baselines. Against Random Search, NSGA-II, and MOTPE, InferOpt leads on both spaces, taking the best proxy hypervolume and the lowest retention on KV and staying closest to the uncompressed reference at the lowest experts on MoE.