Cohere和LG CNS把111B参数模型压到单卡可跑,还能在韩英混合任务里切换推理模式,做SQL和工具调用都变强了。
Cohere与LG CNS联合发布LuckyStar 111B模型,基于Command A后训练,通过preamble conditioning实现简洁非推理与长工具推理切换。采用多语言监督微调、可验证奖励强化学习、语言一致性奖励和4-bit量化四种策略适配韩英企业代理。在数学推理、函数调用和自然语言转SQL任务上显著提升,同时保持通用韩英指令遵循质量。该模型可在单GPU上运行,为内存受限场景下的多语言代理部署提供了实用方案。
Think in English, Answer in Korean: Efficient Adaptation of Multilingual Tool-Using Agents
We present LuckyStar 111B, a 111B-parameter hybrid reasoning model developed through a collaboration between Cohere and LG CNS for Korean-English enterprise agents under practical memory and serving constraints. The model trains from Cohere's fully post-trained Command A model rather than a new pretraining run, and uses preamble conditioning to switch between concise non-reasoning behavior and longer tool-oriented reasoning. We study four choices for scaling tool-using agents efficiently: multilingual supervised fine-tuning, reinforcement learning with verifiable rewards for multi-step tool-use tasks, language-consistency rewards for Korean user-facing responses, and 4-bit quantization for single-GPU serving. The adapted model improves mathematical reasoning, function calling, and agentic natural-language-to-SQL (NL2SQL) performance while preserving general Korean and English instruction-following quality. These results provide a practical recipe and failure-mode analysis for adapting post-trained multilingual models to verifiable agentic workflows under memory-constrained deployment.