LLM-BlockFE:长文本特征工程新框架
Long Text to Predictive Features: LLM-Guided Blockwise Feature Engineering via Executable Program Search
这个新框架让长文本特征工程自动化,比人工设计策略更准,还省了实时调用大模型的成本。
LLM-BlockFE是一种LLM引导的离线特征构建框架,将长文本转换为可执行特征程序,避免在线推理时调用LLM。该方法通过增量追加不可变代码块构建特征程序,使用基于深度校准信用分配的块级回滚机制,交错推进多个独立搜索轨迹。在两个公共和两个私有数据集上,LLM-BlockFE实现AUC绝对提升0.0069至0.0358。在五个金融风控应用中,KS指标绝对提升0.02至1.56个百分点。
Long Text to Predictive Features: LLM-Guided Blockwise Feature Engineering via Executable Program Search
Industrial risk-control systems typically rely on structured-data models for efficient prediction, yet substantial valuable information remains embedded in unstructured long text. Extracting this information through manual feature engineering is labor-intensive, while requiring a large language model (LLM) to process every real-time input may not meet practical deployment requirements. To address this challenge, we propose LLM-BlockFE, an LLM-guided offline feature construction framework that converts long text into executable feature programs, thereby avoiding LLM calls during online inference. LLM-BlockFE constructs feature programs by incrementally appending immutable code blocks and evaluates candidate features using a downstream model. To address the tendency of conventional greedy search to become trapped in suboptimal solutions, our method introduces a block-level rollback mechanism based on depth-calibrated credit allocation and advances multiple independent search trajectories in an interleaved manner, reducing redundant exploration by sharing fixed descriptions of each trajectory's exploration direction. After the search, the resulting programs are frozen and deployed to extract structured features for downstream prediction models. Across two public and two private datasets, LLM-BlockFE achieves absolute AUC improvements of 0.0069 to 0.0358 over the strongest baseline on each dataset in the full-dataset comparison. Post-launch monitoring across five deployed financial risk-control applications shows absolute KS improvements of 0.02 to 1.56 percentage points over the existing manually designed strategy.