论文

RAPO-Sol:检索增强偏好优化框架提升仓库级 Solidity 代码生成

RAPO-Sol: Retrieval-Augmented Preference Optimization for Repository-Level Solidity Code Generation

精选理由

针对智能合约代码生成难保语义一致的问题,这篇论文用 RAFT 加 SAP 扰动构造的 DPO,在 SolidityBench 上让三个 7B 级模型都拿到了最好成绩,做合约生成的人可以看看方法细节。

RAPO-Sol 是一个针对仓库级 Solidity 智能合约代码生成的两阶段训练框架。第一阶段用检索增强微调(RAFT)为训练输入补充相似 Solidity 示例,推理时无需检索。第二阶段通过 DPO 结合语义锚点扰动(SAP)构造拒绝样本,针对校验语句、可见性修饰符、支付操作等维度生成有缺陷的对照合约。在 SolidityBench 上,RAFT+DPO 流水线让 CodeLlama-7B-Instruct、DeepSeek-Coder-6.7B-Instruct、Qwen2.5-Coder-7B-Instruct 三个模型在 BLEU 和 SolidityScore 上均取得最佳成绩。

原文 · arXiv: DeepSeek

RAPO-Sol: Retrieval-Augmented Preference Optimization for Repository-Level Solidity Code Generation

Smart contracts written in Solidity manage assets, permissions, and irreversible state changes, making code generation both useful and security-critical. Repository-level Solidity generation is challenging because models must synthesize complete contracts or libraries while preserving consistency across state variables, modifiers, events, inheritance, external calls, and access-control logic. We present RAPO-Sol, a two-stage training framework for repository-level Solidity code generation. First, Retrieval-Augmented Fine-Tuning (RAFT) augments each training input with similar Solidity examples, helping the model learn recurring contract-level patterns while remaining retrieval-free at inference time. Second, Direct Preference Optimization (DPO) trains the model to prefer reference contracts over close but semantically flawed alternatives. We construct rejected samples using Solidity Semantic-Anchor Perturbation (SAP), which perturbs validation statements, visibility modifiers, data-location keywords, context variables, payment operations, and low-level calls. Experiments on SolidityBench with CodeLlama-7B-Instruct, DeepSeek-Coder-6.7B-Instruct, and Qwen2.5-Coder-7B-Instruct show that RAFT consistently improves over supervised fine-tuning, while SAP-based DPO provides further gains in BLEU and SolidityScore. The full RAFT+DPO pipeline achieves the best performance across all three models, demonstrating complementary benefits from retrieval during training and Solidity-aware preference optimization without adding retrieval cost at inference.