论文

RAPO-Sol:检索增强偏好优化用于仓库级 Solidity 代码生成

RAPO-Sol: Retrieval-Augmented Preference Optimization for Repository-Level Solidity Code Generation

精选理由

一篇做 Solidity 代码生成的训练方法论文,思路挺有意思:训练时检索、推理时不检索,再用语义扰动造 DPO 数据,Qwen2.5-Coder 等模型上验证有效。

RAPO-Sol 是一个两阶段训练框架,面向仓库级 Solidity 智能合约代码生成。第一阶段用 RAFT 在训练时给输入补充相似 Solidity 示例,推理时无需检索。第二阶段用 DPO,通过 Solidity Semantic-Anchor Perturbation 构造rejected样本,扰动校验语句、可见性修饰符、支付操作等内容,让模型区分语义正确的合约与相近但有缺陷的版本。在 SolidityBench 上,RAFT+DPO 流水线使 CodeLlama-7B-Instruct、DeepSeek-Coder-6.7B-Instruct、Qwen2.5-Coder-7B-Instruct 三个模型在 BLEU 和 SolidityScore 上均取得最佳成绩。

原文 · arXiv: DeepSeek

RAPO-Sol: Retrieval-Augmented Preference Optimization for Repository-Level Solidity Code Generation

Smart contracts written in Solidity manage assets, permissions, and irreversible state changes, making code generation both useful and security-critical. Repository-level Solidity generation is challenging because models must synthesize complete contracts or libraries while preserving consistency across state variables, modifiers, events, inheritance, external calls, and access-control logic. We present RAPO-Sol, a two-stage training framework for repository-level Solidity code generation. First, Retrieval-Augmented Fine-Tuning (RAFT) augments each training input with similar Solidity examples, helping the model learn recurring contract-level patterns while remaining retrieval-free at inference time. Second, Direct Preference Optimization (DPO) trains the model to prefer reference contracts over close but semantically flawed alternatives. We construct rejected samples using Solidity Semantic-Anchor Perturbation (SAP), which perturbs validation statements, visibility modifiers, data-location keywords, context variables, payment operations, and low-level calls. Experiments on SolidityBench with CodeLlama-7B-Instruct, DeepSeek-Coder-6.7B-Instruct, and Qwen2.5-Coder-7B-Instruct show that RAFT consistently improves over supervised fine-tuning, while SAP-based DPO provides further gains in BLEU and SolidityScore. The full RAFT+DPO pipeline achieves the best performance across all three models, demonstrating complementary benefits from retrieval during training and Solidity-aware preference optimization without adding retrieval cost at inference.