Executable Research Skills
Objectives, stage scaffolds, permission boundaries, launch procedures, and trusted verifiers become a reusable process prior—not a fixed experiment script.
How little human involvement is sufficient for an agent to develop a frontier model? We concentrate expert knowledge into high-density, low-frequency Research Skills, then let an agent evolve data, coordinate SFT, OPSD, and RLVR, and revise the training strategy through executable evidence.
iCoder-27B leads RTLLM, shares the best TritonBench-G result, and ranks second on KernelBench L2 Fast and CVDP.
Industrial coding performance reported in the paper. Higher is better within every benchmark.
Experts provide executable Research Skills once; the agent owns concrete experiments, diagnoses failures, and revises the recipe.
Objectives, stage scaffolds, permission boundaries, launch procedures, and trusted verifiers become a reusable process prior—not a fixed experiment script.
The model learns from failures it can repair with task-specific guidance while the deployed policy remains conditioned on the original task.
Compilation and execution ground the reward. Unjudgeable trajectories are masked; exploit behavior is ineligible before any scalar reward is assigned.
The final recipe was not written in advance. It emerged from failures, controlled ablations, and verifier-grounded evidence.
Train near the moving capability boundary: tasks that are impossible teach little, and mastered tasks stop producing useful variation.
Privileged context should route credit, not become a deployment dependency. Useful feedback changes what is learned while the policy still sees the bare task.
Reward validity comes before reward scale. Kernel exploits and false RTL verdicts forced eligibility gates and fail-closed supervision.
Trajectory budget is part of the learning rule. Length filtering silently removed valuable responses until prompt and response budgets were coupled.
We open-source the agent-developed training trajectory: iCoder-27B-SFT → iCoder-27B-OPSD → iCoder-27B.
Every value below is reproduced from the paper’s main result table, with repeated benchmark labels grouped for faster comparison.
| Benchmark | Metric | iCoder 27B |
Qwen3.6 27B |
InCoder 32B |
InCoder-32B Thinking |
DeepSeek V4-Pro |
GLM 5.2 |
Kimi K2.6 |
GPT 5.5 |
Claude Opus-4.8 |
Hy3 | Gemini 3.5-Flash |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| VerilogEval | Spec-to-RTL avg@4 | 86.3 | 70.1 | 62.5 | 65.9 | 69.9 | 66.0 | 72.4 | 90.1 | 82.7 | 83.8 | 89.1 |
| Code-complete avg@4 | 86.0 | 70.8 | 58.2 | 54.2 | 79.8 | 74.8 | 78.5 | 91.4 | 81.9 | 81.6 | 83.8 | |
| RTLLM | Functional avg@4 | 68.0 | 49.6 | 48.0 | 44.2 | 67.5 | 64.0 | 59.0 | 66.0 | 64.7 | 53.5 | 63.5 |
| CVDP | Functional avg@5 (%) | 44.1 | 33.9 | 36.9 | 30.3 | 38.5 | 39.5 | 42.1 | 39.5 | 47.7 | 39.7 | 29.7 |
| RealBench | Syntax pass@5 (%) | 61.7 | 38.3 | 60.0 | 55.0 | 36.7 | 43.3 | 58.3 | 80.0 | 83.3 | 41.7 | 68.3 |
| Functional pass@5 (%) | 26.7 | 16.7 | 46.7 | 36.7 | 16.7 | 25.0 | 25.0 | 28.3 | 36.7 | 16.7 | 26.7 | |
| ArchXBench | Functional pass@1 (%) | 49.3 | 35.2 | 36.6 | 29.6 | 50.7 | 50.7 | 42.3 | 56.3 | 54.9 | 47.9 | 50.7 |
| KernelBench L1 | Compiled (%) | 95 | 87 | 88 | 85 | 93 | 96 | 93 | 98 | 95 | 94 | 94 |
| Correct (%) | 61 | 32 | 51 | 47 | 32 | 50 | 32 | 43 | 55 | 42 | 45 | |
| Fast (%) | 25 | 12 | 18 | 18 | 13 | 26 | 5 | 22 | 30 | 21 | 23 | |
| KernelBench L2 | Compiled (%) | 97 | 89 | 90 | 93 | 91 | 98 | 84 | 100 | 97 | 98 | 99 |
| Correct (%) | 74 | 28 | 65 | 63 | 40 | 40 | 17 | 41 | 70 | 56 | 78 | |
| Fast (%) | 40 | 17 | 14 | 15 | 25 | 30 | 7 | 24 | 37 | 29 | 47 | |
| KernelBench L3 | Compiled (%) | 90 | 86 | 60 | 60 | 86 | 90 | 82 | 100 | 84 | 98 | 100 |
| Correct (%) | 34 | 12 | 30 | 20 | 4 | 30 | 18 | 38 | 40 | 18 | 58 | |
| Fast (%) | 10 | 4 | 14 | 12 | 2 | 0 | 0 | 6 | 8 | 2 | 14 | |
| TritonBench-G | Correctness pass@1 (%) | 20.1 | 11.4 | 17.9 | 18.5 | 19.0 | 19.0 | 19.0 | 19.5 | 20.1 | 19.5 | 14.9 |
Bold is best; underline is second-best. KernelBench L1/L2 contain 100 tasks and L3 contains 50; L3 counts are reported as percentages.
In evaluation-driven loops, iCoder repeatedly proposes, executes, measures, and refines complete RTL designs and GPU kernels.

