KernelAgent
Advancing High-performance GPU Kernels Generation with Large-scale Agentic RL System通过大规模 Agentic RL 系统推进高性能 GPU Kernel 生成
- A reliable sandbox with isolated workspaces and rubric-based GRM safeguards against reward hacking.可靠的沙箱环境,结合隔离工作空间与基于评分准则的 GRM 防护,抵御奖励黑客行为。
- A stable and effective RL training framework, combining Ratchet-raw process rewards with a suite of stabilization techniques.稳定高效的强化学习训练框架,将 Ratchet-raw 过程奖励与一系列训练稳定化技术相结合。
- High-throughput infrastructure based on fully asynchronous rollouts and multi-token prediction.基于全异步 rollout 与多 token 预测的高吞吐基础设施。
Scaling Agentic RL to a 1T LLM with 300 turns, 256K context trained on 1,024 GPUs over 150 steps.在 1,024 张 GPU 上训练 150 步,将 Agentic RL 扩展至 1T LLM、300 轮交互与 256K 上下文。
One of the world's largest-scale RL-trained agents.世界上规模最大的 RL 训练智能体之一。
Achieving world-leading GPU kernel generation performance取得世界领先的 GPU Kernel 生成性能
Accomplish in two hours the kernel optimization that takes human experts several days.仅用两小时,完成需要人类专家数天才能完成的 Kernel 优化。














