SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning

July 18, 2026 · 19 min · Episode 2057

About this episode

The episode discusses SEED, a framework for improving reinforcement learning through self-evolving on-policy distillation.

🤗 Upvotes: 71 | cs.CL Authors: Jinyang Wu, Shuo Yang, Zhengxi Lu, Fan Zhang, Yuhao Shen, Lang Feng, Haoran Luo, Zheng Lian, Shuai Zhang, Zhengqi Wen, Jianhua Tao Title: SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Arxiv: http://arxiv.org/abs/2607.14777v1 Abstract: Large language models are increasingly trained as interactive agents for long-horizon tasks involving multi-turn interaction, tool use, and environment feedback. Outcome-based reinforcement learning (RL) provides a practical optimization paradigm, but its sparse trajectory-level rewards offer limited guidance on intermediate decisions, leaving a supervision gap between episode-level outcomes and token-level policy learning. We propose SEED (SElf-Evolving On-Policy Distillation), a self-evolving framework that converts completed on-policy trajectories into training-time hindsight skills and distills their behavioral effect back into the policy model. SEED first fine-tunes the policy to analyze completed trajectories and generate natural-language skills that capture reusable workflows, decisive observations, or failure-avoidance rules. During RL, the current policy both collects…

More episodes of Daily Paper Cast

Explore listener stats, chart rankings, contacts and more on the Daily Paper Cast podcast page.