How Far Has RSI Gotten in Post-Training?
Published:
Self-refine, self-play, self-evolve, RSI: in post-training one variable separates them, how much of the objective is handed to the agent. Order 0 is a plain code agent. Agents complete order 1 today; order 2 is done by pipelines humans design and agents execute; order 3 has never been done by a machine.
