Blog

Occasional notes on evaluation, agents, and working with them. Posts are in English; some carry a Chinese version below the fold.

How Far Has RSI Gotten in Post-Training?

Self-refine, self-play, self-evolve, RSI: in post-training one variable separates them, how much of the objective is handed to the agent. Order 0 is a plain code agent. Agents complete order 1 today; order 2 is done by pipelines humans design and agents execute; order 3 has never been done by a machine.

September 2026