Publications
See also my Google Scholar profile.
An open benchmark of 260 contamination-resistant, real-world tasks spanning Code, Web, Office, and Security.
Technical report on the MiniMax-M2 MoE series (229.9B total / 9.8B active), built end-to-end for agentic deployment. 69% Pass@1 on SWE-bench Verified, #2 on MultiSWE and TerminalBench.
An execution-grounded framework for scalable SWE task and trajectory collection: 2,308 verifiable tasks mined from 311 real repositories.
A contamination-resistant benchmark for agentic repository-level code understanding, plus a training recipe that lets an 8B model surpass GPT-4o.
A unified, easy-to-extend tool-agent RL training framework based on verl, supporting agentic RL across code, search, SQL, and SWE environments.
Technical report on MiniMax-M1 model with focus on software engineering capabilities and test-time compute scaling
A framework for training web agents that directly interact with browser environments for information-seeking tasks
Converting textual reasoning data into images to probe vision-language model reasoning capabilities
A novel approach to tracking dynamic world states and detecting contradictions in story narratives
This research introduces the first benchmark dataset, Corr2Cause, to test large language models (LLMs) pure causal inference skills.
Our paper conducts a post-hoc analysis to check whether large language models can be used to distinguish cause from effect.
Verbalizing three causal mechanisms of sentiment classification into prompts to study how causal structure affects LLM behavior
A dataset for logical fallacy detection (LOGIC and LOGICCLIMATE) with a structure-aware baseline model
