Publications

See also my Google Scholar profile.

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction

An open benchmark of 260 contamination-resistant, real-world tasks spanning Code, Web, Office, and Security.

Technical Report · 2026

The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence

Technical report on the MiniMax-M2 MoE series (229.9B total / 9.8B active), built end-to-end for agentic deployment. 69% Pass@1 on SWE-bench Verified, #2 on MultiSWE and TerminalBench.

Technical Report · 2026

SWE-Next: Scalable Real-World Software Engineering Tasks for Agents

An execution-grounded framework for scalable SWE task and trajectory collection: 2,308 verifiable tasks mined from 311 real repositories.

Under Review · 2026

SWE-QA-Pro: A Representative Benchmark and Scalable Training Recipe for Repository-Level Code Understanding

A contamination-resistant benchmark for agentic repository-level code understanding, plus a training recipe that lets an 8B model surpass GPT-4o.

ACL 2026 Findings · 2026

VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use

A unified, easy-to-extend tool-agent RL training framework based on verl, supporting agentic RL across code, search, SQL, and SWE environments.

TMLR 2026 ยท Best Paper Award @ SPOT Workshop, ICLR 2026 · 2025

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Technical report on MiniMax-M1 model with focus on software engineering capabilities and test-time compute scaling

Technical Report · 2025

BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions

A framework for training web agents that directly interact with browser environments for information-seeking tasks

TMLR · 2025

PixelWorld: Towards Perceiving Everything as Pixels

Converting textual reasoning data into images to probe vision-language model reasoning capabilities

TMLR · 2025

FACTTRACK: Time-Aware World State Tracking in Story Outlines

A novel approach to tracking dynamic world states and detecting contradictions in story narratives

NAACL 2025 (Oral) · 2025

Can Large Language Models Infer Causation from Correlation?

This research introduces the first benchmark dataset, Corr2Cause, to test large language models (LLMs) pure causal inference skills.

ICLR 2024 · 2023

Can Large Language Models Distinguish Cause from Effect?

Our paper conducts a post-hoc analysis to check whether large language models can be used to distinguish cause from effect.

2023

Psychologically-Inspired Causal Prompts.

Verbalizing three causal mechanisms of sentiment classification into prompts to study how causal structure affects LLM behavior

2023

Logical Fallacy Detection

A dataset for logical fallacy detection (LOGIC and LOGICCLIMATE) with a structure-aware baseline model

2022