Page Not Found
Page not found. Your pixels are in another canvas.
A list of all the posts and pages found on the site. For you robots out there is an XML version available for digesting as well.
Page not found. Your pixels are in another canvas.
About me
This is a page not in th emain menu
Published in , 2022
A dataset for logical fallacy detection (LOGIC and LOGICCLIMATE) with a structure-aware baseline model
Recommended citation: Jin, Lalwani, A., Vaidhya, T., Shen, X., Ding, Y., Lyu, Z., Sachan, M., Mihalcea, R., & Schölkopf, B. (2022). Logical Fallacy Detection. https://arxiv.org/abs/2202.13758
Published in , 2023
Verbalizing three causal mechanisms of sentiment classification into prompts to study how causal structure affects LLM behavior
Recommended citation: Lyu, Z., Jin, Z., Mattern, J., Mihalcea, R., Sachan, M., & Schoelkopf, B. (2023). Psychologically-Inspired Causal Prompts. arXiv preprint arXiv:2305.01764. https://arxiv.org/pdf/2305.01764
Published in , 2023
Our paper conducts a post-hoc analysis to check whether large language models can be used to distinguish cause from effect.
Recommended citation: Jin, Lalwani, A., Vaidhya, T., Shen, X., Ding, Y., Lyu, Z., Sachan, M., Mihalcea, R., & Schölkopf, B. (2022). Logical Fallacy Detection. https://openreview.net/forum?id=ucHh-ytUkOH
Published in ICLR 2024, 2023
This research introduces the first benchmark dataset, Corr2Cause, to test large language models (LLMs) pure causal inference skills.
Recommended citation: Jin Z, Liu J, Lyu Z, et al. Can Large Language Models Infer Causation from Correlation? arXiv preprint arXiv:2306.05836, 2023. https://arxiv.org/abs/2306.05836
Published in NAACL 2025 (Oral), 2025
A novel approach to tracking dynamic world states and detecting contradictions in story narratives
Recommended citation: Lyu, Z., Yang, K., Kong, L., & Klein, D. (2025). FACTTRACK: Time-Aware World State Tracking in Story Outlines. NAACL 2025. https://arxiv.org/abs/2407.16347
Published in TMLR, 2025
Converting textual reasoning data into images to probe vision-language model reasoning capabilities
Recommended citation: Lyu, Z., Ma, X., & Chen, W. (2025). PixelWorld: Towards Perceiving Everything as Pixels. Transactions on Machine Learning Research. https://arxiv.org/abs/2501.19339
Published in TMLR, 2025
A framework for training web agents that directly interact with browser environments for information-seeking tasks
Recommended citation: Yu, T., Zhang, Z., Lyu, Z., Gong, J., Yi, H., Wang, X., Zhou, Y., Yang, J., Nie, P., Huang, Y., & Chen, W. (2025). BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions. Transactions on Machine Learning Research. https://arxiv.org/abs/2502.01882
Published in Technical Report, 2025
Technical report on MiniMax-M1 model with focus on software engineering capabilities and test-time compute scaling
Recommended citation: MiniMax. (2025). MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention. Technical Report. https://arxiv.org/abs/2506.13585
Published in TMLR 2026 · Best Paper Award @ SPOT Workshop, ICLR 2026, 2025
A unified, easy-to-extend tool-agent RL training framework based on verl, supporting agentic RL across code, search, SQL, and SWE environments.
Recommended citation: Dongfu Jiang*, Yi Lu*, Zhuofeng Li*, Zhiheng Lyu*, Ping Nie, Haozhe Wang, Alex Su, Hui Chen, Kai Zou, Chao Du, Tianyu Pang, Wenhu Chen (2025). VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use. TMLR 2026. https://arxiv.org/abs/2509.01055
Published in ACL 2026 Findings, 2026
A contamination-resistant benchmark for agentic repository-level code understanding, plus a training recipe that lets an 8B model surpass GPT-4o.
Recommended citation: Songcheng Cai*, Zhiheng Lyu*, Yuansheng Ni, Xiangchao Chen, Baichuan Zhou, Shenzhe Zhu, Yi Lu, Haozhe Wang, Chi Ruan, Benjamin Schneider, Weixu Zhang, Xiang Li, Andy Zheng, Yuyu Zhang, Ping Nie, Wenhu Chen (2026). SWE-QA-Pro: A Representative Benchmark and Scalable Training Recipe for Repository-Level Code Understanding. ACL 2026 Findings. https://arxiv.org/abs/2603.16124
Published in Under Review, 2026
An execution-grounded framework for scalable SWE task and trajectory collection: 2,308 verifiable tasks mined from 311 real repositories.
Recommended citation: Jiarong Liang*, Zhiheng Lyu*, Zijie Liu, Xiangchao Chen, Ping Nie, Kai Zou, Wenhu Chen (2026). SWE-Next: Scalable Real-World Software Engineering Tasks for Agents. Under Review. https://arxiv.org/abs/2603.20691
Published in Technical Report, 2026
Technical report on the MiniMax-M2 MoE series (229.9B total / 9.8B active), built end-to-end for agentic deployment. 69% Pass@1 on SWE-bench Verified, #2 on MultiSWE and TerminalBench.
Recommended citation: MiniMax. (2026). The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence. Technical Report. https://arxiv.org/abs/2605.26494
Published in Technical Report, 2026
An open benchmark of 260 contamination-resistant, real-world tasks spanning Code, Web, Office, and Security.
Recommended citation: Tencent WorkBuddy Bench Team, Siqi Cai, Shaopeng Chen, Xiang Fei, Yong Mao, Zihan Xu, Zhiheng Lyu, et al. (2026). Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction. Technical Report. https://arxiv.org/abs/2607.20911