Posts by Collection

more

publications

Logical Fallacy Detection

Published in , 2022

A dataset for logical fallacy detection (LOGIC and LOGICCLIMATE) with a structure-aware baseline model

Recommended citation: Jin, Lalwani, A., Vaidhya, T., Shen, X., Ding, Y., Lyu, Z., Sachan, M., Mihalcea, R., & Schölkopf, B. (2022). Logical Fallacy Detection. https://arxiv.org/abs/2202.13758

Psychologically-Inspired Causal Prompts.

Published in , 2023

Verbalizing three causal mechanisms of sentiment classification into prompts to study how causal structure affects LLM behavior

Recommended citation: Lyu, Z., Jin, Z., Mattern, J., Mihalcea, R., Sachan, M., & Schoelkopf, B. (2023). Psychologically-Inspired Causal Prompts. arXiv preprint arXiv:2305.01764. https://arxiv.org/pdf/2305.01764

BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions

Published in TMLR, 2025

A framework for training web agents that directly interact with browser environments for information-seeking tasks

Recommended citation: Yu, T., Zhang, Z., Lyu, Z., Gong, J., Yi, H., Wang, X., Zhou, Y., Yang, J., Nie, P., Huang, Y., & Chen, W. (2025). BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions. Transactions on Machine Learning Research. https://arxiv.org/abs/2502.01882

VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use

Published in TMLR 2026 · Best Paper Award @ SPOT Workshop, ICLR 2026, 2025

A unified, easy-to-extend tool-agent RL training framework based on verl, supporting agentic RL across code, search, SQL, and SWE environments.

Recommended citation: Dongfu Jiang*, Yi Lu*, Zhuofeng Li*, Zhiheng Lyu*, Ping Nie, Haozhe Wang, Alex Su, Hui Chen, Kai Zou, Chao Du, Tianyu Pang, Wenhu Chen (2025). VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use. TMLR 2026. https://arxiv.org/abs/2509.01055

SWE-QA-Pro: A Representative Benchmark and Scalable Training Recipe for Repository-Level Code Understanding

Published in ACL 2026 Findings, 2026

A contamination-resistant benchmark for agentic repository-level code understanding, plus a training recipe that lets an 8B model surpass GPT-4o.

Recommended citation: Songcheng Cai*, Zhiheng Lyu*, Yuansheng Ni, Xiangchao Chen, Baichuan Zhou, Shenzhe Zhu, Yi Lu, Haozhe Wang, Chi Ruan, Benjamin Schneider, Weixu Zhang, Xiang Li, Andy Zheng, Yuyu Zhang, Ping Nie, Wenhu Chen (2026). SWE-QA-Pro: A Representative Benchmark and Scalable Training Recipe for Repository-Level Code Understanding. ACL 2026 Findings. https://arxiv.org/abs/2603.16124

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction

Published in Technical Report, 2026

An open benchmark of 260 contamination-resistant, real-world tasks spanning Code, Web, Office, and Security.

Recommended citation: Tencent WorkBuddy Bench Team, Siqi Cai, Shaopeng Chen, Xiang Fei, Yong Mao, Zihan Xu, Zhiheng Lyu, et al. (2026). Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction. Technical Report. https://arxiv.org/abs/2607.20911

statements

teaching