Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction
Published in Technical Report, 2026
Recommended citation: Tencent WorkBuddy Bench Team, Siqi Cai, Shaopeng Chen, Xiang Fei, Yong Mao, Zihan Xu, Zhiheng Lyu, et al. (2026). Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction. Technical Report. https://arxiv.org/abs/2607.20911
Tencent WorkBuddy Bench is an open benchmark for evaluating coding agents across four domains: Code, Web, Office, and Security. Its 260 tasks are reconstructed from real commits, pull requests, and business scenarios, then rewritten as concise user requests to reduce benchmark contamination.
The release provides domain-specific environments, tests, evaluation tools, and reference solutions. Zhiheng is an equal-contribution author. Ke Li is the project lead.
Paper · Code and evaluation framework · Dataset
Recommended citation:
@article{tencentworkbuddybench2026,
title={Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction},
author={{Tencent WorkBuddy Bench Team} and Cai, Siqi and Chen, Shaopeng and Fei, Xiang and Mao, Yong and Xu, Zihan and Lyu, Zhiheng and others},
journal={arXiv preprint arXiv:2607.20911},
year={2026}
}
