Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction

Published in Technical Report, 2026

Recommended citation: Tencent WorkBuddy Bench Team, Siqi Cai, Shaopeng Chen, Xiang Fei, Yong Mao, Zihan Xu, Zhiheng Lyu, et al. (2026). Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction. Technical Report. https://arxiv.org/abs/2607.20911

Tencent WorkBuddy Bench is an open benchmark for evaluating coding agents across four domains: Code, Web, Office, and Security. Its 260 tasks are reconstructed from real commits, pull requests, and business scenarios, then rewritten as concise user requests to reduce benchmark contamination.

The release provides domain-specific environments, tests, evaluation tools, and reference solutions. Zhiheng is an equal-contribution author. Ke Li is the project lead.

Paper · Code and evaluation framework · Dataset

Recommended citation:

@article{tencentworkbuddybench2026,
  title={Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction},
  author={{Tencent WorkBuddy Bench Team} and Cai, Siqi and Chen, Shaopeng and Fei, Xiang and Mao, Yong and Xu, Zihan and Lyu, Zhiheng and others},
  journal={arXiv preprint arXiv:2607.20911},
  year={2026}
}