Best %Resolved (Full)
--
Loading latest results...
Feature-Oriented Agentic Coding Benchmark
End-to-end benchmarking of real-world feature development.
Tasks from real repositories, evaluated by executable tests.
Best %Resolved (Full)
--
Loading latest results...
Best %Passed (Full)
--
Loading latest results...
Latest Update
--
Reading local benchmark data...
Collection Pipeline
FeatureBench builds feature-oriented tasks using an execution-based, test-driven pipeline. It selects fail-to-pass and pass-to-pass tests, traces function dependencies, extracts feature patches, and post-verifies each instance to ensure reproducible and continually refreshable evaluation.
FeatureBench
Full Set
200 tasks
Hover a slice to inspect repository and count.
Citation
Use one of the following formats when citing FeatureBench.
@article{zhou2026featurebench,
title={FeatureBench: Benchmarking Agentic Coding for Complex Feature Development},
author={Zhou, Qixing and Zhang, Jiacheng and Wang, Haiyang and Hao, Rui and Wang, Jiahe and Han, Minghao and Yang, Yuxue and Wu, Shuzhe and Pan, Feiyang and Fan, Lue and others},
journal={arXiv preprint arXiv:2602.10975},
year={2026}
}
Zhou, Q., Zhang, J., Wang, H., Hao, R., Wang, J., Han, M., Yang, Y., Wu, S., Pan, F., Fan, L., Tu, D., & Zhang, Z. (2026). FeatureBench: Benchmarking agentic coding for complex feature development. arXiv. https://doi.org/10.48550/arXiv.2602.10975
Zhou, Qixing, et al. “FeatureBench: Benchmarking Agentic Coding for Complex Feature Development.” arXiv, 2026, https://doi.org/10.48550/arXiv.2602.10975.