Feature-Oriented Agentic Coding Benchmark

FeatureBench: Beyond bug fixing. Ship real features.

End-to-end benchmarking of real-world feature development.

Tasks from real repositories, evaluated by executable tests.

Best %Resolved (Full)

--

Loading latest results...

Best %Passed (Full)

--

Loading latest results...

Latest Update

--

Reading local benchmark data...

Collection Pipeline

Scalable, test-driven instance construction from repositories

FeatureBench builds feature-oriented tasks using an execution-based, test-driven pipeline. It selects fail-to-pass and pass-to-pass tests, traces function dependencies, extracts feature patches, and post-verifies each instance to ensure reproducible and continually refreshable evaluation.

  • Dependency graphs are built via dynamic tracing instead of relying solely on static heuristics.
  • Post-verification enforces fail-to-pass and pass-to-pass conditions before patch replay.
  • Problem statements are synthesized with explicit interfaces and import paths.

FeatureBench

Full Set

200 tasks

Hover a slice to inspect repository and count.

Dataset Composition by Repository (Full Split)

Citation

Use one of the following formats when citing FeatureBench.

BibTeX

@article{zhou2026featurebench,
  title={FeatureBench: Benchmarking Agentic Coding for Complex Feature Development},
  author={Zhou, Qixing and Zhang, Jiacheng and Wang, Haiyang and Hao, Rui and Wang, Jiahe and Han, Minghao and Yang, Yuxue and Wu, Shuzhe and Pan, Feiyang and Fan, Lue and others},
  journal={arXiv preprint arXiv:2602.10975},
  year={2026}
}

APA

Zhou, Q., Zhang, J., Wang, H., Hao, R., Wang, J., Han, M., Yang, Y., Wu, S., Pan, F., Fan, L., Tu, D., & Zhang, Z. (2026). FeatureBench: Benchmarking agentic coding for complex feature development. arXiv. https://doi.org/10.48550/arXiv.2602.10975

MLA

Zhou, Qixing, et al. “FeatureBench: Benchmarking Agentic Coding for Complex Feature Development.” arXiv, 2026, https://doi.org/10.48550/arXiv.2602.10975.
Visitors: --