nooa-bench provides a benchmark agent (BenchAgent) and Harbor runner specifically for the NOOA framework. It is designed to reproduce the results from the technical report on SWE-bench and Terminal-Bench. The package enables standardized evaluation of code-generating AI agents through reproducible benchmarking tools.
In the LLM eval & observability space, nooa-bench takes a focused approach. It focuses on reproducing and running standardized benchmarks for code-generating AI agents in the NOOA framework. It is built as an open-source project for developers. The project is open source (Apache-2.0). It ships for the command line.
Behind nooa-bench is NVIDIA NeMo, and it first shipped in 2026. Development happens publicly on GitHub with 512 stars and 109 commits in the last 90 days. Key capabilities include Benchmark Runner, SWE-Bench, and terminal-Bench.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Same category — not a similarity match