toolkit-eval-harness
PulseGate's liveness check found it on 2 Oct 2026; it is registered on GitHub and PyPI and has been in the index since 2 Oct 2026. How this is checked
Toolkit Eval Harness is an open-source package for gating LLM evaluation regressions with signed evidence. It supports paired-bootstrap comparisons, imports results from Promptfoo, Inspect AI, and DeepEval, and produces in-toto report envelopes for CI workflows.
Inferred · not functionally tested
Overview
6 featuresPurpose: Preventing undetected regressions in LLM evaluations and producing verifiable evidence for CI decisions.
Inferred · not functionally tested
Audience: AI and machine learning developers
Inferred · not functionally tested
Functions: analytics, monitoring
Inferred · not functionally tested
Interfaces: API: unknown · MCP: unknown · CLI: indicated (inferred, not tested) · Self-hosting: indicated (inferred, not tested)
Recorded constraints: pricing: open_source · license: Apache-2.0 · platforms: CLI · deployment: cli, self_hosted
Constraint provenance is unknown; confirm requirements with the publisher.
Record sources: pypi.org · github.com. These links do not verify the individual claims.
toolkit-eval-harness is a LLM evaluation & benchmarks project. Inferred · not functionally tested: It focuses on preventing undetected regressions in LLM evaluations and producing verifiable evidence for CI decisions. Inferred · not functionally tested: toolkit-eval-harness is an open-source project aimed at AI and machine learning developers. Basis unknown · not verified: The project is open source (Apache-2.0). Basis unknown · not verified: It ships for the command line, and it can be self-hosted.
AKIVA-AI builds and maintains toolkit-eval-harness, and it first shipped in 2026. Development happens publicly on GitHub with 2 commits in the last 90 days. Inferred · not functionally tested: Among its 6 catalogued features are regression gates, paired bootstrap, and evaluation importers.
Summary written by a language model from the project’s public pages.
Tasks: Inferred · not functionally tested
- Regression gates
- Paired bootstrap
- Evaluation importers
- Signed evidence
- In-toto reports
- CI integration
Topics: Inferred · not functionally tested
Built with & integrations
- Claude Code
- commit 2609b84a7e5d · since Oct 2026
Trust & compliance
Indexing history
1What PulseGate has recorded for this listing
- Indexed2 Oct · 03:08 UTCtoolkit-eval-harness seen via PyPI Bulk EnumeratorSource: PyPI Bulk Enumerator · Open
Frequently asked questions about toolkit-eval-harness
- What does toolkit-eval-harness do?
- Inferred · not functionally tested: Toolkit-eval-harness focuses on preventing undetected regressions in LLM evaluations and producing verifiable evidence for CI decisions. It is catalogued under LLM evaluation & benchmarks on PulseGate.
- Who is toolkit-eval-harness for?
- Inferred · not functionally tested: toolkit-eval-harness is an open-source project built for AI and machine learning developers.
- Is toolkit-eval-harness free?
- Basis unknown · not verified: Yes — toolkit-eval-harness is open source under the Apache-2.0 license and free to use.
- What platforms does toolkit-eval-harness run on?
- Basis unknown · not verified: toolkit-eval-harness runs on the command line. It can also be self-hosted.
- Is toolkit-eval-harness still active?
- PulseGate's liveness check found it on 2 Oct 2026. Its GitHub repository shows 2 commits in the last 90 days.
- Who makes toolkit-eval-harness?
- toolkit-eval-harness is developed by AKIVA-AI.
- When did toolkit-eval-harness launch?
- toolkit-eval-harness first shipped in 2026.
- Is toolkit-eval-harness open source?
- Basis unknown · not verified: Yes — toolkit-eval-harness is open source under the Apache-2.0 license, developed on GitHub.
Similar projects
Closest matches by what these projects do