Skip to content
Back to the index

Terminal-Bench-Science 0.1

terminal-bench-science.aiInfrastructure🇺🇸

PulseGate's liveness check found it on 3 Oct 2026; it is registered on GitHub and has been in the index since 28 Aug 2026. How this is checked

Terminal-Bench-Science is a benchmark for evaluating AI agents on challenging scientific research workflows. It provides expert-curated tasks across life, physical, Earth, mathematical, and engineering sciences, with leaderboard-based resolution-rate comparisons for researchers and model developers.

Inferred · not functionally tested

FreeApache-2.0WebCLICloud-managed
324stars
286forks
1alternative
7features
2026since

Overview

6 features

Purpose: Measuring how effectively AI agents complete real scientific research workflows.

Inferred · not functionally tested

Audience: AI researchers and scientists evaluating agent capabilities

Inferred · not functionally tested

Functions: Unknown

Interfaces: API: unknown · MCP: unknown · CLI: indicated (inferred, not tested) · Self-hosting: unknown

Recorded constraints: pricing: free · license: Apache-2.0 · platforms: CLI, WEB · deployment: browser, cli, cloud_managed

Constraint provenance is unknown; confirm requirements with the publisher.

Record sources: terminal-bench-science.ai · github.com. These links do not verify the individual claims.

In the LLM evaluation & benchmarks space, Terminal-Bench-Science 0.1 takes a focused approach. Inferred · not functionally tested: It focuses on measuring how effectively AI agents complete real scientific research workflows. Inferred · not functionally tested: Terminal-Bench-Science 0.1 is an open-source project aimed at AI researchers and scientists evaluating agent capabilities. Basis unknown · not verified: Terminal-Bench-Science 0.1 costs nothing to use. Basis unknown · not verified: It runs on the web and the command line.

Behind Terminal-Bench-Science 0.1 is Stanford University researchers, based in the United States, and it first shipped in 2026. Development happens publicly on GitHub with 324 stars and 702 commits in the last 90 days. Inferred · not functionally tested: Key capabilities include agent benchmarking, scientific workflows, and expert-curated tasks.

Summary written by a language model from the project’s public pages.

Tasks: Inferred · not functionally tested

  • Agent benchmarking
  • Scientific workflows
  • Expert-curated tasks
  • Multi-discipline tasks
  • Leaderboard
  • Continuous evaluation

Topics: Inferred · not functionally tested

Tags
scientific-benchmarkagent-evaluationresearch-workflowsleaderboard
AI capabilities
TextCodeStructured
Inference: Cloud API

JSON profile · Text profile · Access guide

Built with & integrations

Framework
Next.js
Hosting
Vercel
AI providers
multipleanthropicopenai
Runs on
BrowserCLICloud-managed
Detected from
Next.js
x-nextjs-prerender header · /_next/static/ in the HTML · __next_f in the HTML
openai
bgpt- in the HTML
Vercel
x-vercel-id header · x-vercel-cache header
anthropic
bclaude in the HTML

Trust & compliance

License
Apache-2.0
Public signals
HTTPSFree tierGitHub · ★ 324Active maintenance

Indexing history

1

What PulseGate has recorded for this listing

  1. Indexed28 Aug · 03:31 UTC
    Terminal-Bench-Science: Evaluating AI agents on scientific research workflows seen via Hacker News firehose (Algolia)
    Source: Hacker News firehose (Algolia) · Open

Frequently asked questions about Terminal-Bench-Science 0.1

What does Terminal-Bench-Science 0.1 do?
Inferred · not functionally tested: Terminal-Bench-Science 0.1 focuses on measuring how effectively AI agents complete real scientific research workflows. It is catalogued under LLM evaluation & benchmarks on PulseGate.
Who should use Terminal-Bench-Science 0.1?
Inferred · not functionally tested: Terminal-Bench-Science 0.1 is an open-source project built for AI researchers and scientists evaluating agent capabilities.
Does Terminal-Bench-Science 0.1 have a free plan?
Basis unknown · not verified: Yes — Terminal-Bench-Science 0.1 is free to use.
What platforms does Terminal-Bench-Science 0.1 run on?
Basis unknown · not verified: Terminal-Bench-Science 0.1 runs on the web and the command line.
Is Terminal-Bench-Science 0.1 still active?
PulseGate's liveness check found it on 3 Oct 2026. Its GitHub repository shows 702 commits in the last 90 days.
Who makes Terminal-Bench-Science 0.1?
Terminal-Bench-Science 0.1 is developed by Stanford University researchers, based in the United States.
When did Terminal-Bench-Science 0.1 launch?
Terminal-Bench-Science 0.1 first shipped in 2026.
Is Terminal-Bench-Science 0.1 open source?
Basis unknown · not verified: Terminal-Bench-Science 0.1 has a public GitHub repository.

Similar projects

Closest matches by what these projects do