OpenMark is a web-based tool for comparing AI models and benchmarking LLMs on a user’s actual task. It is built around custom benchmarks rather than general-purpose leaderboards, and its interface is organized around creating a task, running tests against selected models, and reviewing results.
The editor supports three task-creation modes: Simple, Advanced, and Manual. In Simple mode, a user describes what they want to test in plain language and the AI agent generates test cases, expected answers, and a scoring configuration. Advanced mode uses structured forms to build tests with prompts, expected answers, attachments, and scoring modes, while Manual mode allows direct YAML editing with full control over fields. The interface also shows validation and preflight steps, task previews, and options to copy from preview or apply edits back to the advanced form.
After a task is prepared, OpenMark lets users select models, configure a benchmark, and run it. The benchmark controls shown in the interface include stability runs, maximum tokens, temperature preferences, timeout profiles, and a fail-fast option. Results can be viewed in a table or chart, sorted, and shared or exported as CSV, JSON, or TXT. The results view displays model-related columns such as score, stability, temperature, pricing, cost, time, accuracy per dollar, accuracy per minute, average output, and completion count. The page also mentions “smart pick” for selecting models, and a quick benchmark flow that can auto-select models and launch a run in one step.
OpenMark also describes uses around monitoring model drift, validating prompt robustness, and preparing fallback options for API issues. It states that benchmarking can help compare cost, speed, and scores for a specific use case, and that its scoring is deterministic and objective rather than subjective. The service includes a guided benchmark, demo, help, billing, and task management areas, and it offers an audit service that will design a test, benchmark it on OpenMark, and send back a model recommendation in 48 hours. The footer identifies the maker as OpenMark AI, Lda., and the site includes terms of service and privacy policy links.
OpenMark sits in PulseGate's LLM eval & observability category. It focuses on comparing and benchmarking multiple AI language models for specific user tasks without relying on generic leaderboards. OpenMark is a B2B product aimed at AI researchers and developers. There is a free tier, and paid plans start at $9. OpenMark is available on the web.
OpenMark AI, Lda builds and maintains OpenMark, and it first shipped in 2026. Development happens publicly on GitHub with 32 commits in the last 90 days. Key capabilities include model comparison, task benchmarking, and API cost analysis.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do