adversemed-gen is an MIT-licensed Python package that uses language models to generate adversarial medical multiple-choice questions with false premises and abstention-focused answers. It is intended for researchers evaluating calibration, hallucination, and clinical safety in LLMs.
In the Agent evaluation & testing space, adversemed-gen takes a focused approach. It focuses on creating adversarial medical benchmarks to test whether language models recognize uncertainty and abstain safely. adversemed-gen is an open-source project aimed at AI safety researchers and medical LLM evaluators. The project is open source (MIT). It ships for the command line, and it can be self-hosted.
calnugget builds and maintains adversemed-gen, and it first shipped in 2026. The project is developed in the open on GitHub with 9 commits in the last 90 days. Among its 7 catalogued features are question generation, medical benchmarks, and adversarial prompts.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Same category — not a similarity match