Research

Health Optimization Bench

84.9
83.8
81.3
78.3
77.9
77.8
77.7
68.4
57.9
53.9
50.1
35.5
33.0
29.2
14.4
7.8
Anthropic logo
Anthropic logo
xAI logo
Anthropic logo
OpenAI logo
Moonshot AI logo
OpenAI logo
Meta logo
Google logo
TM
Anthropic logo
MiniMax logo
Microsoft AI logo
Zhipu logo
Mistral logo
NVIDIA logo
Claude Fable 5.1
Claude Fable 5
Grok 4.6
Claude Opus 5
GPT-5.6 Sol (max)
Kimi K3
GPT-5.6 Sol (high)
Muse Spark
Gemini 3.6
Inkling
Claude Sonnet 5
MiniMax M3
MAI Thinking
GLM 5.2
Mistral Medium 3.5
Nemotron 3.5 Lightning

Health Optimization Bench ranks frontier language models on rubric-graded questions of current clinical evidence, authored and verified across independent model families.

The evaluation harness is open source, and the public task sample is published on Hugging Face.

View more