Research
Health Optimization Bench
84.9
83.8
81.3
78.3
77.9
77.8
77.7
68.4
57.9
53.9
50.1
35.5
33.0
29.2
14.4
7.8
TM
Claude Fable 5.1
Claude Fable 5
Grok 4.6
Claude Opus 5
GPT-5.6 Sol (max)
Kimi K3
GPT-5.6 Sol (high)
Muse Spark
Gemini 3.6
Inkling
Claude Sonnet 5
MiniMax M3
MAI Thinking
GLM 5.2
Mistral Medium 3.5
Nemotron 3.5 Lightning
Health Optimization Bench ranks frontier language models on rubric-graded questions of current clinical evidence, authored and verified across independent model families.
The evaluation harness is open source, and the public task sample is published on Hugging Face.
View more