Human Phenotype Project · predictive benchmark

What can each measurement predict?

PhenoBench maps predictive information across 90 clinically grounded tasks, from blood markers and body composition to sleep, glucose dynamics, and longitudinal physiology.

measurement>representation>task>matched evaluation
No universal score. Methods are ranked only within a Benchmark Track that fixes the task, cohort, split, metric, and allowed information.

160 matched cells · 52 tasks

Matched model-family leaderboard

Mean rank compares methods only where task, cohort, feature set, split, and metric are fixed. Lower is better. Averaged within each cell, the three pretrained models improved over ridge by a median of 0.004 R² (95% bootstrap CI, 0.002–0.006).

#MethodMean rankMatched cells
1TabDPT2.34160
2TabICL2.88160
3TabSwift3.17160
4RealMLP3.66160
5Ridge4.11160
6GBDT4.84160

Benchmark atlas

90 clinical tasks across 15 domains

3 framework diagnostics test benchmark mechanics alongside the clinical task set.

Each card links a clinical question to its valid comparison routes. Aggregate result rows contain no participant-level records or private execution identifiers.

Proof of concept

PhenoBench-LLM

Descriptive validation results on 11 shared tasks. Win rate is the fraction of same-task model comparisons won. It does not measure participant-level prediction accuracy. Models fitted on the same fields generally performed better. The all-available analysis remains downloadable with task coverage.

Win rate with task-bootstrap 95% confidence intervals
82.5%Gemini 3.7 Flash
73.4%Claude Opus 4.8
72.7%Claude Sonnet 5
63.3%Gemini 3.1 Pro
62.2%GPT 5.6 Luna
56.6%Gemini 3.1 Flash Lite
47.2%Claude Haiku 4.5
46.9%Nova Pro
46.2%Grok 4.3
39.5%Llama 4 Maverick
32.9%Llama 3.3 70B
30.1%DeepSeek V4 Pro
25.9%Pixtral Large
20.6%GPT 5.4 Nano

Maintainer-run evaluation

Submit a container

Develop against fabricated, schema-faithful examples. Then send an OCI image pinned by digest. Maintainers run accepted images on private evaluation data when capacity allows.

  1. Choose a synthetic task bundle
  2. Package the prediction contract
  3. Open a submission issue with the immutable image digest

Source, Task Cards, and contribution guide on GitHub.

Manual intake. No automated queue, turnaround guarantee, or guarantee of publication.