real-chart-bench leaderboard (v0)

📌 Latest evaluated set: v0-eval-pilot-n111-llm-full (most recent run: 2026-09-12T13:29:16.461416+00:00 UTC). Rank is only meaningful within a section below -- each section is its own figure set (dataset_version), scored runs are never ranked against a different section's runs, and a higher raw score in a smaller section does NOT mean that model beat a model ranked #1 in a different section. Always check which section a row is in (and its own Run at column) before comparing scores.

⚠️ pre-alpha: evaluation set is a small manually-verified pilot gated on data/verified_pairs/registry.json (real-figure count varies by run, see each section's heading below) + 3 synthetic fixtures, not the full v0 dataset. See docs/experiments/ and docs/design/benchmark-architecture.md §7.19/§7.21/§7.27 for methodology and known limitations (automatic image↔figure pairing is unsolved outside the verified registry; naive baselines cannot see black/gray line series or achromatic markers).

Dataset: v0-eval-pilot-n111 — 111 figures

RankModelMean score#figures Run at (UTC)Breakdown
1Naive CV (hue-bucket baseline)0.7241112026-09-12T12:31:18.088056+00:00
by figure type
TypeMean#
Real figures (linear x-axis)0.724110
Real figures (log x-axis)0.7751
Synthetic fixtures0.6653
2Achromatic CV (grey-level clustering baseline)0.6511112026-09-12T12:31:39.958943+00:00
by figure type
TypeMean#
Real figures (linear x-axis)0.651110
Real figures (log x-axis)0.6651
Synthetic fixtures0.8113

Dataset: v0-eval-pilot-n111-llm-full — 111 figures

RankModelMean score#figures Run at (UTC)Breakdown
1Claude Opus 5(採点対象111図・全件)0.9541112026-09-12T13:29:16.455931+00:00
by figure type
TypeMean#
Real figures (linear x-axis)0.954110
Real figures (log x-axis)0.9821
2Claude Fable 5(採点対象111図・全件)0.9521112026-09-12T13:29:16.459626+00:00
by figure type
TypeMean#
Real figures (linear x-axis)0.951110
Real figures (log x-axis)0.9821
3Claude Sonnet 5(採点対象111図・全件)0.9371112026-09-12T13:29:16.457834+00:00
by figure type
TypeMean#
Real figures (linear x-axis)0.937110
Real figures (log x-axis)0.9451
4Claude Haiku 4.5(採点対象111図・全件)0.5531112026-09-12T13:29:16.461416+00:00
by figure type
TypeMean#
Real figures (linear x-axis)0.554110
Real figures (log x-axis)0.4491

Dataset: v0-eval-pilot-n111-llm-subset-n101-rest — 101 figures

RankModelMean score#figures Run at (UTC)Breakdown
1Claude Opus 5(軸レンジなし)0.9521012026-09-12T13:28:45.311659+00:00
by figure type
TypeMean#
Real figures (linear x-axis)0.952100
Real figures (log x-axis)0.9821
2Claude Fable 5(軸レンジなし)0.9481012026-09-12T13:28:46.099631+00:00
by figure type
TypeMean#
Real figures (linear x-axis)0.948100
Real figures (log x-axis)0.9821
3Claude Sonnet 5(軸レンジなし)0.9331012026-09-12T13:28:45.714969+00:00
by figure type
TypeMean#
Real figures (linear x-axis)0.933100
Real figures (log x-axis)0.9451
4Claude Haiku 4.5(軸レンジなし)0.5221012026-09-12T13:28:46.908451+00:00
by figure type
TypeMean#
Real figures (linear x-axis)0.523100
Real figures (log x-axis)0.4491

Dataset: v0-eval-pilot-n42 — 45 figures

RankModelMean score#figures Run at (UTC)Breakdown
1LineFormer (pretrained, ICDAR2023)0.647452026-08-29T02:16:22.185776+00:00
by figure type
TypeMean#
Real figures (linear x-axis)0.62741
Real figures (log x-axis)0.6421
Synthetic fixtures0.9173

Dataset: v0-eval-pilot-n42-comparable-n38 — 38 figures

RankModelMean score#figures Run at (UTC)Breakdown
1LineFormer (pretrained) -- recomputed on the current LineFormer-comparable subset0.633382026-08-29T02:16:22.185776+00:00
by figure type
TypeMean#
Real figures (linear x-axis)0.60834
Real figures (log x-axis)0.6421
Synthetic fixtures0.9173
2Naive CV (hue-bucket baseline) -- LineFormer-comparable subset0.602382026-09-12T12:31:18.088056+00:00
by figure type
TypeMean#
Real figures (linear x-axis)0.59134
Real figures (log x-axis)0.7751
Synthetic fixtures0.6653

Dataset: v0-eval-pilot-n111-llm-subset-n10 — 10 figures

RankModelMean score#figures Run at (UTC)Breakdown
1Claude Fable 50.982102026-09-12T12:32:57.011417+00:00
by figure type
TypeMean#
Real figures (linear x-axis)0.98210
2Claude Sonnet 50.979102026-09-12T12:32:56.964409+00:00
by figure type
TypeMean#
Real figures (linear x-axis)0.97910
3Claude Opus 50.976102026-09-12T12:32:56.910896+00:00
by figure type
TypeMean#
Real figures (linear x-axis)0.97610
4Claude Haiku 4.50.873102026-09-12T12:32:57.045037+00:00
by figure type
TypeMean#
Real figures (linear x-axis)0.87310

Dataset: v0-eval-pilot-n111-llm-subset-n10-noaxis — 10 figures

RankModelMean score#figures Run at (UTC)Breakdown
1Claude Fable 5(軸レンジなし)0.880102026-09-12T12:33:06.082521+00:00
by figure type
TypeMean#
Real figures (linear x-axis)0.88010
2Claude Opus 5(軸レンジなし)0.849102026-09-12T12:33:05.993145+00:00
by figure type
TypeMean#
Real figures (linear x-axis)0.84910
3Claude Sonnet 5(軸レンジなし)0.848102026-09-12T12:33:06.035742+00:00
by figure type
TypeMean#
Real figures (linear x-axis)0.84810
4Claude Haiku 4.5(軸レンジなし)0.758102026-09-12T12:33:06.118658+00:00
by figure type
TypeMean#
Real figures (linear x-axis)0.75810

Dataset: v0-eval-pilot-n111-hue-zero-subset-n8 — 8 figures

RankModelMean score#figures Run at (UTC)Breakdown
1Achromatic CV (grey-level clustering baseline) -- restricted to naive-cv-v0's zero-score figures0.78382026-09-12T12:31:39.958943+00:00
by figure type
TypeMean#
Real figures (linear x-axis)0.7847
Synthetic fixtures0.7721

Pending -- not yet run

Registered but not yet scored. Never ranked against any section above.

ModelStatus
Human ceiling (independent re-digitization agreement)pending external run — no annotations yet under data/human_ceiling/annotations/ -- run scripts/eval/select_human_ceiling_subset.py to choose the figures to re-digitize, add independent digitizations there (see data/human_ceiling/FORMAT.md), then re-run scripts/eval/compute_human_ceiling.py.
Head-to-head, same figures only: the sections above are each scored against a different figure count (see each section's own heading) -- comparing a row's raw score across sections is misleading. The one comparison below IS apples-to-apples: both rows are scored on the exact same 38-figure set (dataset_version v0-eval-pilot-n42-comparable-n38).
ModelMean score#figures
LineFormer (pretrained) -- recomputed on the current LineFormer-comparable subset0.63338
Naive CV (hue-bucket baseline) -- LineFormer-comparable subset0.60238