real-chart-bench leaderboard (v0)
📌 Latest evaluated set: v0-eval-pilot-n111-llm-full
(most recent run: 2026-09-12T13:29:16.461416+00:00 UTC). Rank is only meaningful within a section
below -- each section is its own figure set (dataset_version), scored runs are
never ranked against a different section's runs, and a higher raw score in a smaller
section does NOT mean that model beat a model ranked #1 in a different section. Always
check which section a row is in (and its own Run at column) before
comparing scores.
⚠️ pre-alpha: evaluation set is a small manually-verified pilot
gated on data/verified_pairs/registry.json (real-figure count varies by run, see each
section's heading below) + 3 synthetic fixtures, not the full v0 dataset. See
docs/experiments/ and docs/design/benchmark-architecture.md
§7.19/§7.21/§7.27 for methodology and known limitations (automatic
image↔figure pairing is unsolved outside the verified registry; naive baselines
cannot see black/gray line series or achromatic markers).
Dataset: v0-eval-pilot-n111 — 111 figures
| Rank | Model | Mean score | #figures |
Run at (UTC) | Breakdown |
| 1 | Naive CV (hue-bucket baseline) | 0.724 | 111 | 2026-09-12T12:31:18.088056+00:00 | by figure type| Type | Mean | # |
|---|
| Real figures (linear x-axis) | 0.724 | 110 |
| Real figures (log x-axis) | 0.775 | 1 |
| Synthetic fixtures | 0.665 | 3 |
|
| 2 | Achromatic CV (grey-level clustering baseline) | 0.651 | 111 | 2026-09-12T12:31:39.958943+00:00 | by figure type| Type | Mean | # |
|---|
| Real figures (linear x-axis) | 0.651 | 110 |
| Real figures (log x-axis) | 0.665 | 1 |
| Synthetic fixtures | 0.811 | 3 |
|
Dataset: v0-eval-pilot-n111-llm-full — 111 figures
| Rank | Model | Mean score | #figures |
Run at (UTC) | Breakdown |
| 1 | Claude Opus 5(採点対象111図・全件) | 0.954 | 111 | 2026-09-12T13:29:16.455931+00:00 | by figure type| Type | Mean | # |
|---|
| Real figures (linear x-axis) | 0.954 | 110 |
| Real figures (log x-axis) | 0.982 | 1 |
|
| 2 | Claude Fable 5(採点対象111図・全件) | 0.952 | 111 | 2026-09-12T13:29:16.459626+00:00 | by figure type| Type | Mean | # |
|---|
| Real figures (linear x-axis) | 0.951 | 110 |
| Real figures (log x-axis) | 0.982 | 1 |
|
| 3 | Claude Sonnet 5(採点対象111図・全件) | 0.937 | 111 | 2026-09-12T13:29:16.457834+00:00 | by figure type| Type | Mean | # |
|---|
| Real figures (linear x-axis) | 0.937 | 110 |
| Real figures (log x-axis) | 0.945 | 1 |
|
| 4 | Claude Haiku 4.5(採点対象111図・全件) | 0.553 | 111 | 2026-09-12T13:29:16.461416+00:00 | by figure type| Type | Mean | # |
|---|
| Real figures (linear x-axis) | 0.554 | 110 |
| Real figures (log x-axis) | 0.449 | 1 |
|
Dataset: v0-eval-pilot-n111-llm-subset-n101-rest — 101 figures
| Rank | Model | Mean score | #figures |
Run at (UTC) | Breakdown |
| 1 | Claude Opus 5(軸レンジなし) | 0.952 | 101 | 2026-09-12T13:28:45.311659+00:00 | by figure type| Type | Mean | # |
|---|
| Real figures (linear x-axis) | 0.952 | 100 |
| Real figures (log x-axis) | 0.982 | 1 |
|
| 2 | Claude Fable 5(軸レンジなし) | 0.948 | 101 | 2026-09-12T13:28:46.099631+00:00 | by figure type| Type | Mean | # |
|---|
| Real figures (linear x-axis) | 0.948 | 100 |
| Real figures (log x-axis) | 0.982 | 1 |
|
| 3 | Claude Sonnet 5(軸レンジなし) | 0.933 | 101 | 2026-09-12T13:28:45.714969+00:00 | by figure type| Type | Mean | # |
|---|
| Real figures (linear x-axis) | 0.933 | 100 |
| Real figures (log x-axis) | 0.945 | 1 |
|
| 4 | Claude Haiku 4.5(軸レンジなし) | 0.522 | 101 | 2026-09-12T13:28:46.908451+00:00 | by figure type| Type | Mean | # |
|---|
| Real figures (linear x-axis) | 0.523 | 100 |
| Real figures (log x-axis) | 0.449 | 1 |
|
Dataset: v0-eval-pilot-n42 — 45 figures
| Rank | Model | Mean score | #figures |
Run at (UTC) | Breakdown |
| 1 | LineFormer (pretrained, ICDAR2023) | 0.647 | 45 | 2026-08-29T02:16:22.185776+00:00 | by figure type| Type | Mean | # |
|---|
| Real figures (linear x-axis) | 0.627 | 41 |
| Real figures (log x-axis) | 0.642 | 1 |
| Synthetic fixtures | 0.917 | 3 |
|
Dataset: v0-eval-pilot-n42-comparable-n38 — 38 figures
| Rank | Model | Mean score | #figures |
Run at (UTC) | Breakdown |
| 1 | LineFormer (pretrained) -- recomputed on the current LineFormer-comparable subset | 0.633 | 38 | 2026-08-29T02:16:22.185776+00:00 | by figure type| Type | Mean | # |
|---|
| Real figures (linear x-axis) | 0.608 | 34 |
| Real figures (log x-axis) | 0.642 | 1 |
| Synthetic fixtures | 0.917 | 3 |
|
| 2 | Naive CV (hue-bucket baseline) -- LineFormer-comparable subset | 0.602 | 38 | 2026-09-12T12:31:18.088056+00:00 | by figure type| Type | Mean | # |
|---|
| Real figures (linear x-axis) | 0.591 | 34 |
| Real figures (log x-axis) | 0.775 | 1 |
| Synthetic fixtures | 0.665 | 3 |
|
Dataset: v0-eval-pilot-n111-llm-subset-n10 — 10 figures
| Rank | Model | Mean score | #figures |
Run at (UTC) | Breakdown |
| 1 | Claude Fable 5 | 0.982 | 10 | 2026-09-12T12:32:57.011417+00:00 | by figure type| Type | Mean | # |
|---|
| Real figures (linear x-axis) | 0.982 | 10 |
|
| 2 | Claude Sonnet 5 | 0.979 | 10 | 2026-09-12T12:32:56.964409+00:00 | by figure type| Type | Mean | # |
|---|
| Real figures (linear x-axis) | 0.979 | 10 |
|
| 3 | Claude Opus 5 | 0.976 | 10 | 2026-09-12T12:32:56.910896+00:00 | by figure type| Type | Mean | # |
|---|
| Real figures (linear x-axis) | 0.976 | 10 |
|
| 4 | Claude Haiku 4.5 | 0.873 | 10 | 2026-09-12T12:32:57.045037+00:00 | by figure type| Type | Mean | # |
|---|
| Real figures (linear x-axis) | 0.873 | 10 |
|
Dataset: v0-eval-pilot-n111-llm-subset-n10-noaxis — 10 figures
| Rank | Model | Mean score | #figures |
Run at (UTC) | Breakdown |
| 1 | Claude Fable 5(軸レンジなし) | 0.880 | 10 | 2026-09-12T12:33:06.082521+00:00 | by figure type| Type | Mean | # |
|---|
| Real figures (linear x-axis) | 0.880 | 10 |
|
| 2 | Claude Opus 5(軸レンジなし) | 0.849 | 10 | 2026-09-12T12:33:05.993145+00:00 | by figure type| Type | Mean | # |
|---|
| Real figures (linear x-axis) | 0.849 | 10 |
|
| 3 | Claude Sonnet 5(軸レンジなし) | 0.848 | 10 | 2026-09-12T12:33:06.035742+00:00 | by figure type| Type | Mean | # |
|---|
| Real figures (linear x-axis) | 0.848 | 10 |
|
| 4 | Claude Haiku 4.5(軸レンジなし) | 0.758 | 10 | 2026-09-12T12:33:06.118658+00:00 | by figure type| Type | Mean | # |
|---|
| Real figures (linear x-axis) | 0.758 | 10 |
|
Dataset: v0-eval-pilot-n111-hue-zero-subset-n8 — 8 figures
| Rank | Model | Mean score | #figures |
Run at (UTC) | Breakdown |
| 1 | Achromatic CV (grey-level clustering baseline) -- restricted to naive-cv-v0's zero-score figures | 0.783 | 8 | 2026-09-12T12:31:39.958943+00:00 | by figure type| Type | Mean | # |
|---|
| Real figures (linear x-axis) | 0.784 | 7 |
| Synthetic fixtures | 0.772 | 1 |
|
Pending -- not yet run
Registered but not yet scored. Never ranked against any section above.
| Model | Status |
| Human ceiling (independent re-digitization agreement) | pending external run — no annotations yet under data/human_ceiling/annotations/ -- run scripts/eval/select_human_ceiling_subset.py to choose the figures to re-digitize, add independent digitizations there (see data/human_ceiling/FORMAT.md), then re-run scripts/eval/compute_human_ceiling.py. |
Head-to-head, same figures only: the sections above are
each scored against a different figure count (see each section's own
heading) -- comparing a row's raw score across sections is misleading. The
one comparison below IS apples-to-apples: both rows are scored on the exact
same 38-figure set (dataset_version
v0-eval-pilot-n42-comparable-n38).
| Model | Mean score | #figures |
| LineFormer (pretrained) -- recomputed on the current LineFormer-comparable subset | 0.633 | 38 |
| Naive CV (hue-bucket baseline) -- LineFormer-comparable subset | 0.602 | 38 |