RISE Humanities Data Benchmark, 0.6.1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'magazine_pages__true' with Search Hidden 'False' returned 178 results, showing page 1 of 18.
Result 1 of 178

Test T1822 at 2026-09-29

newspaper-page document-understanding printed latin 20 en prose, columns

anthropic claude-opus-5-5 · temp 0.0 · dataclass MagazinePage

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

F1100.0%
Precision100.0%
Recall100.0%
Mean IoU97.1%
True positives1
False positives0
False negatives0

{"advertisements":[{"box":[351.0,331.0,2313.0,3135.0]}]}

Scoring

F196.8%
Precision98.2%
Recall96.8%
Mean IoU95.9%
Pages46
True positives120
False positives2
False negatives6

Speed

Model time, all inputs272.5 s
Mean per input5.92 s
Slowest input10.31 s
Inputs timed46

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens216K
Output tokens12.1K
Total tokens228.1K
Input cost$0.8639
Output cost$0.2423
Total cost$1.11

Priced from 2026-09-27 · Input $/M $4.00 · Output $/M $20.00

Result 2 of 178

Test T1882 at 2026-09-29

newspaper-page document-understanding printed latin 20 en prose, columns

anthropic claude-sonnet-5-5 · temp 0.0 · dataclass MagazinePage

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

F1100.0%
Precision100.0%
Recall100.0%
Mean IoU97.2%
True positives1
False positives0
False negatives0

{"advertisements":[{"box":[351,331,2313,3139]}]}

Scoring

F182.5%
Precision84.1%
Recall82.1%
Mean IoU79.7%
Pages46
True positives105
False positives17
False negatives21

Speed

Model time, all inputs152.9 s
Mean per input3.32 s
Slowest input5.83 s
Inputs timed46

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens216K
Output tokens5.3K
Total tokens221.2K
Input cost$0.4319
Output cost$0.0527
Total cost$0.4846

Priced from 2026-09-28 · Input $/M $2.00 · Output $/M $10.00

Result 3 of 178

Test T1807 at 2026-09-29

newspaper-page document-understanding printed latin 20 en prose, columns

scicore Qwen3.8-Flash-Next-FP8 · temp 0.0 · dataclass MagazinePage

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

F1100.0%
Precision100.0%
Recall100.0%
Mean IoU89.4%
True positives1
False positives0
False negatives0

{"advertisements":[{"box":[420.0,330.0,2230.0,3130.0]}]}

Scoring

F171.2%
Precision71.4%
Recall71.8%
Mean IoU70.8%
Pages46
True positives106
False positives17
False negatives20

Speed

Model time, all inputs295.4 s
Mean per input6.42 s
Slowest input37.49 s
Inputs timed46

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens142.6K
Output tokens31.5K
Total tokens174K
Input cost$0.00
Output cost$0.00
Total cost$0.00

Priced from 2026-09-27 · Input $/M $0.00 · Output $/M $0.00

Result 4 of 178

Test T1687 at 2026-09-29

newspaper-page document-understanding printed latin 20 en prose, columns

anthropic claude-fable-5-1 · temp 0.0 · dataclass MagazinePage

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

F1100.0%
Precision100.0%
Recall100.0%
Mean IoU92.5%
True positives1
False positives0
False negatives0

{"advertisements":[{"box":[440.0,330.0,2320.0,3175.0]}]}

Scoring

F196.8%
Precision98.2%
Recall96.8%
Mean IoU94.0%
Pages46
True positives120
False positives2
False negatives6

Speed

Model time, all inputs411.1 s
Mean per input8.94 s
Slowest input12.76 s
Inputs timed46

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens209K
Output tokens19K
Total tokens228K
Input cost$2.09
Output cost$0.9483
Total cost$3.04

Priced from 2026-09-08 · Input $/M $10.00 · Output $/M $50.00

Result 5 of 178

Test T1882 at 2026-09-28

newspaper-page document-understanding printed latin 20 en prose, columns

anthropic claude-sonnet-5-5 · temp 0.0 · dataclass MagazinePage

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

F1100.0%
Precision100.0%
Recall100.0%
Mean IoU97.1%
True positives1
False positives0
False negatives0

{"advertisements":[{"box":[351,331,2313,3135]}]}

Scoring

F164.4%
Precision65.9%
Recall64.2%
Mean IoU64.5%
Pages46
True positives85
False positives3
False negatives41

Speed

Model time, all inputs173.9 s
Mean per input3.78 s
Slowest input5.72 s
Inputs timed46

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens187.5K
Output tokens11.4K
Total tokens198.9K
Input cost$0.375
Output cost$0.1139
Total cost$0.4889

Priced from 2026-09-28 · Input $/M $2.00 · Output $/M $10.00

Result 6 of 178

Test T1852 at 2026-09-27

newspaper-page document-understanding printed latin 20 en prose, columns

openai gpt-6-luna · temp 1.0 · dataclass MagazinePage

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

F1100.0%
Precision100.0%
Recall100.0%
Mean IoU87.8%
True positives1
False positives0
False negatives0

{"advertisements":[{"box":[447.0,333.0,2221.0,3138.0]}]}

Scoring

F196.6%
Precision97.8%
Recall96.8%
Mean IoU95.3%
Pages46
True positives120
False positives3
False negatives6

Speed

Model time, all inputs227.6 s
Mean per input4.95 s
Slowest input13.54 s
Inputs timed46

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens176.6K
Output tokens18.1K
Total tokens194.7K
Input cost$0.0177
Output cost$0.009
Total cost$0.0267

Priced from 2026-09-27 · Input $/M $0.10 · Output $/M $0.50

Result 7 of 178

Test T1867 at 2026-09-27

newspaper-page document-understanding printed latin 20 en prose, columns

x-ai grok-4.7 · temp 0.0 · dataclass MagazinePage

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

F1100.0%
Precision100.0%
Recall100.0%
Mean IoU89.2%
True positives1
False positives0
False negatives0

{"advertisements":[{"box":[420.0,350.0,2235.0,3160.0]}]}

Scoring

F185.9%
Precision87.0%
Recall86.2%
Mean IoU74.3%
Pages46
True positives101
False positives22
False negatives25

Speed

Model time, all inputs2266.5 s
Mean per input49.27 s
Slowest input146.33 s
Inputs timed46

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens183.6K
Output tokens2.2K
Total tokens185.8K
Input cost$0.3673
Output cost$0.0132
Total cost$1.34
Reasoning cost usd$0.961
Total reasoning tokens160.2K

Priced from 2026-09-27 · Input $/M $2.00 · Output $/M $6.00

Result 8 of 178

Test T1837 at 2026-09-27

newspaper-page document-understanding printed latin 20 en prose, columns

openai gpt-6-sol · temp 1.0 · dataclass MagazinePage

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

F1100.0%
Precision100.0%
Recall100.0%
Mean IoU97.4%
True positives1
False positives0
False negatives0

{"advertisements":[{"box":[350.0,333.0,2318.0,3137.0]}]}

Scoring

F195.3%
Precision96.9%
Recall96.8%
Mean IoU96.6%
Pages46
True positives120
False positives24
False negatives6

Speed

Model time, all inputs275.6 s
Mean per input5.99 s
Slowest input24.23 s
Inputs timed46

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens176.6K
Output tokens14K
Total tokens190.7K
Input cost$0.3533
Output cost$0.1401
Total cost$0.4934

Priced from 2026-09-27 · Input $/M $2.00 · Output $/M $10.00

Result 9 of 178

Test T1822 at 2026-09-27

newspaper-page document-understanding printed latin 20 en prose, columns

anthropic claude-opus-5-5 · temp 0.0 · dataclass MagazinePage

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

F1100.0%
Precision100.0%
Recall100.0%
Mean IoU97.0%
True positives1
False positives0
False negatives0

{
    "advertisements": [
        {"box": [352, 330, 2316, 3127]}
    ]
}

Scoring

F196.8%
Precision98.2%
Recall96.8%
Mean IoU95.0%
Pages46
True positives120
False positives2
False negatives6

Speed

Model time, all inputs197.4 s
Mean per input4.29 s
Slowest input9.47 s
Inputs timed46

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens187.5K
Output tokens6K
Total tokens193.5K
Input cost$0.75
Output cost$0.1208
Total cost$0.8708

Priced from 2026-09-27 · Input $/M $4.00 · Output $/M $20.00

Result 10 of 178

Test T0789 at 2026-09-18

newspaper-page document-understanding printed latin 20 en prose, columns

openrouter meta-llama/llama-4-maverick · temp 0.0 · dataclass MagazinePage

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

F10.0%
Precision0.0%
Recall0.0%
Mean IoU0.0%
True positives0
False positives1
False negatives1

{"advertisements":[{"box":[0.177,0.088,0.836,0.912]}]}

Scoring

F10.0%
Precision0.0%
Recall0.0%
Mean IoU0.0%
Pages46
True positives0
False positives119
False negatives126

Speed

Model time, all inputs143.2 s
Mean per input3.11 s
Slowest input16.44 s
Inputs timed46

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens109.2K
Output tokens3.4K
Total tokens112.7K
Input cost$0.0205
Output cost$0.0022
Total cost$0.0227

Priced from 2026-09-16 · Input $/M $0.1875 · Output $/M $0.6525