RISE Humanities Data Benchmark, 0.6.0-pre1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'library_cards__true' with Search Hidden 'False' returned 180 results, showing page 5 of 18.
Result 41 of 180

Test T1081 at 2026-09-16

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

anthropic claude-opus-4-8 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"","place":null,"pages":null,"publisher":null,"format":null,"editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro89.5%
F1 macro88.7%
Micro precision86.8%
Micro recall92.3%
Instances263
True positives2,228
False positives338
False negatives187

Speed

Model time, all inputs1017.4 s
Mean per input3.87 s
Slowest input7.54 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens770.4K
Output tokens69.8K
Total tokens840.2K
Input cost$3.85
Output cost$1.74
Total cost$5.60

Priced from 2026-09-16 · Input $/M $5.00 · Output $/M $25.00

Result 42 of 180

Test T1027 at 2026-09-16

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

anthropic claude-opus-4-7 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"","place":null,"pages":null,"publisher":null,"format":null,"editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro89.7%
F1 macro88.8%
Micro precision87.4%
Micro recall92.2%
Instances263
True positives2,227
False positives322
False negatives188

Speed

Model time, all inputs1133.8 s
Mean per input4.31 s
Slowest input11.44 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens875.3K
Output tokens71.7K
Total tokens947K
Input cost$4.38
Output cost$1.79
Total cost$6.17

Priced from 2026-09-16 · Input $/M $5.00 · Output $/M $25.00

Result 43 of 180

Test T1590 at 2026-09-16

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

genai gemini-3.7-flash · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"","place":null,"pages":null,"publisher":null,"format":null,"editor":null},"library_reference":{"shelfmark":null,"subjects":null}}

Scoring

F1 micro88.2%
F1 macro87.3%
Micro precision86.6%
Micro recall90.0%
Instances263
True positives2,173
False positives337
False negatives242

Speed

Model time, all inputs1140.0 s
Mean per input4.33 s
Slowest input20.37 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens406.2K
Output tokens49.4K
Total tokens455.6K
Input cost$0.3047
Output cost$0.1852
Total cost$1.03
Reasoning cost usd$0.5384
Total reasoning tokens143.6K

Priced from 2026-08-18 · Input $/M $0.75 · Output $/M $3.75

Result 44 of 180

Test T1770 at 2026-09-16

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

deepseek deepseek-flash · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"","place":"","pages":"","publisher":"","format":"","editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro86.7%
F1 macro85.9%
Micro precision84.7%
Micro recall88.9%
Instances263
True positives2,146
False positives388
False negatives269

Speed

Model time, all inputs2144.0 s
Mean per input8.15 s
Slowest input41.99 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens425.6K
Output tokens391.4K
Total tokens817K
Input cost$0.1277
Output cost$0.4697
Total cost$0.5974

Priced from 2026-09-16 · Input $/M $0.30 · Output $/M $1.20

Result 45 of 180

Test T0515 at 2026-09-16

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

anthropic claude-opus-4-5-20251101 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision20.0%
Recall33.3%
True positives1
False positives4
False negatives2
F1 score25.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"<UNKNOWN>","place":null,"pages":null,"publisher":null,"format":null,"editor":null},"library_reference":{"shelfmark":null,"subjects":null}}

Scoring

F1 micro86.2%
F1 macro85.4%
Micro precision82.7%
Micro recall90.0%
Instances263
True positives2,173
False positives456
False negatives242

Speed

Model time, all inputs1108.1 s
Mean per input4.21 s
Slowest input7.38 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens706K
Output tokens56K
Total tokens762K
Input cost$3.53
Output cost$1.40
Total cost$4.93

Priced from 2026-09-16 · Input $/M $5.00 · Output $/M $25.00

Result 46 of 180

Test T1650 at 2026-09-16

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

genai gemini-3.8-flash · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro90.6%
F1 macro89.8%
Micro precision89.1%
Micro recall92.1%
Instances263
True positives2,225
False positives271
False negatives190

Speed

Model time, all inputs1470.8 s
Mean per input5.59 s
Slowest input54.37 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens406.2K
Output tokens48.4K
Total tokens454.7K
Input cost$0.3047
Output cost$0.1816
Total cost$1.59
Reasoning cost usd$1.10
Total reasoning tokens293.7K

Priced from 2026-09-02 · Input $/M $0.75 · Output $/M $3.75

Result 47 of 180

Test T0686 at 2026-09-16

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

genai gemini-3.1-pro-preview · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro88.7%
F1 macro87.7%
Micro precision87.2%
Micro recall90.2%
Instances263
True positives2,178
False positives320
False negatives237

Speed

Model time, all inputs2483.2 s
Mean per input9.44 s
Slowest input30.94 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens406.2K
Output tokens49.1K
Total tokens455.3K
Input cost$0.8125
Output cost$0.5892
Total cost$3.77
Reasoning cost usd$2.36
Total reasoning tokens197K

Priced from 2026-09-16 · Input $/M $2.00 · Output $/M $12.00

Result 48 of 180

Test T0633 at 2026-09-16

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

anthropic claude-opus-4-6 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision20.0%
Recall33.3%
True positives1
False positives4
False negatives2
F1 score25.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"<UNKNOWN>","place":null,"pages":null,"publisher":null,"format":null,"editor":null},"library_reference":{"shelfmark":null,"subjects":null}}

Scoring

F1 micro85.6%
F1 macro85.0%
Micro precision81.8%
Micro recall89.8%
Instances263
True positives2,168
False positives483
False negatives247

Speed

Model time, all inputs1737.1 s
Mean per input6.60 s
Slowest input17.50 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens706.2K
Output tokens57.4K
Total tokens763.6K
Input cost$3.53
Output cost$1.43
Total cost$4.97

Priced from 2026-09-16 · Input $/M $5.00 · Output $/M $25.00

Result 49 of 180

Test T0657 at 2026-09-16

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-5.4-2026-03-05 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":null},"library_reference":{"shelfmark":"","subjects":"Sonnenbichler - Montaghemi, Nourhomadeh"}}

Scoring

F1 micro83.9%
F1 macro83.2%
Micro precision81.4%
Micro recall86.5%
Instances263
True positives2,090
False positives476
False negatives325

Speed

Model time, all inputs948.1 s
Mean per input3.60 s
Slowest input67.23 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens444.1K
Output tokens40.2K
Total tokens484.3K
Input cost$1.11
Output cost$0.6027
Total cost$1.71

Priced from 2026-09-16 · Input $/M $2.50 · Output $/M $15.00

Result 50 of 180

Test T1094 at 2026-09-16

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

mistral mistral-medium-3.5 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi, Nourhomadeh","first_name":""},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"","place":null,"pages":null,"publisher":null,"format":null,"editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro83.7%
F1 macro83.1%
Micro precision81.3%
Micro recall86.1%
Instances263
True positives2,080
False positives478
False negatives335

Speed

Model time, all inputs518.0 s
Mean per input1.97 s
Slowest input4.23 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens524.2K
Output tokens44.8K
Total tokens569K
Input cost$0.7864
Output cost$0.3359
Total cost$1.12

Priced from 2026-09-16 · Input $/M $1.50 · Output $/M $7.50