RISE Humanities Data Benchmark, 0.6.1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'library_cards__true' with Search Hidden 'False' returned 187 results, showing page 6 of 19.
Result 51 of 187

Test T1425 at 2026-09-16

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

genai gemini-3.6-flash · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"","place":"","pages":"","publisher":"","format":"","editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro88.1%
F1 macro87.0%
Micro precision86.3%
Micro recall89.9%
Instances263
True positives2,170
False positives344
False negatives245

Speed

Model time, all inputs2070.2 s
Mean per input7.87 s
Slowest input16.43 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens406.2K
Output tokens40.5K
Total tokens446.8K
Input cost$0.3047
Output cost$0.152
Total cost$1.67
Reasoning cost usd$1.22
Total reasoning tokens324.6K

Priced from 2026-09-16 · Input $/M $0.75 · Output $/M $3.75

Result 52 of 187

Test T1605 at 2026-09-16

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

x-ai grok-4.6 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":null,"pages":null,"publisher":null,"format":null,"editor":null},"library_reference":{"shelfmark":null,"subjects":null}}

Scoring

F1 micro91.6%
F1 macro90.8%
Micro precision90.5%
Micro recall92.7%
Instances263
True positives2,239
False positives234
False negatives176

Speed

Model time, all inputs10404.0 s
Mean per input39.56 s
Slowest input108.22 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens636.2K
Output tokens32.7K
Total tokens668.9K
Input cost$1.27
Output cost$0.1962
Total cost$5.15
Reasoning cost usd$3.68
Total reasoning tokens614.1K

Priced from 2026-08-18 · Input $/M $2.00 · Output $/M $6.00

Result 53 of 187

Test T0657 at 2026-09-16

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-5.4-2026-03-05 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":null},"library_reference":{"shelfmark":"","subjects":"Sonnenbichler - Montaghemi, Nourhomadeh"}}

Scoring

F1 micro83.9%
F1 macro83.2%
Micro precision81.4%
Micro recall86.5%
Instances263
True positives2,090
False positives476
False negatives325

Speed

Model time, all inputs948.1 s
Mean per input3.60 s
Slowest input67.23 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens444.1K
Output tokens40.2K
Total tokens484.3K
Input cost$1.11
Output cost$0.6027
Total cost$1.71

Priced from 2026-09-16 · Input $/M $2.50 · Output $/M $15.00

Result 54 of 187

Test T1365 at 2026-09-16

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-5.6-terra · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler-Montaghemi, Nourhomadeh","year":"","place":"","pages":"","publisher":"","format":"","editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro76.8%
F1 macro76.2%
Micro precision75.9%
Micro recall77.8%
Instances263
True positives1,878
False positives597
False negatives537

Speed

Model time, all inputs7343.0 s
Mean per input27.92 s
Slowest input802.34 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens444.1K
Output tokens172.7K
Total tokens616.8K
Input cost$0.8883
Output cost$2.07
Total cost$2.96

Priced from 2026-09-16 · Input $/M $2.00 · Output $/M $12.00

Result 55 of 187

Test T1770 at 2026-09-16

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

deepseek deepseek-flash · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"","place":"","pages":"","publisher":"","format":"","editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro86.7%
F1 macro85.9%
Micro precision84.7%
Micro recall88.9%
Instances263
True positives2,146
False positives388
False negatives269

Speed

Model time, all inputs2144.0 s
Mean per input8.15 s
Slowest input41.99 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens425.6K
Output tokens391.4K
Total tokens817K
Input cost$0.1277
Output cost$0.4697
Total cost$0.5974

Priced from 2026-09-16 · Input $/M $0.30 · Output $/M $1.20

Result 56 of 187

Test T1455 at 2026-09-16

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

x-ai grok-4.5 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro91.1%
F1 macro90.3%
Micro precision90.1%
Micro recall92.1%
Instances263
True positives2,224
False positives244
False negatives191

Speed

Model time, all inputs5475.3 s
Mean per input20.82 s
Slowest input70.78 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens598.9K
Output tokens37.2K
Total tokens636.1K
Input cost$1.20
Output cost$0.2232
Total cost$3.22
Reasoning cost usd$1.80
Total reasoning tokens300.2K

Priced from 2026-09-16 · Input $/M $2.00 · Output $/M $6.00

Result 57 of 187

Test T1169 at 2026-09-16

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

genai gemini-3.1-flash-lite · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"","place":"","pages":"","publisher":"","format":"","editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro88.7%
F1 macro87.9%
Micro precision85.5%
Micro recall92.0%
Instances263
True positives2,223
False positives377
False negatives192

Speed

Model time, all inputs357.9 s
Mean per input1.36 s
Slowest input2.83 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens406.2K
Output tokens47.9K
Total tokens454.2K
Input cost$0.1016
Output cost$0.0719
Total cost$0.1734

Priced from 2026-09-16 · Input $/M $0.25 · Output $/M $1.50

Result 58 of 187

Test T1440 at 2026-09-16

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

genai gemini-3.5-flash-lite · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"s. Sonnenbichler - Montaghemi, Nourhomadeh","year":"","place":null,"pages":null,"publisher":null,"format":null,"editor":null},"library_reference":{"shelfmark":null,"subjects":null}}

Scoring

F1 micro87.4%
F1 macro86.7%
Micro precision84.5%
Micro recall90.6%
Instances263
True positives2,187
False positives402
False negatives228

Speed

Model time, all inputs339.9 s
Mean per input1.29 s
Slowest input2.80 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens406.2K
Output tokens49.3K
Total tokens455.5K
Input cost$0.1219
Output cost$0.1232
Total cost$0.2451

Priced from 2026-09-16 · Input $/M $0.30 · Output $/M $2.50

Result 59 of 187

Test T1650 at 2026-09-16

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

genai gemini-3.8-flash · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro90.6%
F1 macro89.8%
Micro precision89.1%
Micro recall92.1%
Instances263
True positives2,225
False positives271
False negatives190

Speed

Model time, all inputs1470.8 s
Mean per input5.59 s
Slowest input54.37 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens406.2K
Output tokens48.4K
Total tokens454.7K
Input cost$0.3047
Output cost$0.1816
Total cost$1.59
Reasoning cost usd$1.10
Total reasoning tokens293.7K

Priced from 2026-09-02 · Input $/M $0.75 · Output $/M $3.75

Result 60 of 187

Test T1055 at 2026-09-16

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

genai gemini-3.5-flash · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"","place":null,"pages":null,"publisher":null,"format":null,"editor":null},"library_reference":{"shelfmark":null,"subjects":null}}

Scoring

F1 micro86.9%
F1 macro85.9%
Micro precision84.0%
Micro recall90.1%
Instances263
True positives2,175
False positives413
False negatives240

Speed

Model time, all inputs1713.1 s
Mean per input6.51 s
Slowest input10.45 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens406.2K
Output tokens49.9K
Total tokens456.1K
Input cost$0.6094
Output cost$0.4489
Total cost$4.04
Reasoning cost usd$2.98
Total reasoning tokens331.3K

Priced from 2026-09-16 · Input $/M $1.50 · Output $/M $9.00