RISE Humanities Data Benchmark, 0.6.0-pre1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'library_cards__true' with Search Hidden 'False' returned 180 results, showing page 3 of 18.
Result 21 of 180

Test T1620 at 2026-09-17

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter meta/muse-spark-1.2 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":" ","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro90.1%
F1 macro89.1%
Micro precision90.3%
Micro recall89.9%
Instances263
True positives2,170
False positives233
False negatives245

Speed

Model time, all inputs10484.2 s
Mean per input39.86 s
Slowest input65.55 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens499.5K
Output tokens762.2K
Total tokens1.3M
Input cost$0.414
Output cost$2.47
Total cost$3.72

Priced from 2026-08-18 · Input $/M $1.25 · Output $/M $4.25

Result 22 of 180

Test T1635 at 2026-09-17

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter z-ai/glm-5v-turbo · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":null,"place":"","pages":"","publisher":"","format":"","editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro82.9%
F1 macro82.4%
Micro precision80.9%
Micro recall85.0%
Instances262
True positives2,044
False positives482
False negatives361

Speed

Model time, all inputs16237.7 s
Mean per input61.74 s
Slowest input131.07 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens523K
Output tokens334K
Total tokens857K
Input cost$0.575
Output cost$1.23
Total cost$1.86

Priced from 2026-08-18 · Input $/M $1.20 · Output $/M $4.00

Result 23 of 180

Test T1001 at 2026-09-17

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter google/gemma-4-26b-a4b-it · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro81.6%
F1 macro81.0%
Micro precision76.7%
Micro recall87.1%
Instances263
True positives2,104
False positives639
False negatives311

Speed

Model time, all inputs2684.4 s
Mean per input10.21 s
Slowest input521.63 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens201.3K
Output tokens53.9K
Total tokens255.2K
Input cost$0.0181
Output cost$0.0162
Total cost$0.0343

Priced from 2026-09-16 · Input $/M $0.09 · Output $/M $0.30

Result 24 of 180

Test T0923 at 2026-09-17

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter qwen/qwen3.5-27b · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghehmi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro85.6%
F1 macro84.8%
Micro precision84.1%
Micro recall87.2%
Instances260
True positives2,082
False positives393
False negatives305

Speed

Model time, all inputs51342.2 s
Mean per input195.96 s
Slowest input1665.15 s
Inputs timed262

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens444.1K
Output tokens317.3K
Total tokens761.4K
Input cost$0.0656
Output cost$0.3265
Total cost$0.676

Priced from 2026-09-16 · Input $/M $0.195 · Output $/M $1.56

Result 25 of 180

Test T1515 at 2026-09-17

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

huggingface Qwen/Qwen3-VL-235B-A22B-Instruct · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro75.9%
F1 macro75.2%
Micro precision71.5%
Micro recall80.9%
Instances263
True positives1,953
False positives780
False negatives462

Speed

Model time, all inputs58644.5 s
Mean per input222.98 s
Slowest input389.03 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens351.3K
Output tokens50K
Total tokens401.3K
Input cost$0.0703
Output cost$0.044
Total cost$0.1142

Priced from 2026-09-16 · Input $/M $0.20 · Output $/M $0.88

Result 26 of 180

Test T0247 at 2026-09-17

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter qwen/qwen3-vl-8b-thinking · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghem i","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro77.9%
F1 macro77.3%
Micro precision75.0%
Micro recall80.9%
Instances263
True positives1,954
False positives650
False negatives461

Speed

Model time, all inputs13066.4 s
Mean per input49.68 s
Slowest input171.27 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens470.7K
Output tokens1.1M
Total tokens1.5M
Input cost$0.0528
Output cost$1.35
Total cost$2.30

Priced from 2026-09-16 · Input $/M $0.18 · Output $/M $2.10

Result 27 of 180

Test T0264 at 2026-09-17

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter qwen/qwen3-vl-8b-instruct · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":""},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro73.3%
F1 macro72.6%
Micro precision74.3%
Micro recall72.4%
Instances263
True positives1,748
False positives605
False negatives667

Speed

Model time, all inputs3277.1 s
Mean per input12.46 s
Slowest input309.82 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens275.8K
Output tokens44K
Total tokens319.8K
Input cost$0.0296
Output cost$0.0191
Total cost$0.0534

Priced from 2026-09-16 · Input $/M $0.117 · Output $/M $0.455

Result 28 of 180

Test T0252 at 2026-09-17

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter meta-llama/llama-4-maverick · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro85.6%
F1 macro84.7%
Micro precision85.0%
Micro recall86.1%
Instances263
True positives2,080
False positives367
False negatives335

Speed

Model time, all inputs71356.0 s
Mean per input271.32 s
Slowest input1393.40 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens516.9K
Output tokens107.7K
Total tokens624.6K
Input cost$0.0496
Output cost$0.0608
Total cost$0.2052

Priced from 2026-09-16 · Input $/M $0.1875 · Output $/M $0.6525

Result 29 of 180

Test T1133 at 2026-09-17

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter stepfun/step-3.7-flash · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro84.1%
F1 macro83.5%
Micro precision82.7%
Micro recall85.6%
Instances262
True positives2,058
False positives432
False negatives347

Speed

Model time, all inputs5136.2 s
Mean per input19.53 s
Slowest input50.96 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens220.2K
Output tokens747.5K
Total tokens967.7K
Input cost$0.0437
Output cost$0.857
Total cost$0.9035

Priced from 2026-09-16 · Input $/M $0.20 · Output $/M $1.15

Result 30 of 180

Test T1410 at 2026-09-17

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

anthropic claude-opus-5 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision0.0%
Recall0.0%
True positives0
False positives4
False negatives3
F1 score0.0%
Total fields23

{"input_data": {"type": {"type": "Reference"}, "author": {"last_name": "Montaghemi", "first_name": "Nourhomadeh"}, "publication": {"title": "Sonnenbichler - Montaghemi, Nourhomadeh", "year": "", "place": "", "pages": "", "publisher": "", "format": ""}, "library_reference": {"shelfmark": "", "subjects": ""}}}

Scoring

F1 micro35.4%
F1 macro34.1%
Micro precision33.9%
Micro recall37.2%
Instances263
True positives898
False positives1,754
False negatives1,517

Speed

Model time, all inputs1201.2 s
Mean per input4.57 s
Slowest input12.03 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens769.4K
Output tokens69.1K
Total tokens838.5K
Input cost$3.85
Output cost$1.73
Total cost$5.57

Priced from 2026-09-16 · Input $/M $5.00 · Output $/M $25.00