RISE Humanities Data Benchmark, 0.6.1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'library_cards__true' with Search Hidden 'False' returned 188 results, showing page 4 of 19.
Result 31 of 188

Test T0910 at 2026-09-17

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter qwen/qwen3.5-122b-a10b · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro86.5%
F1 macro85.7%
Micro precision84.2%
Micro recall88.9%
Instances263
True positives2,146
False positives403
False negatives269

Speed

Model time, all inputs6535.5 s
Mean per input24.85 s
Slowest input1703.73 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens403.3K
Output tokens329K
Total tokens732.2K
Input cost$0.1052
Output cost$0.6864
Total cost$0.7916

Priced from 2026-09-16 · Input $/M $0.26 · Output $/M $2.08

Result 32 of 188

Test T1395 at 2026-09-17

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter moonshotai/kimi-k3 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":null,"place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro91.8%
F1 macro90.9%
Micro precision90.5%
Micro recall93.1%
Instances263
True positives2,248
False positives236
False negatives167

Speed

Model time, all inputs10686.3 s
Mean per input40.63 s
Slowest input242.38 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens517.9K
Output tokens744.3K
Total tokens1.3M
Input cost$1.37
Output cost$9.89
Total cost$11.26

Priced from 2026-09-16 · Input $/M $2.65 · Output $/M $13.28

Result 33 of 188

Test T0975 at 2026-09-17

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter qwen/qwen3.5-flash-02-23 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro76.8%
F1 macro76.2%
Micro precision74.9%
Micro recall78.8%
Instances261
True positives1,889
False positives634
False negatives507

Speed

Model time, all inputs15126.6 s
Mean per input57.96 s
Slowest input191.39 s
Inputs timed261

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens475.4K
Output tokens1.1M
Total tokens1.6M
Input cost$0.0309
Output cost$0.2864
Total cost$0.3173

Priced from 2026-09-16 · Input $/M $0.065 · Output $/M $0.26

Result 34 of 188

Test T1380 at 2026-09-17

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-5.6-luna · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montagnemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montagnemi, Nourhomadeh","year":"","place":null,"pages":null,"publisher":null,"format":null,"editor":null},"library_reference":{"shelfmark":null,"subjects":null}}

Scoring

F1 micro83.4%
F1 macro82.9%
Micro precision81.7%
Micro recall85.3%
Instances263
True positives2,059
False positives461
False negatives356

Speed

Model time, all inputs2784.1 s
Mean per input10.59 s
Slowest input213.20 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens445.6K
Output tokens118.4K
Total tokens564K
Input cost$0.0891
Output cost$0.1421
Total cost$0.2312

Priced from 2026-09-16 · Input $/M $0.20 · Output $/M $1.20

Result 35 of 188

Test T0897 at 2026-09-17

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter qwen/qwen3.6-plus · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro89.4%
F1 macro88.9%
Micro precision87.8%
Micro recall91.1%
Instances252
True positives2,125
False positives296
False negatives207

Speed

Model time, all inputs13252.4 s
Mean per input52.59 s
Slowest input212.26 s
Inputs timed252

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens451.8K
Output tokens400.8K
Total tokens852.6K
Input cost$0.1468
Output cost$0.7816
Total cost$0.9284

Priced from 2026-09-16 · Input $/M $0.325 · Output $/M $1.95

Result 36 of 188

Test T1665 at 2026-09-17

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-6-astra · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro92.1%
F1 macro91.3%
Micro precision91.1%
Micro recall93.2%
Instances263
True positives2,250
False positives221
False negatives165

Speed

Model time, all inputs1924.1 s
Mean per input7.32 s
Slowest input15.08 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens444.1K
Output tokens71.1K
Total tokens515.3K
Input cost$4.44
Output cost$3.56
Total cost$8.00

Priced from 2026-09-03 · Input $/M $10.00 · Output $/M $50.00

Result 37 of 188

Test T1695 at 2026-09-17

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter meta/muse-spark-1.3 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":null,"place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro91.1%
F1 macro90.2%
Micro precision91.3%
Micro recall91.0%
Instances263
True positives2,197
False positives209
False negatives218

Speed

Model time, all inputs9170.1 s
Mean per input34.87 s
Slowest input122.58 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens458.9K
Output tokens872.1K
Total tokens1.3M
Input cost$0.5716
Output cost$3.71
Total cost$4.28

Priced from 2026-09-08 · Input $/M $1.25 · Output $/M $4.25

Result 38 of 188

Test T0962 at 2026-09-17

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter qwen/qwen3.5-plus-02-15 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro88.5%
F1 macro87.7%
Micro precision86.1%
Micro recall91.0%
Instances256
True positives2,139
False positives346
False negatives211

Speed

Model time, all inputs21348.3 s
Mean per input82.75 s
Slowest input253.59 s
Inputs timed258

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens463.8K
Output tokens906.2K
Total tokens1.4M
Input cost$0.1206
Output cost$1.41
Total cost$1.53

Priced from 2026-09-16 · Input $/M $0.26 · Output $/M $1.56

Result 39 of 188

Test T0247 at 2026-09-17

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter qwen/qwen3-vl-8b-thinking · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghem i","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro77.9%
F1 macro77.3%
Micro precision75.0%
Micro recall80.9%
Instances263
True positives1,954
False positives650
False negatives461

Speed

Model time, all inputs13066.4 s
Mean per input49.68 s
Slowest input171.27 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens470.7K
Output tokens1.1M
Total tokens1.5M
Input cost$0.0847
Output cost$2.22
Total cost$2.30

Priced from 2026-09-16 · Input $/M $0.18 · Output $/M $2.10

Result 40 of 188

Test T0949 at 2026-09-17

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter qwen/qwen3.5-397b-a17b · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"type","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro80.5%
F1 macro80.0%
Micro precision78.4%
Micro recall82.6%
Instances263
True positives1,994
False positives548
False negatives421

Speed

Model time, all inputs27316.0 s
Mean per input103.86 s
Slowest input3608.08 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens282.8K
Output tokens259.5K
Total tokens542.4K
Input cost$0.1544
Output cost$0.8929
Total cost$1.05

Priced from 2026-09-16 · Input $/M $0.55 · Output $/M $3.50