RISE Humanities Data Benchmark, 0.5.5

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'library_cards__true' with Search Hidden 'False' returned 126 results, showing page 1 of 13.
Result 1 of 126

Test T1740 at 2026-09-09

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter qwen/qwen3.8-27b · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision100.0%
Recall100.0%
True positives3
False positives0
False negatives0
F1 score100.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Sonnenbichler","first_name":""},"publication":{"title":"Montaghem i, Nourhomadeh","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro86.0%
F1 macro84.2%
Micro precision85.6%
Micro recall86.5%
Instances263
True positives2,088
False positives352
False negatives327

Speed

Model time, all inputs477274.0 s
Mean per input1814.73 s
Slowest input21985.50 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens460.7K
Output tokens1.3M
Total tokens1.8M
Input cost$0.1266
Output cost$2.54
Total cost$4.11

Priced from 2026-09-08 · Input $/M $0.42 · Output $/M $3.00

Result 2 of 126

Test T1710 at 2026-09-08

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter z-ai/glm-5.3-flash · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghem","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro79.2%
F1 macro78.2%
Micro precision78.2%
Micro recall80.2%
Instances263
True positives1,937
False positives539
False negatives478

Speed

Model time, all inputs7443.4 s
Mean per input28.30 s
Slowest input941.20 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens323.8K
Output tokens66.6K
Total tokens390.5K
Input cost$0.0232
Output cost$0.0149
Total cost$0.0411

Priced from 2026-09-08 · Input $/M $0.075 · Output $/M $0.25

Result 3 of 126

Test T1725 at 2026-09-08

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter qwen/qwen3.8-flash · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision20.0%
Recall33.3%
True positives1
False positives4
False negatives2
F1 score25.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"S","year":"","place":"","pages":"","publisher":"","format":"","editor":null},"library_reference":{"shelfmark":"","subjects":"Sonnenbichler - Montaghemi, Nourhomadeh"}}

Scoring

F1 micro85.3%
F1 macro83.3%
Micro precision84.2%
Micro recall86.4%
Instances258
True positives2,046
False positives384
False negatives321

Speed

Model time, all inputs32273.5 s
Mean per input122.71 s
Slowest input955.87 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens484.4K
Output tokens848.8K
Total tokens1.3M
Input cost$0.0262
Output cost$0.2523
Total cost$0.4649

Priced from 2026-09-08 · Input $/M $0.15 · Output $/M $0.47

Result 4 of 126

Test T1695 at 2026-09-08

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter meta/muse-spark-1.3 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro91.1%
F1 macro90.2%
Micro precision91.3%
Micro recall90.9%
Instances263
True positives2,196
False positives210
False negatives219

Speed

Model time, all inputs5234.2 s
Mean per input19.90 s
Slowest input47.48 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens454K
Output tokens630.8K
Total tokens1.1M
Input cost$0.5675
Output cost$2.68
Total cost$3.25

Priced from 2026-09-08 · Input $/M $1.25 · Output $/M $4.25

Result 5 of 126

Test T1755 at 2026-09-08

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

deepseek deepseek-v4-flash-vision-exp · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro88.4%
F1 macro88.0%
Micro precision87.5%
Micro recall89.4%
Instances263
True positives2,158
False positives309
False negatives257

Speed

Model time, all inputs4906.1 s
Mean per input18.65 s
Slowest input52.93 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens429.7K
Output tokens523.6K
Total tokens953.3K
Input cost$0.1891
Output cost$0.6912
Total cost$0.8802

Priced from 2026-09-08 · Input $/M $0.44 · Output $/M $1.32

Result 6 of 126

Test T1665 at 2026-09-04

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-6-astra · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro91.7%
F1 macro90.9%
Micro precision90.8%
Micro recall92.6%
Instances263
True positives2,237
False positives228
False negatives178

Speed

Model time, all inputs1565.3 s
Mean per input5.95 s
Slowest input15.14 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens444.1K
Output tokens73K
Total tokens517.2K
Input cost$4.44
Output cost$3.65
Total cost$8.09

Priced from 2026-09-03 · Input $/M $10.00 · Output $/M $50.00

Result 7 of 126

Test T1650 at 2026-09-03

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

genai gemini-3.8-flash · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"","place":"","pages":"","publisher":"","format":null,"editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro90.5%
F1 macro89.6%
Micro precision89.0%
Micro recall91.9%
Instances263
True positives2,220
False positives273
False negatives195

Speed

Model time, all inputs1726.5 s
Mean per input6.56 s
Slowest input113.52 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens406.2K
Output tokens48.8K
Total tokens455K
Input cost$0.3047
Output cost$0.183
Total cost$0.4876

Priced from 2026-09-02 · Input $/M $0.75 · Output $/M $3.75

Result 8 of 126

Test T1620 at 2026-08-18

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter meta/muse-spark-1.2 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro90.3%
F1 macro89.2%
Micro precision90.5%
Micro recall90.2%
Instances262
True positives2,170
False positives229
False negatives237

Speed

Model time, all inputs128221.0 s
Mean per input487.53 s
Slowest input1370.89 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens489.3K
Output tokens979K
Total tokens1.5M
Input cost$0.6117
Output cost$4.16
Total cost$4.77

Priced from 2026-08-18 · Input $/M $1.25 · Output $/M $4.25

Result 9 of 126

Test T1590 at 2026-08-18

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

genai gemini-3.7-flash · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"","place":"","pages":"","publisher":"","format":"","editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro88.4%
F1 macro87.5%
Micro precision86.7%
Micro recall90.2%
Instances263
True positives2,178
False positives333
False negatives237

Speed

Model time, all inputs910.4 s
Mean per input3.46 s
Slowest input10.15 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens406.2K
Output tokens49.6K
Total tokens455.8K
Input cost$0.3047
Output cost$0.1858
Total cost$0.4905

Priced from 2026-08-18 · Input $/M $0.75 · Output $/M $3.75

Result 10 of 126

Test T1635 at 2026-08-18

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter z-ai/glm-5v-turbo · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{
  "type": {
    "type": "Reference"
  },
  "author": {
    "last_name": "Montaghemi",
    "first_name": "Nourhomadeh"
  },
  "publication": {
    "title": "Sonnenbichler - Montaghemi, Nourhomadeh",
    "place": "",
    "pages": "",
    "publisher": "",
    "format": ""
  },
  "library_reference": {
    "shelfmark": "",
    "subjects": ""
  }
}

Scoring

F1 micro84.8%
F1 macro84.2%
Micro precision82.9%
Micro recall86.8%
Instances261
True positives2,079
False positives428
False negatives317

Speed

Model time, all inputs9303.7 s
Mean per input35.38 s
Slowest input66.46 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens519.6K
Output tokens324.4K
Total tokens843.9K
Input cost$0.6235
Output cost$1.30
Total cost$1.92

Priced from 2026-08-18 · Input $/M $1.20 · Output $/M $4.00