RISE Humanities Data Benchmark, 0.6.1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'library_cards__true' with Search Hidden 'False' returned 188 results, showing page 16 of 19.
Result 151 of 188

Test T0161 at 2026-01-24

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-4.1-mini-2025-04-14 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision0.0%
Recall0.0%
True positives0
False positives0
False negatives3
F1 score0.0%
Total fields13

{
  "type": {
    "type": "Reference"
  },
  "author": {
    "last_name": "Montaghemi",
    "first_name": "Nourhomadeh"
  },
  "publication": {
    "title": "",
    "year": null,
    "place": "",
    "pages": "",
    "publisher": "",
    "format": ""
  },
  "library_reference": {
    "shelfmark": "",
    "subjects": ""
  }
}

Scoring

F1 micro58.2%
F1 macro46.5%
Micro precision76.2%
Micro recall47.0%
Instances263
True positives1,136
False positives355
False negatives1,279

Speed

Model time, all inputs49684.2 s
Mean per input188.91 s
Slowest input1118.96 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens531.9K
Output tokens72.3K
Total tokens604.2K
Input cost$0.2127
Output cost$0.1157
Total cost$0.3285

Priced from 2026-01-23 · Input $/M $0.40 · Output $/M $1.60

Result 152 of 188

Test T0526 at 2026-01-24

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

anthropic claude-haiku-4-5-20251001 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler-Montaghemi, Nourhomadeh","year":"","place":null,"pages":null,"publisher":null,"format":null,"editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro84.4%
F1 macro83.5%
Micro precision83.2%
Micro recall85.7%
Instances263
True positives2,069
False positives418
False negatives346

Speed

Model time, all inputs572.3 s
Mean per input2.18 s
Slowest input9.42 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens706K
Output tokens53.1K
Total tokens759.1K
Input cost$0.706
Output cost$0.2655
Total cost$0.9714

Priced from 2026-01-23 · Input $/M $1.00 · Output $/M $5.00

Result 153 of 188

Test T0247 at 2026-01-24

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter qwen/qwen3-vl-8b-thinking · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision0.0%
Recall0.0%
True positives0
False positives0
False negatives3
F1 score0.0%
Total fields13

Scoring

F1 micro17.5%
F1 macro9.9%
Micro precision71.6%
Micro recall9.9%
Instances263
True positives240
False positives95
False negatives2,175

Speed

Model time, all inputs6748.5 s
Mean per input25.66 s
Slowest input622.78 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens509.2K
Output tokens96K
Total tokens605.2K
Input cost$0.0917
Output cost$0.2016
Total cost$0.2933

Priced from 2026-01-23 · Input $/M $0.18 · Output $/M $2.10

Result 154 of 188

Test T0258 at 2026-01-24

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter qwen/qwen3-vl-30b-a3b-instruct · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"MontaghemI","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - MontaghemI, Nourhomadeh","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro61.3%
F1 macro53.6%
Micro precision66.4%
Micro recall56.8%
Instances263
True positives1,372
False positives693
False negatives1,043

Speed

Model time, all inputs18678.2 s
Mean per input71.02 s
Slowest input608.05 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens345.3K
Output tokens43.7K
Total tokens389K
Input cost$0.0551
Output cost$0.027
Total cost$0.0821

Priced from 2026-01-23 · Input $/M $0.15 · Output $/M $0.60

Result 155 of 188

Test T0066 at 2026-01-24

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-4o-2024-08-06 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro85.7%
F1 macro85.0%
Micro precision86.1%
Micro recall85.3%
Instances263
True positives2,059
False positives332
False negatives356

Speed

Model time, all inputs3615.8 s
Mean per input13.75 s
Slowest input572.35 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens462.1K
Output tokens30.8K
Total tokens492.8K
Input cost$1.16
Output cost$0.3077
Total cost$1.46

Priced from 2026-01-23 · Input $/M $2.50 · Output $/M $10.00

Result 156 of 188

Test T0230 at 2026-01-24

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

anthropic claude-sonnet-4-5-20250929 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"<UNKNOWN>","place":"","pages":"","publisher":"","format":"","editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro85.8%
F1 macro84.7%
Micro precision84.2%
Micro recall87.5%
Instances263
True positives2,112
False positives395
False negatives303

Speed

Model time, all inputs1230.6 s
Mean per input4.68 s
Slowest input10.75 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens706K
Output tokens56.2K
Total tokens762.1K
Input cost$2.12
Output cost$0.8423
Total cost$2.96

Priced from 2026-01-23 · Input $/M $3.00 · Output $/M $15.00

Result 157 of 188

Test T0155 at 2026-01-24

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

genai gemini-2.5-pro · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"s. Sonnenbichler - Montaghemi, Nourhomadeh","year":"","place":null,"pages":null,"publisher":null,"format":null,"editor":null},"library_reference":{"shelfmark":null,"subjects":null}}

Scoring

F1 micro86.8%
F1 macro86.4%
Micro precision84.8%
Micro recall88.9%
Instances263
True positives2,147
False positives385
False negatives268

Speed

Model time, all inputs3309.6 s
Mean per input12.58 s
Slowest input31.71 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens187.3K
Output tokens48.9K
Total tokens236.2K
Input cost$0.2341
Output cost$0.489
Total cost$4.13
Reasoning cost usd$3.40
Total reasoning tokens340.4K

Priced from 2026-01-23 · Input $/M $1.25 · Output $/M $10.00

Result 158 of 188

Test T0200 at 2026-01-24

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

genai gemini-2.5-flash · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":null,"pages":null,"publisher":null,"format":null,"editor":null},"library_reference":{"shelfmark":null,"subjects":null}}

Scoring

F1 micro87.0%
F1 macro86.4%
Micro precision85.0%
Micro recall89.0%
Instances263
True positives2,149
False positives379
False negatives266

Speed

Model time, all inputs1679.4 s
Mean per input6.39 s
Slowest input13.89 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens187.3K
Output tokens48.1K
Total tokens235.4K
Input cost$0.0562
Output cost$0.1203
Total cost$0.8743
Reasoning cost usd$0.6978
Total reasoning tokens279.1K

Priced from 2026-01-23 · Input $/M $0.30 · Output $/M $2.50

Result 159 of 188

Test T0160 at 2026-01-24

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-4.1-2025-04-14 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro84.6%
F1 macro84.2%
Micro precision85.1%
Micro recall84.2%
Instances263
True positives2,033
False positives356
False negatives382

Speed

Model time, all inputs18179.3 s
Mean per input69.12 s
Slowest input1810.44 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens457.8K
Output tokens361K
Total tokens818.9K
Input cost$0.9157
Output cost$2.89
Total cost$3.80

Priced from 2026-01-23 · Input $/M $2.00 · Output $/M $8.00

Result 160 of 188

Test T0408 at 2025-11-25

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-5.1-2025-11-13 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":null},"library_reference":{"shelfmark":"","subjects":"Sonnenbichler - Montaghemi, Nourhomadeh"}}

Scoring

F1 micro82.2%
F1 macro81.8%
Micro precision82.4%
Micro recall81.9%
Instances263
True positives1,979
False positives422
False negatives436

Speed

Model time, all inputs17931.9 s
Mean per input68.18 s
Slowest input1805.73 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens365.2K
Output tokens33.2K
Total tokens398.4K
Input cost$0.4565
Output cost$0.3322
Total cost$0.7886

Priced from 2025-11-24 · Input $/M $1.25 · Output $/M $10.00