RISE Humanities Data Benchmark, 0.6.1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'library_cards__true' with Search Hidden 'False' returned 187 results, showing page 8 of 19.
Result 71 of 187

Test T1725 at 2026-09-08

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter qwen/qwen3.8-flash · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision20.0%
Recall33.3%
True positives1
False positives4
False negatives2
F1 score25.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"S","year":"","place":"","pages":"","publisher":"","format":"","editor":null},"library_reference":{"shelfmark":"","subjects":"Sonnenbichler - Montaghemi, Nourhomadeh"}}

Scoring

F1 micro85.3%
F1 macro83.3%
Micro precision84.2%
Micro recall86.4%
Instances258
True positives2,046
False positives384
False negatives321

Speed

Model time, all inputs32273.5 s
Mean per input122.71 s
Slowest input955.87 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens484.4K
Output tokens848.8K
Total tokens1.3M
Input cost$0.0262
Output cost$0.2523
Total cost$0.4649

Priced from 2026-09-08 · Input $/M $0.15 · Output $/M $0.47

Result 72 of 187

Test T1665 at 2026-09-04

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-6-astra · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro91.7%
F1 macro90.9%
Micro precision90.8%
Micro recall92.6%
Instances263
True positives2,237
False positives228
False negatives178

Speed

Model time, all inputs1565.3 s
Mean per input5.95 s
Slowest input15.14 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens444.1K
Output tokens73K
Total tokens517.2K
Input cost$4.44
Output cost$3.65
Total cost$8.09

Priced from 2026-09-03 · Input $/M $10.00 · Output $/M $50.00

Result 73 of 187

Test T1650 at 2026-09-03

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

genai gemini-3.8-flash · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"","place":"","pages":"","publisher":"","format":null,"editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro90.5%
F1 macro89.6%
Micro precision89.0%
Micro recall91.9%
Instances263
True positives2,220
False positives273
False negatives195

Speed

Model time, all inputs1726.5 s
Mean per input6.56 s
Slowest input113.52 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens406.2K
Output tokens48.8K
Total tokens455K
Input cost$0.3047
Output cost$0.183
Total cost$1.65
Reasoning cost usd$1.16
Total reasoning tokens310K

Priced from 2026-09-02 · Input $/M $0.75 · Output $/M $3.75

Result 74 of 187

Test T1605 at 2026-08-18

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

x-ai grok-4.6 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"","place":null,"pages":null,"publisher":null,"format":null,"editor":null},"library_reference":{"shelfmark":null,"subjects":null}}

Scoring

F1 micro90.4%
F1 macro89.4%
Micro precision89.5%
Micro recall91.3%
Instances263
True positives2,205
False positives259
False negatives210

Speed

Model time, all inputs7609.8 s
Mean per input28.93 s
Slowest input85.28 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens523.1K
Output tokens33.1K
Total tokens556.3K
Input cost$1.05
Output cost$0.1988
Total cost$3.53
Reasoning cost usd$2.29
Total reasoning tokens380.9K

Priced from 2026-08-18 · Input $/M $2.00 · Output $/M $6.00

Result 75 of 187

Test T1635 at 2026-08-18

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter z-ai/glm-5v-turbo · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{
  "type": {
    "type": "Reference"
  },
  "author": {
    "last_name": "Montaghemi",
    "first_name": "Nourhomadeh"
  },
  "publication": {
    "title": "Sonnenbichler - Montaghemi, Nourhomadeh",
    "place": "",
    "pages": "",
    "publisher": "",
    "format": ""
  },
  "library_reference": {
    "shelfmark": "",
    "subjects": ""
  }
}

Scoring

F1 micro84.8%
F1 macro84.2%
Micro precision82.9%
Micro recall86.8%
Instances261
True positives2,079
False positives428
False negatives317

Speed

Model time, all inputs9303.7 s
Mean per input35.38 s
Slowest input66.46 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens519.6K
Output tokens324.4K
Total tokens843.9K
Input cost$0.6235
Output cost$1.30
Total cost$1.92

Priced from 2026-08-18 · Input $/M $1.20 · Output $/M $4.00

Result 76 of 187

Test T1620 at 2026-08-18

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter meta/muse-spark-1.2 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro90.3%
F1 macro89.2%
Micro precision90.5%
Micro recall90.2%
Instances262
True positives2,170
False positives229
False negatives237

Speed

Model time, all inputs128221.0 s
Mean per input487.53 s
Slowest input1370.89 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens489.3K
Output tokens979K
Total tokens1.5M
Input cost$0.6117
Output cost$4.16
Total cost$4.77

Priced from 2026-08-18 · Input $/M $1.25 · Output $/M $4.25

Result 77 of 187

Test T1590 at 2026-08-18

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

genai gemini-3.7-flash · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"","place":"","pages":"","publisher":"","format":"","editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro88.4%
F1 macro87.5%
Micro precision86.7%
Micro recall90.2%
Instances263
True positives2,178
False positives333
False negatives237

Speed

Model time, all inputs910.4 s
Mean per input3.46 s
Slowest input10.15 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens406.2K
Output tokens49.6K
Total tokens455.8K
Input cost$0.3047
Output cost$0.1858
Total cost$1.03
Reasoning cost usd$0.542
Total reasoning tokens144.5K

Priced from 2026-08-18 · Input $/M $0.75 · Output $/M $3.75

Result 78 of 187

Test T1545 at 2026-08-15

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

huggingface thinkingmachines/Inkling-Small:deepinfra · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision0.0%
Recall0.0%
True positives0
False positives0
False negatives3
F1 score0.0%
Total fields13

Scoring

F1 micro71.9%
F1 macro62.3%
Micro precision80.6%
Micro recall64.9%
Instances263
True positives1,567
False positives378
False negatives848

Speed

Model time, all inputs68538.9 s
Mean per input260.60 s
Slowest input662.22 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens508.9K
Output tokens410.5K
Total tokens919.4K
Input cost$0.229
Output cost$0.4926
Total cost$0.7216

Priced from 2026-08-15 · Input $/M $0.45 · Output $/M $1.20

Result 79 of 187

Test T1575 at 2026-08-15

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

huggingface meta-models/Muse-Glimmer-30B · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro89.6%
F1 macro88.9%
Micro precision88.6%
Micro recall90.6%
Instances263
True positives2,188
False positives282
False negatives227

Speed

Model time, all inputs101690.5 s
Mean per input386.66 s
Slowest input1300.78 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens503.1K
Output tokens664.1K
Total tokens1.2M
Input cost$0.1761
Output cost$0.9962
Total cost$1.17

Priced from 2026-08-15 · Input $/M $0.35 · Output $/M $1.50

Result 80 of 187

Test T1530 at 2026-08-15

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

huggingface MiniMaxAI/MiniMax-M3 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro76.5%
F1 macro76.1%
Micro precision72.6%
Micro recall80.7%
Instances263
True positives1,949
False positives734
False negatives466

Speed

Model time, all inputs6250.8 s
Mean per input23.77 s
Slowest input146.06 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens323.6K
Output tokens52K
Total tokens375.6K
Input cost$0.0906
Output cost$0.0572
Total cost$0.1478

Priced from 2026-08-15 · Input $/M $0.28 · Output $/M $1.10