RISE Humanities Data Benchmark, 0.6.1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'library_cards__true' with Search Hidden 'False' returned 188 results, showing page 1 of 19.
Result 1 of 188

Test T1815 at 2026-09-29

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

anthropic claude-opus-5-5 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"s. Sonnenbichler-Montaghemi, Nourhomadeh","year":null,"place":"","pages":"","publisher":"","format":"","editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro91.6%
F1 macro90.8%
Micro precision90.6%
Micro recall92.6%
Instances263
True positives2,236
False positives233
False negatives179

Speed

Model time, all inputs1332.9 s
Mean per input5.07 s
Slowest input10.31 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens749.6K
Output tokens109.8K
Total tokens859.4K
Input cost$3.00
Output cost$2.20
Total cost$5.19

Priced from 2026-09-27 · Input $/M $4.00 · Output $/M $20.00

Result 2 of 188

Test T1875 at 2026-09-29

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

anthropic claude-sonnet-5-5 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler-Montaghemi, Nourhomadeh","year":null,"place":"","pages":"","publisher":"","format":"","editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro91.8%
F1 macro91.0%
Micro precision90.7%
Micro recall93.0%
Instances263
True positives2,246
False positives230
False negatives169

Speed

Model time, all inputs677.8 s
Mean per input2.58 s
Slowest input5.79 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens749.6K
Output tokens73.1K
Total tokens822.7K
Input cost$1.50
Output cost$0.7312
Total cost$2.23

Priced from 2026-09-28 · Input $/M $2.00 · Output $/M $10.00

Result 3 of 188

Test T1800 at 2026-09-29

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

scicore Qwen3.8-Flash-Next-FP8 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":"Sonnenbichler"}}

Scoring

F1 micro79.9%
F1 macro79.3%
Micro precision76.9%
Micro recall83.3%
Instances263
True positives2,011
False positives605
False negatives404

Speed

Model time, all inputs3803.2 s
Mean per input14.46 s
Slowest input166.34 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens288.2K
Output tokens400.8K
Total tokens689K
Input cost$0.00
Output cost$0.00
Total cost$0.00

Priced from 2026-09-27 · Input $/M $0.00 · Output $/M $0.00

Result 4 of 188

Test T1680 at 2026-09-29

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

anthropic claude-fable-5-1 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":null,"place":"","pages":"","publisher":"","format":"","editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro90.0%
F1 macro89.1%
Micro precision86.0%
Micro recall94.2%
Instances263
True positives2,276
False positives369
False negatives139

Speed

Model time, all inputs2042.2 s
Mean per input7.77 s
Slowest input17.46 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens749.6K
Output tokens100.9K
Total tokens850.5K
Input cost$7.50
Output cost$5.05
Total cost$12.54

Priced from 2026-09-08 · Input $/M $10.00 · Output $/M $50.00

Result 5 of 188

Test T1875 at 2026-09-28

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

anthropic claude-sonnet-5-5 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

```json
{
  "type": {
    "type": "Reference"
  },
  "author": {
    "last_name": "Montaghemi",
    "first_name": "Nourhomadeh"
  },
  "publication": {
    "title": "",
    "year": "",
    "place": "",
    "pages": "",
    "publisher": "",
    "format": ""
  },
  "library_reference": {
    "shelfmark": "",
    "subjects": ""
  }
}
```

Scoring

F1 micro92.2%
F1 macro91.4%
Micro precision91.4%
Micro recall93.1%
Instances263
True positives2,248
False positives211
False negatives167

Speed

Model time, all inputs762.9 s
Mean per input2.90 s
Slowest input8.97 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens364.9K
Output tokens59.5K
Total tokens424.3K
Input cost$0.7297
Output cost$0.5946
Total cost$1.32

Priced from 2026-09-28 · Input $/M $2.00 · Output $/M $10.00

Result 6 of 188

Test T1860 at 2026-09-27

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

x-ai grok-4.7 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":null,"place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro91.7%
F1 macro90.8%
Micro precision91.0%
Micro recall92.3%
Instances263
True positives2,230
False positives221
False negatives185

Speed

Model time, all inputs15382.7 s
Mean per input58.49 s
Slowest input239.49 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens807.7K
Output tokens39.9K
Total tokens847.6K
Input cost$1.62
Output cost$0.2396
Total cost$9.07
Reasoning cost usd$7.22
Total reasoning tokens1.2M

Priced from 2026-09-27 · Input $/M $2.00 · Output $/M $6.00

Result 7 of 188

Test T1845 at 2026-09-27

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-6-luna · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro86.7%
F1 macro86.3%
Micro precision86.2%
Micro recall87.2%
Instances263
True positives2,105
False positives338
False negatives310

Speed

Model time, all inputs1511.0 s
Mean per input5.75 s
Slowest input76.82 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens447.3K
Output tokens144K
Total tokens591.3K
Input cost$0.0447
Output cost$0.072
Total cost$0.1167

Priced from 2026-09-27 · Input $/M $0.10 · Output $/M $0.50

Result 8 of 188

Test T1830 at 2026-09-27

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-6-sol · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro90.8%
F1 macro90.1%
Micro precision90.7%
Micro recall90.9%
Instances263
True positives2,196
False positives224
False negatives219

Speed

Model time, all inputs1673.9 s
Mean per input6.36 s
Slowest input175.70 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens447.5K
Output tokens100K
Total tokens547.5K
Input cost$0.895
Output cost$1.00
Total cost$1.89

Priced from 2026-09-27 · Input $/M $2.00 · Output $/M $10.00

Result 9 of 188

Test T1815 at 2026-09-27

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

anthropic claude-opus-5-5 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

```json
{
  "type": {
    "type": "Reference"
  },
  "author": {
    "last_name": "Montaghemi",
    "first_name": "Nourhomadeh"
  },
  "publication": {
    "title": "",
    "place": "",
    "pages": "",
    "publisher": "",
    "format": ""
  },
  "library_reference": {
    "shelfmark": "",
    "subjects": ""
  }
}
```

Scoring

F1 micro91.8%
F1 macro91.1%
Micro precision91.2%
Micro recall92.4%
Instances263
True positives2,231
False positives216
False negatives184

Speed

Model time, all inputs1573.0 s
Mean per input5.98 s
Slowest input14.53 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens364.9K
Output tokens99.8K
Total tokens464.7K
Input cost$1.46
Output cost$2.00
Total cost$3.46

Priced from 2026-09-27 · Input $/M $4.00 · Output $/M $20.00

Result 10 of 188

Test T1740 at 2026-09-18

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter qwen/qwen3.8-27b · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghe mi","first_name":"Nourhomadeh"},"publication":{"title":"","year":null,"place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro86.8%
F1 macro85.5%
Micro precision86.1%
Micro recall87.4%
Instances252
True positives2,023
False positives326
False negatives291

Speed

Model time, all inputs114293.1 s
Mean per input439.59 s
Slowest input7371.09 s
Inputs timed260

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens442.4K
Output tokens937.7K
Total tokens1.4M
Input cost$0.0889
Output cost$2.17
Total cost$2.26

Priced from 2026-09-18 · Input $/M $0.214 · Output $/M $2.55