RISE Humanities Data Benchmark, 0.6.0-pre1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'library_cards__true' with Search Hidden 'False' returned 180 results, showing page 4 of 18.
Result 31 of 180

Test T0949 at 2026-09-17

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter qwen/qwen3.5-397b-a17b · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"type","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro80.5%
F1 macro80.0%
Micro precision78.4%
Micro recall82.6%
Instances263
True positives1,994
False positives548
False negatives421

Speed

Model time, all inputs27316.0 s
Mean per input103.86 s
Slowest input3608.08 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens282.8K
Output tokens259.5K
Total tokens542.4K
Input cost$0.1438
Output cost$0.7597
Total cost$1.05

Priced from 2026-09-16 · Input $/M $0.55 · Output $/M $3.50

Result 32 of 180

Test T1195 at 2026-09-17

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

anthropic claude-fable-5 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"s. Sonnenbichler-Montaghemi, Nourhomadeh","year":"","place":null,"pages":null,"publisher":null,"format":null,"editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro88.7%
F1 macro87.7%
Micro precision85.9%
Micro recall91.7%
Instances263
True positives2,214
False positives364
False negatives201

Speed

Model time, all inputs1612.0 s
Mean per input6.13 s
Slowest input22.19 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens770.4K
Output tokens71K
Total tokens841.4K
Input cost$7.70
Output cost$3.55
Total cost$11.25

Priced from 2026-09-16 · Input $/M $10.00 · Output $/M $50.00

Result 33 of 180

Test T0258 at 2026-09-17

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter qwen/qwen3-vl-30b-a3b-instruct · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghem","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":"Sonnenbichler - Montaghem"},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro74.2%
F1 macro73.4%
Micro precision73.2%
Micro recall75.2%
Instances260
True positives1,794
False positives657
False negatives592

Speed

Model time, all inputs28299.0 s
Mean per input107.60 s
Slowest input2310.42 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens308.7K
Output tokens49.7K
Total tokens358.4K
Input cost$0.0313
Output cost$0.0214
Total cost$0.0782

Priced from 2026-09-16 · Input $/M $0.15 · Output $/M $0.60

Result 34 of 180

Test T1001 at 2026-09-17

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter google/gemma-4-26b-a4b-it · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro81.6%
F1 macro81.0%
Micro precision76.7%
Micro recall87.1%
Instances263
True positives2,104
False positives639
False negatives311

Speed

Model time, all inputs2684.4 s
Mean per input10.21 s
Slowest input521.63 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens201.3K
Output tokens53.9K
Total tokens255.2K
Input cost$0.0181
Output cost$0.0162
Total cost$0.0343

Priced from 2026-09-16 · Input $/M $0.09 · Output $/M $0.30

Result 35 of 180

Test T1182 at 2026-09-16

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

anthropic claude-sonnet-5 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision0.0%
Recall0.0%
True positives0
False positives5
False negatives3
F1 score0.0%
Total fields23

{"document": {"type": {"type": "Reference"}, "author": {"last_name": "Montaghemi", "first_name": "Nourhomadeh"}, "publication": {"title": "", "year": 0, "place": "", "pages": "", "publisher": "", "format": ""}, "library_reference": {"shelfmark": "", "subjects": "Sonnenbichler-Montaghemi, Nourhomadeh"}}}

Scoring

F1 micro9.2%
F1 macro9.2%
Micro precision9.0%
Micro recall9.5%
Instances263
True positives230
False positives2,339
False negatives2,185

Speed

Model time, all inputs864.3 s
Mean per input3.29 s
Slowest input6.36 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens787.2K
Output tokens68.3K
Total tokens855.6K
Input cost$1.57
Output cost$0.6832
Total cost$2.26

Priced from 2026-09-16 · Input $/M $2.00 · Output $/M $10.00

Result 36 of 180

Test T0493 at 2026-09-16

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-5.2-2025-12-11 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"","place":"","pages":"","publisher":"","format":"","editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring
Speed

Model time, all inputs5994.5 s
Mean per input352.62 s
Slowest input603.66 s
Inputs timed17

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

No cost recorded for this run

Priced from 2026-09-16 · Input $/M $1.75 · Output $/M $14.00

Result 37 of 180

Test T0645 at 2026-09-16

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

anthropic claude-sonnet-4-6 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision20.0%
Recall33.3%
True positives1
False positives4
False negatives2
F1 score25.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"<UNKNOWN>","place":null,"pages":null,"publisher":null,"format":null,"editor":null},"library_reference":{"shelfmark":null,"subjects":null}}

Scoring

F1 micro87.3%
F1 macro86.3%
Micro precision82.6%
Micro recall92.6%
Instances263
True positives2,236
False positives470
False negatives179

Speed

Model time, all inputs1361.2 s
Mean per input5.18 s
Slowest input15.20 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens706.2K
Output tokens57.2K
Total tokens763.5K
Input cost$2.12
Output cost$0.8584
Total cost$2.98

Priced from 2026-09-16 · Input $/M $3.00 · Output $/M $15.00

Result 38 of 180

Test T0526 at 2026-09-16

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

anthropic claude-haiku-4-5-20251001 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler-Montaghemi, Nourhomadeh","year":"","place":null,"pages":null,"publisher":null,"format":null,"editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro84.4%
F1 macro83.5%
Micro precision83.3%
Micro recall85.6%
Instances263
True positives2,067
False positives414
False negatives348

Speed

Model time, all inputs647.2 s
Mean per input2.46 s
Slowest input5.34 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens706K
Output tokens53K
Total tokens759K
Input cost$0.706
Output cost$0.265
Total cost$0.971

Priced from 2026-09-16 · Input $/M $1.00 · Output $/M $5.00

Result 39 of 180

Test T1605 at 2026-09-16

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

x-ai grok-4.6 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":null,"pages":null,"publisher":null,"format":null,"editor":null},"library_reference":{"shelfmark":null,"subjects":null}}

Scoring

F1 micro91.6%
F1 macro90.8%
Micro precision90.5%
Micro recall92.7%
Instances263
True positives2,239
False positives234
False negatives176

Speed

Model time, all inputs10404.0 s
Mean per input39.56 s
Slowest input108.22 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens636.2K
Output tokens32.7K
Total tokens668.9K
Input cost$1.27
Output cost$0.1962
Total cost$5.15
Reasoning cost usd$3.68
Total reasoning tokens614.1K

Priced from 2026-08-18 · Input $/M $2.00 · Output $/M $6.00

Result 40 of 180

Test T1027 at 2026-09-16

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

anthropic claude-opus-4-7 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"","place":null,"pages":null,"publisher":null,"format":null,"editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro89.7%
F1 macro88.8%
Micro precision87.4%
Micro recall92.2%
Instances263
True positives2,227
False positives322
False negatives188

Speed

Model time, all inputs1133.8 s
Mean per input4.31 s
Slowest input11.44 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens875.3K
Output tokens71.7K
Total tokens947K
Input cost$4.38
Output cost$1.79
Total cost$6.17

Priced from 2026-09-16 · Input $/M $5.00 · Output $/M $25.00