RISE Humanities Data Benchmark, 0.6.1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'library_cards__true' with Search Hidden 'False' returned 187 results, showing page 7 of 19.
Result 61 of 187

Test T1605 at 2026-09-16

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

x-ai grok-4.6 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":null,"pages":null,"publisher":null,"format":null,"editor":null},"library_reference":{"shelfmark":null,"subjects":null}}

Scoring

F1 micro91.6%
F1 macro90.8%
Micro precision90.5%
Micro recall92.7%
Instances263
True positives2,239
False positives234
False negatives176

Speed

Model time, all inputs10404.0 s
Mean per input39.56 s
Slowest input108.22 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens636.2K
Output tokens32.7K
Total tokens668.9K
Input cost$1.27
Output cost$0.1962
Total cost$5.15
Reasoning cost usd$3.68
Total reasoning tokens614.1K

Priced from 2026-08-18 · Input $/M $2.00 · Output $/M $6.00

Result 62 of 187

Test T0493 at 2026-09-16

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-5.2-2025-12-11 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"","place":"","pages":"","publisher":"","format":"","editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring
Speed

Model time, all inputs5994.5 s
Mean per input352.62 s
Slowest input603.66 s
Inputs timed17

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

No cost recorded for this run

Priced from 2026-09-16 · Input $/M $1.75 · Output $/M $14.00

Result 63 of 187

Test T0645 at 2026-09-16

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

anthropic claude-sonnet-4-6 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision20.0%
Recall33.3%
True positives1
False positives4
False negatives2
F1 score25.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"<UNKNOWN>","place":null,"pages":null,"publisher":null,"format":null,"editor":null},"library_reference":{"shelfmark":null,"subjects":null}}

Scoring

F1 micro87.3%
F1 macro86.3%
Micro precision82.6%
Micro recall92.6%
Instances263
True positives2,236
False positives470
False negatives179

Speed

Model time, all inputs1361.2 s
Mean per input5.18 s
Slowest input15.20 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens706.2K
Output tokens57.2K
Total tokens763.5K
Input cost$2.12
Output cost$0.8584
Total cost$2.98

Priced from 2026-09-16 · Input $/M $3.00 · Output $/M $15.00

Result 64 of 187

Test T1650 at 2026-09-16

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

genai gemini-3.8-flash · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro90.6%
F1 macro89.8%
Micro precision89.1%
Micro recall92.1%
Instances263
True positives2,225
False positives271
False negatives190

Speed

Model time, all inputs1470.8 s
Mean per input5.59 s
Slowest input54.37 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens406.2K
Output tokens48.4K
Total tokens454.7K
Input cost$0.3047
Output cost$0.1816
Total cost$1.59
Reasoning cost usd$1.10
Total reasoning tokens293.7K

Priced from 2026-09-02 · Input $/M $0.75 · Output $/M $3.75

Result 65 of 187

Test T1440 at 2026-09-16

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

genai gemini-3.5-flash-lite · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"s. Sonnenbichler - Montaghemi, Nourhomadeh","year":"","place":null,"pages":null,"publisher":null,"format":null,"editor":null},"library_reference":{"shelfmark":null,"subjects":null}}

Scoring

F1 micro87.4%
F1 macro86.7%
Micro precision84.5%
Micro recall90.6%
Instances263
True positives2,187
False positives402
False negatives228

Speed

Model time, all inputs339.9 s
Mean per input1.29 s
Slowest input2.80 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens406.2K
Output tokens49.3K
Total tokens455.5K
Input cost$0.1219
Output cost$0.1232
Total cost$0.2451

Priced from 2026-09-16 · Input $/M $0.30 · Output $/M $2.50

Result 66 of 187

Test T0526 at 2026-09-16

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

anthropic claude-haiku-4-5-20251001 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler-Montaghemi, Nourhomadeh","year":"","place":null,"pages":null,"publisher":null,"format":null,"editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro84.4%
F1 macro83.5%
Micro precision83.3%
Micro recall85.6%
Instances263
True positives2,067
False positives414
False negatives348

Speed

Model time, all inputs647.2 s
Mean per input2.46 s
Slowest input5.34 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens706K
Output tokens53K
Total tokens759K
Input cost$0.706
Output cost$0.265
Total cost$0.971

Priced from 2026-09-16 · Input $/M $1.00 · Output $/M $5.00

Result 67 of 187

Test T0633 at 2026-09-16

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

anthropic claude-opus-4-6 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision20.0%
Recall33.3%
True positives1
False positives4
False negatives2
F1 score25.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"<UNKNOWN>","place":null,"pages":null,"publisher":null,"format":null,"editor":null},"library_reference":{"shelfmark":null,"subjects":null}}

Scoring

F1 micro85.6%
F1 macro85.0%
Micro precision81.8%
Micro recall89.8%
Instances263
True positives2,168
False positives483
False negatives247

Speed

Model time, all inputs1737.1 s
Mean per input6.60 s
Slowest input17.50 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens706.2K
Output tokens57.4K
Total tokens763.6K
Input cost$3.53
Output cost$1.43
Total cost$4.97

Priced from 2026-09-16 · Input $/M $5.00 · Output $/M $25.00

Result 68 of 187

Test T1740 at 2026-09-09

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter qwen/qwen3.8-27b · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision100.0%
Recall100.0%
True positives3
False positives0
False negatives0
F1 score100.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Sonnenbichler","first_name":""},"publication":{"title":"Montaghem i, Nourhomadeh","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro86.0%
F1 macro84.2%
Micro precision85.6%
Micro recall86.5%
Instances263
True positives2,088
False positives352
False negatives327

Speed

Model time, all inputs477274.0 s
Mean per input1814.73 s
Slowest input21985.50 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens460.7K
Output tokens1.3M
Total tokens1.8M
Input cost$0.1231
Output cost$3.79
Total cost$3.91

Priced from 2026-09-08 · Input $/M $0.42 · Output $/M $3.00

Result 69 of 187

Test T1710 at 2026-09-08

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter z-ai/glm-5.3-flash · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghem","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro79.2%
F1 macro78.2%
Micro precision78.2%
Micro recall80.2%
Instances263
True positives1,937
False positives539
False negatives478

Speed

Model time, all inputs7443.4 s
Mean per input28.30 s
Slowest input941.20 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens323.8K
Output tokens66.6K
Total tokens390.5K
Input cost$0.0244
Output cost$0.0167
Total cost$0.0411

Priced from 2026-09-08 · Input $/M $0.075 · Output $/M $0.25

Result 70 of 187

Test T1695 at 2026-09-08

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter meta/muse-spark-1.3 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro91.1%
F1 macro90.2%
Micro precision91.3%
Micro recall90.9%
Instances263
True positives2,196
False positives210
False negatives219

Speed

Model time, all inputs5234.2 s
Mean per input19.90 s
Slowest input47.48 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens454K
Output tokens630.8K
Total tokens1.1M
Input cost$0.5675
Output cost$2.68
Total cost$3.25

Priced from 2026-09-08 · Input $/M $1.25 · Output $/M $4.25