RISE Humanities Data Benchmark, 0.6.0-pre1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'library_cards__true' with Search Hidden 'False' returned 180 results, showing page 9 of 18.
Result 81 of 180

Test T1350 at 2026-07-25

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-5.6-sol · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler-Montaghemi, Nourhomadeh","year":"","place":"","pages":"","publisher":"","format":"","editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro87.3%
F1 macro86.8%
Micro precision86.6%
Micro recall88.1%
Instances263
True positives2,127
False positives330
False negatives288

Speed

Model time, all inputs1742.6 s
Mean per input6.63 s
Slowest input19.54 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens444.4K
Output tokens90.5K
Total tokens534.9K
Input cost$2.22
Output cost$2.71
Total cost$4.94

Priced from 2026-07-25 · Input $/M $5.00 · Output $/M $30.00

Result 82 of 180

Test T1425 at 2026-07-25

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

genai gemini-3.6-flash · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro88.5%
F1 macro87.4%
Micro precision86.7%
Micro recall90.3%
Instances263
True positives2,180
False positives333
False negatives235

Speed

Model time, all inputs3031.3 s
Mean per input11.53 s
Slowest input39.24 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens406.2K
Output tokens39.8K
Total tokens446K
Input cost$0.6094
Output cost$0.2985
Total cost$3.43
Reasoning cost usd$2.53
Total reasoning tokens336.8K

Priced from 2026-07-25 · Input $/M $1.50 · Output $/M $7.50

Result 83 of 180

Test T1380 at 2026-07-25

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-5.6-luna · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro84.9%
F1 macro84.4%
Micro precision83.0%
Micro recall87.0%
Instances263
True positives2,101
False positives431
False negatives314

Speed

Model time, all inputs1209.2 s
Mean per input4.60 s
Slowest input21.68 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens444.4K
Output tokens124.5K
Total tokens568.9K
Input cost$0.4444
Output cost$0.7468
Total cost$1.19

Priced from 2026-07-25 · Input $/M $1.00 · Output $/M $6.00

Result 84 of 180

Test T1395 at 2026-07-25

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter moonshotai/kimi-k3 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro89.1%
F1 macro88.5%
Micro precision88.3%
Micro recall89.8%
Instances263
True positives2,168
False positives286
False negatives247

Speed

Model time, all inputs27341.2 s
Mean per input103.96 s
Slowest input315.24 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens516.1K
Output tokens850.7K
Total tokens1.4M
Input cost$1.54
Output cost$12.68
Total cost$14.30

Priced from 2026-07-25 · Input $/M $3.00 · Output $/M $15.00

Result 85 of 180

Test T1365 at 2026-07-25

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-5.6-terra · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler-Montaghemi, Nourhomadeh","year":"","place":"","pages":"","publisher":"","format":"","editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro85.2%
F1 macro84.6%
Micro precision84.3%
Micro recall86.0%
Instances263
True positives2,078
False positives387
False negatives337

Speed

Model time, all inputs1095.4 s
Mean per input4.16 s
Slowest input164.07 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens444.4K
Output tokens75K
Total tokens519.4K
Input cost$1.11
Output cost$1.13
Total cost$2.24

Priced from 2026-07-25 · Input $/M $2.50 · Output $/M $15.00

Result 86 of 180

Test T1195 at 2026-07-02

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

anthropic claude-fable-5 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"s. Sonnenbichler-Montaghemi, Nourhomadeh","year":"","place":null,"pages":null,"publisher":null,"format":null,"editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro87.1%
F1 macro84.6%
Micro precision86.0%
Micro recall88.2%
Instances263
True positives2,129
False positives346
False negatives286

Speed

Model time, all inputs1770.3 s
Mean per input6.73 s
Slowest input10.86 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens770.4K
Output tokens68K
Total tokens838.4K
Input cost$7.70
Output cost$3.40
Total cost$11.10

Priced from 2026-07-02 · Input $/M $10.00 · Output $/M $50.00

Result 87 of 180

Test T1182 at 2026-07-01

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

anthropic claude-sonnet-5 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision0.0%
Recall0.0%
True positives0
False positives5
False negatives3
F1 score0.0%
Total fields23

{"document": {"type": {"type": "Reference"}, "author": {"last_name": "Montaghemi", "first_name": "Nourhomadeh"}, "publication": {"title": "", "year": 0, "place": "", "pages": "", "publisher": "", "format": ""}, "library_reference": {"shelfmark": "", "subjects": "Sonnenbichler-Montaghemi, Nourhomadeh"}}}

Scoring

F1 micro4.5%
F1 macro4.7%
Micro precision4.4%
Micro recall4.6%
Instances263
True positives112
False positives2,456
False negatives2,302

Speed

Model time, all inputs1082.7 s
Mean per input4.12 s
Slowest input11.81 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens787.2K
Output tokens69.5K
Total tokens856.7K
Input cost$1.57
Output cost$0.6948
Total cost$2.27

Priced from 2026-07-01 · Input $/M $2.00 · Output $/M $10.00

Result 88 of 180

Test T1169 at 2026-06-29

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

genai gemini-3.1-flash-lite · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"","place":"","pages":"","publisher":"","format":"","editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro81.7%
F1 macro75.8%
Micro precision84.8%
Micro recall78.8%
Instances263
True positives1,902
False positives340
False negatives513

Speed

Model time, all inputs797.1 s
Mean per input3.03 s
Slowest input14.38 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens352.2K
Output tokens41K
Total tokens393.2K
Input cost$0.0881
Output cost$0.0615
Total cost$0.1496

Priced from 2026-06-29 · Input $/M $0.25 · Output $/M $1.50

Result 89 of 180

Test T1156 at 2026-06-22

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

scicore qwen35-397b-a17b-fp8 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro81.6%
F1 macro81.0%
Micro precision80.4%
Micro recall83.0%
Instances263
True positives2,004
False positives490
False negatives411

Speed

Model time, all inputs25138.3 s
Mean per input95.58 s
Slowest input1967.02 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens274.5K
Output tokens969.7K
Total tokens1.2M
Input cost$0.00
Output cost$0.00
Total cost$0.00

Priced from 2026-06-05 · Input $/M $0.00 · Output $/M $0.00

Result 90 of 180

Test T1107 at 2026-06-08

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

x-ai grok-4.3 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro85.7%
F1 macro84.6%
Micro precision85.5%
Micro recall85.9%
Instances263
True positives2,074
False positives353
False negatives341

Speed

Model time, all inputs166125.2 s
Mean per input631.65 s
Slowest input1819.58 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens468.1K
Output tokens41.3K
Total tokens509.4K
Input cost$0.5851
Output cost$0.1033
Total cost$1.18
Reasoning cost usd$0.4873
Total reasoning tokens194.9K

Priced from 2026-06-05 · Input $/M $1.25 · Output $/M $2.50