RISE Humanities Data Benchmark, 0.6.1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'library_cards__true' with Search Hidden 'False' returned 188 results, showing page 18 of 19.
Result 171 of 188

Test T0167 at 2025-10-01

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-5-nano · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

F1 micro77.1%
F1 macro76.6%
Micro precision76.0%
Micro recall78.1%
Instances263
True positives1,887
False positives596
False negatives528

Speed

Model time, all inputs5049.1 s
Mean per input19.20 s
Slowest input83.88 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens434.6K
Output tokens807.7K
Total tokens1.2M
Input cost$0.0217
Output cost$0.3231
Total cost$0.3448

Priced from 2025-10-01 · Input $/M $0.05 · Output $/M $0.40

Result 172 of 188

Test T0166 at 2025-10-01

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-5-mini · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

F1 micro79.1%
F1 macro78.9%
Micro precision77.2%
Micro recall81.1%
Instances263
True positives1,959
False positives579
False negatives456

Speed

Model time, all inputs22744.8 s
Mean per input86.48 s
Slowest input783.16 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens387.5K
Output tokens609.5K
Total tokens997K
Input cost$0.0969
Output cost$1.22
Total cost$1.32

Priced from 2025-10-01 · Input $/M $0.25 · Output $/M $2.00

Result 173 of 188

Test T0155 at 2025-10-01

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

genai gemini-2.5-pro · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

{
  "type": {
    "type": "Reference"
  },
  "author": {
    "last_name": "Montaghemi",
    "first_name": "Nourhomadeh"
  },
  "publication": {
    "title": "Sonnenbichler - Montaghemi, Nourhomadeh",
    "year": 
    "place",
    "pages": "",
    "publisher": "",
    "format": ""
  },
  "library_reference": {
    "shelfmark": "",
    "subjects": ""
  }
}

Scoring

F1 micro87.0%
F1 macro86.5%
Micro precision85.4%
Micro recall88.7%
Instances263
True positives2,141
False positives367
False negatives274

Speed

Model time, all inputs3744.4 s
Mean per input14.24 s
Slowest input31.98 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens185.4K
Output tokens47.8K
Total tokens233.2K
Input cost$0.2318
Output cost$0.4777
Total cost$3.66
Reasoning cost usd$2.96
Total reasoning tokens295.5K

Priced from 2025-10-01 · Input $/M $1.25 · Output $/M $10.00

Result 174 of 188

Test T0200 at 2025-09-30

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

genai gemini-2.5-flash · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro85.7%
F1 macro85.0%
Micro precision83.5%
Micro recall88.1%
Instances263
True positives2,128
False positives422
False negatives287

Speed

Model time, all inputs3027.5 s
Mean per input11.51 s
Slowest input53.14 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens185.4K
Output tokens45.1K
Total tokens230.5K
Input cost$0.0556
Output cost$0.1128
Total cost$1.11
Reasoning cost usd$0.9378
Total reasoning tokens375.1K

Priced from 2025-09-30 · Input $/M $0.30 · Output $/M $2.50

Result 175 of 188

Test T0160 at 2025-09-30

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-4.1 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

F1 micro85.9%
F1 macro85.4%
Micro precision86.7%
Micro recall85.2%
Instances263
True positives2,057
False positives316
False negatives358

Speed

Model time, all inputs11597.0 s
Mean per input44.10 s
Slowest input607.63 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens401.3K
Output tokens245.2K
Total tokens646.5K
Input cost$0.8027
Output cost$1.96
Total cost$2.76

Priced from 2025-09-30 · Input $/M $2.00 · Output $/M $8.00

Result 176 of 188

Test T0066 at 2025-09-30

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-4o · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

F1 micro82.5%
F1 macro82.1%
Micro precision82.5%
Micro recall82.5%
Instances263
True positives1,992
False positives423
False negatives423

Speed

Model time, all inputs3291.1 s
Mean per input12.51 s
Slowest input319.51 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens404.8K
Output tokens29.1K
Total tokens433.9K
Input cost$1.01
Output cost$0.2912
Total cost$1.30

Priced from 2025-09-30 · Input $/M $2.50 · Output $/M $10.00

Result 177 of 188

Test T0230 at 2025-09-30

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

anthropic claude-sonnet-4-5-20250929 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

F1 micro80.6%
F1 macro79.3%
Micro precision79.8%
Micro recall81.5%
Instances263
True positives1,968
False positives499
False negatives447

Speed

Model time, all inputs1428.8 s
Mean per input5.43 s
Slowest input54.98 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens638.4K
Output tokens57.2K
Total tokens695.6K
Input cost$1.92
Output cost$0.8579
Total cost$2.77

Priced from 2025-09-30 · Input $/M $3.00 · Output $/M $15.00

Result 178 of 188

Test T0162 at 2025-09-24

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-4.1-nano · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

F1 micro61.2%
F1 macro60.4%
Micro precision65.3%
Micro recall57.6%
Instances263
True positives1,390
False positives738
False negatives1,025

Speed

Model time, all inputs1867.0 s
Mean per input7.10 s
Slowest input304.95 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

No cost recorded for this run

Result 179 of 188

Test T0167 at 2025-09-02

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-5-nano · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

F1 micro80.2%
F1 macro79.7%
Micro precision78.9%
Micro recall81.5%
Instances263
True positives1,969
False positives526
False negatives446

Speed

Model time, all inputs4991.5 s
Mean per input18.98 s
Slowest input36.17 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

No cost recorded for this run

Result 180 of 188

Test T0162 at 2025-09-02

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-4.1-nano · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

F1 micro73.0%
F1 macro71.8%
Micro precision77.2%
Micro recall69.2%
Instances263
True positives1,671
False positives494
False negatives744

Speed

Model time, all inputs573.1 s
Mean per input2.18 s
Slowest input6.54 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

No cost recorded for this run