RISE Humanities Data Benchmark, 0.5.5

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'library_cards__true' with Search Hidden 'False' returned 126 results, showing page 11 of 13.
Result 101 of 126

Test T0161 at 2025-10-02

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-4.1-mini · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

F1 micro75.9%
F1 macro75.7%
Micro precision75.8%
Micro recall75.9%
Instances263
True positives1,834
False positives584
False negatives581

Speed

Model time, all inputs29647.3 s
Mean per input112.73 s
Slowest input1506.66 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

No cost recorded for this run

Priced from 2025-10-01 · Input $/M $0.40 · Output $/M $1.60

Result 102 of 126

Test T0162 at 2025-10-02

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-4.1-nano · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

F1 micro67.6%
F1 macro66.7%
Micro precision72.2%
Micro recall63.6%
Instances263
True positives1,536
False positives592
False negatives879

Speed

Model time, all inputs471.3 s
Mean per input1.79 s
Slowest input5.28 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

No cost recorded for this run

Priced from 2025-10-01 · Input $/M $0.10 · Output $/M $0.40

Result 103 of 126

Test T0168 at 2025-10-01

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai o3 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

F1 micro83.4%
F1 macro83.1%
Micro precision83.7%
Micro recall83.1%
Instances263
True positives2,006
False positives390
False negatives409

Speed

Model time, all inputs9024.0 s
Mean per input34.31 s
Slowest input620.80 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens376.6K
Output tokens223.3K
Total tokens599.9K
Input cost$0.7532
Output cost$1.79
Total cost$2.54

Priced from 2025-10-01 · Input $/M $2.00 · Output $/M $8.00

Result 104 of 126

Test T0160 at 2025-10-01

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-4.1 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

F1 micro85.5%
F1 macro84.8%
Micro precision86.1%
Micro recall85.0%
Instances263
True positives2,052
False positives331
False negatives363

Speed

Model time, all inputs12936.4 s
Mean per input49.19 s
Slowest input788.73 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens401.3K
Output tokens236.6K
Total tokens637.9K
Input cost$0.8027
Output cost$1.89
Total cost$2.70

Priced from 2025-10-01 · Input $/M $2.00 · Output $/M $8.00

Result 105 of 126

Test T0208 at 2025-10-01

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

genai gemini-2.5-flash-lite · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

{
  "type": {
    "type": "Reference"
  },
  "author": {
    "last_name": "Montaghemi",
    "first_name": "Nourhomadeh"
  },
  "publication": {
    "title": "Sonnenbichler - Montaghemi",
    "year": 
    "place"
    ,
    "pages": "",
    "publisher": "",
    "format": ""
  },
  "library_reference": {
    "shelfmark": "",
    "subjects": ""
  }
}

Scoring

F1 micro69.7%
F1 macro68.9%
Micro precision71.3%
Micro recall68.2%
Instances263
True positives1,648
False positives663
False negatives767

Speed

Model time, all inputs587.8 s
Mean per input2.24 s
Slowest input111.93 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens185.4K
Output tokens214.7K
Total tokens400.1K
Input cost$0.0185
Output cost$0.0859
Total cost$0.1044

Priced from 2025-10-01 · Input $/M $0.10 · Output $/M $0.40

Result 106 of 126

Test T0165 at 2025-10-01

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-5 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

F1 micro87.0%
F1 macro86.3%
Micro precision87.3%
Micro recall86.7%
Instances263
True positives2,095
False positives305
False negatives320

Speed

Model time, all inputs21108.3 s
Mean per input80.26 s
Slowest input1912.47 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens364.8K
Output tokens666.8K
Total tokens1M
Input cost$0.456
Output cost$6.67
Total cost$7.12

Priced from 2025-10-01 · Input $/M $1.25 · Output $/M $10.00

Result 107 of 126

Test T0146 at 2025-10-01

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

anthropic claude-opus-4-1-20250805 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

F1 micro82.6%
F1 macro81.5%
Micro precision81.8%
Micro recall83.4%
Instances263
True positives2,014
False positives448
False negatives401

Speed

Model time, all inputs1581.4 s
Mean per input6.01 s
Slowest input10.10 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens566.6K
Output tokens54.3K
Total tokens620.8K
Input cost$8.50
Output cost$4.07
Total cost$12.57

Priced from 2025-10-01 · Input $/M $15.00 · Output $/M $75.00

Result 108 of 126

Test T0167 at 2025-10-01

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-5-nano · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

F1 micro77.1%
F1 macro76.6%
Micro precision76.0%
Micro recall78.1%
Instances263
True positives1,887
False positives596
False negatives528

Speed

Model time, all inputs5049.1 s
Mean per input19.20 s
Slowest input83.88 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens434.6K
Output tokens807.7K
Total tokens1.2M
Input cost$0.0217
Output cost$0.3231
Total cost$0.3448

Priced from 2025-10-01 · Input $/M $0.05 · Output $/M $0.40

Result 109 of 126

Test T0166 at 2025-10-01

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-5-mini · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

F1 micro79.1%
F1 macro78.9%
Micro precision77.2%
Micro recall81.1%
Instances263
True positives1,959
False positives579
False negatives456

Speed

Model time, all inputs22744.8 s
Mean per input86.48 s
Slowest input783.16 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens387.5K
Output tokens609.5K
Total tokens997K
Input cost$0.0969
Output cost$1.22
Total cost$1.32

Priced from 2025-10-01 · Input $/M $0.25 · Output $/M $2.00

Result 110 of 126

Test T0155 at 2025-10-01

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

genai gemini-2.5-pro · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

{
  "type": {
    "type": "Reference"
  },
  "author": {
    "last_name": "Montaghemi",
    "first_name": "Nourhomadeh"
  },
  "publication": {
    "title": "Sonnenbichler - Montaghemi, Nourhomadeh",
    "year": 
    "place",
    "pages": "",
    "publisher": "",
    "format": ""
  },
  "library_reference": {
    "shelfmark": "",
    "subjects": ""
  }
}

Scoring

F1 micro87.0%
F1 macro86.5%
Micro precision85.4%
Micro recall88.7%
Instances263
True positives2,141
False positives367
False negatives274

Speed

Model time, all inputs3744.4 s
Mean per input14.24 s
Slowest input31.98 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens185.4K
Output tokens47.8K
Total tokens233.2K
Input cost$0.2318
Output cost$0.4777
Total cost$0.7095

Priced from 2025-10-01 · Input $/M $1.25 · Output $/M $10.00