RISE Humanities Data Benchmark, 0.5.5

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'library_cards__true' with Search Hidden 'False' returned 126 results, showing page 12 of 13.
Result 111 of 126

Test T0230 at 2025-09-30

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

anthropic claude-sonnet-4-5-20250929 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

F1 micro80.6%
F1 macro79.3%
Micro precision79.8%
Micro recall81.5%
Instances263
True positives1,968
False positives499
False negatives447

Speed

Model time, all inputs1428.8 s
Mean per input5.43 s
Slowest input54.98 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens638.4K
Output tokens57.2K
Total tokens695.6K
Input cost$1.92
Output cost$0.8579
Total cost$2.77

Priced from 2025-09-30 · Input $/M $3.00 · Output $/M $15.00

Result 112 of 126

Test T0200 at 2025-09-30

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

genai gemini-2.5-flash · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro85.7%
F1 macro85.0%
Micro precision83.5%
Micro recall88.1%
Instances263
True positives2,128
False positives422
False negatives287

Speed

Model time, all inputs3027.5 s
Mean per input11.51 s
Slowest input53.14 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens185.4K
Output tokens45.1K
Total tokens230.5K
Input cost$0.0556
Output cost$0.1128
Total cost$0.1684

Priced from 2025-09-30 · Input $/M $0.30 · Output $/M $2.50

Result 113 of 126

Test T0160 at 2025-09-30

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-4.1 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

F1 micro85.9%
F1 macro85.4%
Micro precision86.7%
Micro recall85.2%
Instances263
True positives2,057
False positives316
False negatives358

Speed

Model time, all inputs11597.0 s
Mean per input44.10 s
Slowest input607.63 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens401.3K
Output tokens245.2K
Total tokens646.5K
Input cost$0.8027
Output cost$1.96
Total cost$2.76

Priced from 2025-09-30 · Input $/M $2.00 · Output $/M $8.00

Result 114 of 126

Test T0066 at 2025-09-30

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-4o · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

F1 micro82.5%
F1 macro82.1%
Micro precision82.5%
Micro recall82.5%
Instances263
True positives1,992
False positives423
False negatives423

Speed

Model time, all inputs3291.1 s
Mean per input12.51 s
Slowest input319.51 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens404.8K
Output tokens29.1K
Total tokens433.9K
Input cost$1.01
Output cost$0.2912
Total cost$1.30

Priced from 2025-09-30 · Input $/M $2.50 · Output $/M $10.00

Result 115 of 126

Test T0162 at 2025-09-24

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-4.1-nano · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

F1 micro61.2%
F1 macro60.4%
Micro precision65.3%
Micro recall57.6%
Instances263
True positives1,390
False positives738
False negatives1,025

Speed

Model time, all inputs1867.0 s
Mean per input7.10 s
Slowest input304.95 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

No cost recorded for this run

Result 116 of 126

Test T0167 at 2025-09-02

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-5-nano · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

F1 micro80.2%
F1 macro79.7%
Micro precision78.9%
Micro recall81.5%
Instances263
True positives1,969
False positives526
False negatives446

Speed

Model time, all inputs4991.5 s
Mean per input18.98 s
Slowest input36.17 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

No cost recorded for this run

Result 117 of 126

Test T0066 at 2025-09-02

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-4o · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

F1 micro89.6%
F1 macro89.4%
Micro precision91.3%
Micro recall88.0%
Instances263
True positives2,125
False positives203
False negatives290

Speed

Model time, all inputs1291.3 s
Mean per input4.91 s
Slowest input12.29 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

No cost recorded for this run

Result 118 of 126

Test T0166 at 2025-09-02

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-5-mini · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

F1 micro84.4%
F1 macro84.1%
Micro precision90.2%
Micro recall79.3%
Instances263
True positives1,916
False positives209
False negatives499

Speed

Model time, all inputs4748.4 s
Mean per input18.05 s
Slowest input41.21 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

No cost recorded for this run

Result 119 of 126

Test T0161 at 2025-09-02

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-4.1-mini · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

F1 micro87.0%
F1 macro86.7%
Micro precision88.8%
Micro recall85.3%
Instances263
True positives2,061
False positives261
False negatives354

Speed

Model time, all inputs796.3 s
Mean per input3.03 s
Slowest input7.14 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

No cost recorded for this run

Result 120 of 126

Test T0146 at 2025-09-02

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

anthropic claude-opus-4-1-20250805 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

```json
{
  "type": {
    "type": "Reference"
  },
  "author": {
    "last_name": "Montaghemi",
    "first_name": "Nourhomadeh"
  },
  "publication": {
    "title": "",
    "year": "",
    "place": "",
    "pages": "",
    "publisher": "",
    "format": "",
    "reprint_note": ""
  },
  "examination": {
    "place": "",
    "year": ""
  },
  "library_reference": {
    "shelfmark": "",
    "publication_number": "",
    "subjects": ""
  }
}
```

Scoring

F1 micro84.5%
F1 macro83.3%
Micro precision86.4%
Micro recall82.7%
Instances263
True positives1,996
False positives315
False negatives419

Speed

Model time, all inputs1903.0 s
Mean per input7.24 s
Slowest input27.20 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

No cost recorded for this run