RISE Humanities Data Benchmark, 0.5.5

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'library_cards__true' with Search Hidden 'False' returned 126 results, showing page 13 of 13.
Result 121 of 126

Test T0165 at 2025-09-02

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-5 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

F1 micro89.7%
F1 macro89.5%
Micro precision91.7%
Micro recall87.9%
Instances263
True positives2,122
False positives193
False negatives293

Speed

Model time, all inputs12744.7 s
Mean per input48.46 s
Slowest input342.21 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

No cost recorded for this run

Result 122 of 126

Test T0160 at 2025-09-02

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-4.1 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

F1 micro89.7%
F1 macro89.4%
Micro precision91.2%
Micro recall88.2%
Instances263
True positives2,131
False positives205
False negatives284

Speed

Model time, all inputs1883.8 s
Mean per input7.16 s
Slowest input49.47 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

No cost recorded for this run

Result 123 of 126

Test T0167 at 2025-09-02

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-5-nano · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

F1 micro80.2%
F1 macro79.7%
Micro precision78.9%
Micro recall81.5%
Instances263
True positives1,969
False positives526
False negatives446

Speed

Model time, all inputs4991.5 s
Mean per input18.98 s
Slowest input36.17 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

No cost recorded for this run

Result 124 of 126

Test T0166 at 2025-09-02

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-5-mini · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

F1 micro84.4%
F1 macro84.1%
Micro precision90.2%
Micro recall79.3%
Instances263
True positives1,916
False positives209
False negatives499

Speed

Model time, all inputs4748.4 s
Mean per input18.05 s
Slowest input41.21 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

No cost recorded for this run

Result 125 of 126

Test T0155 at 2025-09-02

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

genai gemini-2.5-pro · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

```json
{
  "type": {
    "type": "Reference"
  },
  "author": {
    "last_name": "Montaghemi",
    "first_name": "Nourhomadeh"
  },
  "publication": {
    "title": "",
    "place": "",
    "pages": "",
    "publisher": "",
    "format": "",
    "reprint_note": ""
  },
  "examination": {
    "place": ""
  },
  "library_reference": {
    "shelfmark": "",
    "publication_number": "",
    "subjects": ""
  }
}
```

Scoring

F1 micro88.6%
F1 macro88.0%
Micro precision89.6%
Micro recall87.6%
Instances263
True positives2,115
False positives245
False negatives300

Speed

Model time, all inputs5030.7 s
Mean per input19.13 s
Slowest input60.38 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

No cost recorded for this run

Result 126 of 126

Test T0164 at 2025-09-02

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-4o-mini · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

F1 micro80.7%
F1 macro80.3%
Micro precision78.7%
Micro recall82.9%
Instances263
True positives2,002
False positives543
False negatives413

Speed

Model time, all inputs1249.9 s
Mean per input4.75 s
Slowest input15.19 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

No cost recorded for this run