RISE Humanities Data Benchmark, 0.6.1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'library_cards__true' with Search Hidden 'False' returned 188 results, showing page 17 of 19.
Result 161 of 188

Test T0200 at 2025-11-24

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

genai gemini-2.5-flash · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro86.6%
F1 macro86.0%
Micro precision84.6%
Micro recall88.7%
Instances263
True positives2,141
False positives389
False negatives274

Speed

Model time, all inputs1985.2 s
Mean per input7.55 s
Slowest input25.85 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens187.3K
Output tokens38.2K
Total tokens225.5K
Input cost$0.0562
Output cost$0.0955
Total cost$0.9765
Reasoning cost usd$0.8248
Total reasoning tokens329.9K

Priced from 2025-11-24 · Input $/M $0.30 · Output $/M $2.50

Result 162 of 188

Test T0252 at 2025-10-17

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter meta-llama/llama-4-maverick · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

F1 micro67.5%
F1 macro62.3%
Micro precision78.2%
Micro recall59.3%
Instances263
True positives1,433
False positives399
False negatives982

Speed

Model time, all inputs19155.5 s
Mean per input72.83 s
Slowest input701.36 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens390.9K
Output tokens819.9K
Total tokens1.2M
Input cost$0.0586
Output cost$0.492
Total cost$0.5506

Priced from 2025-10-17 · Input $/M $0.15 · Output $/M $0.60

Result 163 of 188

Test T0247 at 2025-10-17

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter qwen/qwen3-vl-8b-thinking · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

F1 micro10.3%
F1 macro5.9%
Micro precision32.7%
Micro recall6.1%
Instances263
True positives148
False positives305
False negatives2,267

Speed

Model time, all inputs1036.0 s
Mean per input3.94 s
Slowest input41.03 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens495.2K
Output tokens44.6K
Total tokens539.8K
Input cost$0.0891
Output cost$0.0936
Total cost$0.1827

Priced from 2025-10-17 · Input $/M $0.18 · Output $/M $2.10

Result 164 of 188

Test T0164 at 2025-10-03

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-4o-mini · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

F1 micro1.5%
F1 macro1.0%
Micro precision6.8%
Micro recall0.8%
Instances263
True positives20
False positives275
False negatives2,395

Speed

Model time, all inputs1229.5 s
Mean per input245.89 s
Slowest input607.93 s
Inputs timed5

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens131.3K
Output tokens727
Total tokens132K
Input cost$0.0197
Output cost$0.0004
Total cost$0.0201

Priced from 2025-10-01 · Input $/M $0.15 · Output $/M $0.60

Result 165 of 188

Test T0162 at 2025-10-02

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-4.1-nano · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

F1 micro67.6%
F1 macro66.7%
Micro precision72.2%
Micro recall63.6%
Instances263
True positives1,536
False positives592
False negatives879

Speed

Model time, all inputs471.3 s
Mean per input1.79 s
Slowest input5.28 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

No cost recorded for this run

Priced from 2025-10-01 · Input $/M $0.10 · Output $/M $0.40

Result 166 of 188

Test T0161 at 2025-10-02

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-4.1-mini · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

F1 micro75.9%
F1 macro75.7%
Micro precision75.8%
Micro recall75.9%
Instances263
True positives1,834
False positives584
False negatives581

Speed

Model time, all inputs29647.3 s
Mean per input112.73 s
Slowest input1506.66 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

No cost recorded for this run

Priced from 2025-10-01 · Input $/M $0.40 · Output $/M $1.60

Result 167 of 188

Test T0168 at 2025-10-01

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai o3 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

F1 micro83.4%
F1 macro83.1%
Micro precision83.7%
Micro recall83.1%
Instances263
True positives2,006
False positives390
False negatives409

Speed

Model time, all inputs9024.0 s
Mean per input34.31 s
Slowest input620.80 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens376.6K
Output tokens223.3K
Total tokens599.9K
Input cost$0.7532
Output cost$1.79
Total cost$2.54

Priced from 2025-10-01 · Input $/M $2.00 · Output $/M $8.00

Result 168 of 188

Test T0160 at 2025-10-01

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-4.1 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

F1 micro85.5%
F1 macro84.8%
Micro precision86.1%
Micro recall85.0%
Instances263
True positives2,052
False positives331
False negatives363

Speed

Model time, all inputs12936.4 s
Mean per input49.19 s
Slowest input788.73 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens401.3K
Output tokens236.6K
Total tokens637.9K
Input cost$0.8027
Output cost$1.89
Total cost$2.70

Priced from 2025-10-01 · Input $/M $2.00 · Output $/M $8.00

Result 169 of 188

Test T0208 at 2025-10-01

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

genai gemini-2.5-flash-lite · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

{
  "type": {
    "type": "Reference"
  },
  "author": {
    "last_name": "Montaghemi",
    "first_name": "Nourhomadeh"
  },
  "publication": {
    "title": "Sonnenbichler - Montaghemi",
    "year": 
    "place"
    ,
    "pages": "",
    "publisher": "",
    "format": ""
  },
  "library_reference": {
    "shelfmark": "",
    "subjects": ""
  }
}

Scoring

F1 micro69.7%
F1 macro68.9%
Micro precision71.3%
Micro recall68.2%
Instances263
True positives1,648
False positives663
False negatives767

Speed

Model time, all inputs587.8 s
Mean per input2.24 s
Slowest input111.93 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens185.4K
Output tokens214.7K
Total tokens400.1K
Input cost$0.0185
Output cost$0.0859
Total cost$0.1044

Priced from 2025-10-01 · Input $/M $0.10 · Output $/M $0.40

Result 170 of 188

Test T0165 at 2025-10-01

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-5 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

F1 micro87.0%
F1 macro86.3%
Micro precision87.3%
Micro recall86.7%
Instances263
True positives2,095
False positives305
False negatives320

Speed

Model time, all inputs21108.3 s
Mean per input80.26 s
Slowest input1912.47 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens364.8K
Output tokens666.8K
Total tokens1M
Input cost$0.456
Output cost$6.67
Total cost$7.12

Priced from 2025-10-01 · Input $/M $1.25 · Output $/M $10.00