RISE Humanities Data Benchmark, 0.6.1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'library_cards__true' with Search Hidden 'False' returned 187 results, showing page 14 of 19.
Result 131 of 187

Test T0657 at 2026-03-16

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-5.4-2026-03-05 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro83.1%
F1 macro82.4%
Micro precision80.5%
Micro recall85.8%
Instances263
True positives2,073
False positives501
False negatives342

Speed

Model time, all inputs2346.8 s
Mean per input8.92 s
Slowest input602.89 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens444.7K
Output tokens55.7K
Total tokens500.4K
Input cost$1.11
Output cost$0.836
Total cost$1.95

Priced from 2026-03-16 · Input $/M $2.50 · Output $/M $15.00

Result 132 of 187

Test T0645 at 2026-03-16

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

anthropic claude-sonnet-4-6 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision20.0%
Recall33.3%
True positives1
False positives4
False negatives2
F1 score25.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"<UNKNOWN>","place":null,"pages":null,"publisher":null,"format":null,"editor":null},"library_reference":{"shelfmark":null,"subjects":null}}

Scoring

F1 micro87.1%
F1 macro86.3%
Micro precision82.4%
Micro recall92.5%
Instances263
True positives2,233
False positives477
False negatives182

Speed

Model time, all inputs1385.0 s
Mean per input5.27 s
Slowest input15.94 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens706.2K
Output tokens57.3K
Total tokens763.5K
Input cost$2.12
Output cost$0.8598
Total cost$2.98

Priced from 2026-03-16 · Input $/M $3.00 · Output $/M $15.00

Result 133 of 187

Test T0504 at 2026-03-12

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

genai gemini-3-flash-preview · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"","place":"","pages":"","publisher":"","format":"","editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro86.1%
F1 macro85.1%
Micro precision82.2%
Micro recall90.3%
Instances263
True positives2,181
False positives471
False negatives234

Speed

Model time, all inputs14239.9 s
Mean per input54.14 s
Slowest input298.78 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens406.2K
Output tokens39.9K
Total tokens446.1K
Input cost$0.2031
Output cost$0.1196
Total cost$9.62
Reasoning cost usd$9.29
Total reasoning tokens3.1M

Priced from 2026-03-02 · Input $/M $0.50 · Output $/M $3.00

Result 134 of 187

Test T0166 at 2026-02-17

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-5-mini-2025-08-07 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":"Sonnenbichler - Montaghemi, Nourhomadeh"}}

Scoring

F1 micro62.9%
F1 macro51.1%
Micro precision79.7%
Micro recall51.9%
Instances263
True positives1,254
False positives320
False negatives1,161

Speed

Model time, all inputs55323.3 s
Mean per input210.35 s
Slowest input16015.51 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens282.3K
Output tokens467.4K
Total tokens749.8K
Input cost$0.0706
Output cost$0.9349
Total cost$1.01

Priced from 2026-01-23 · Input $/M $0.25 · Output $/M $2.00

Result 135 of 187

Test T0168 at 2026-01-26

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai o3-2025-04-16 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro83.6%
F1 macro83.0%
Micro precision84.1%
Micro recall83.1%
Instances263
True positives2,006
False positives378
False negatives409

Speed

Model time, all inputs14263.6 s
Mean per input54.23 s
Slowest input1218.46 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens433.7K
Output tokens190.9K
Total tokens624.6K
Input cost$0.8674
Output cost$1.53
Total cost$2.39

Priced from 2026-01-23 · Input $/M $2.00 · Output $/M $8.00

Result 136 of 187

Test T0493 at 2026-01-26

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-5.2-2025-12-11 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision20.0%
Recall33.3%
True positives1
False positives4
False negatives2
F1 score25.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"place  ","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro46.9%
F1 macro32.5%
Micro precision82.2%
Micro recall32.8%
Instances263
True positives792
False positives172
False negatives1,623

Speed

Model time, all inputs79056.4 s
Mean per input300.59 s
Slowest input1814.02 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens185.5K
Output tokens56K
Total tokens241.5K
Input cost$0.3246
Output cost$0.7835
Total cost$1.11

Priced from 2026-01-23 · Input $/M $1.75 · Output $/M $14.00

Result 137 of 187

Test T0165 at 2026-01-26

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-5 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision0.0%
Recall0.0%
True positives0
False positives0
False negatives3
F1 score0.0%
Total fields13

Scoring

F1 micro71.1%
F1 macro59.6%
Micro precision87.3%
Micro recall59.9%
Instances263
True positives1,447
False positives210
False negatives968

Speed

Model time, all inputs21125.9 s
Mean per input80.33 s
Slowest input1226.89 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens287.1K
Output tokens414.4K
Total tokens701.5K
Input cost$0.3589
Output cost$4.14
Total cost$4.50

Priced from 2026-01-23 · Input $/M $1.25 · Output $/M $10.00

Result 138 of 187

Test T0167 at 2026-01-26

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-5-nano-2025-08-07 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro73.8%
F1 macro70.2%
Micro precision77.0%
Micro recall70.8%
Instances263
True positives1,709
False positives510
False negatives706

Speed

Model time, all inputs8006.2 s
Mean per input30.44 s
Slowest input617.52 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens447K
Output tokens669K
Total tokens1.1M
Input cost$0.0223
Output cost$0.2676
Total cost$0.2899

Priced from 2026-01-23 · Input $/M $0.05 · Output $/M $0.40

Result 139 of 187

Test T0559 at 2026-01-25

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

mistral ministral-8b-2512 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision0.0%
Recall0.0%
True positives0
False positives0
False negatives3
F1 score0.0%
Total fields13

{
  "type": {
    "type": "Dissertation or thesis"
  },
  "author": {
    "last_name": "Smith",
    "first_name": "John"
  },
  "publication": {
    "title": "The Role of Industrialization in 19th Century Economic Growth",
    "year": 1895,
    "place": "Berlin",
    "pages": "210",
    "publisher": "Universitätsverlag Berlin",
    "format": "8°"
  },
  "library_reference": {
    "shelfmark": "Diss. 1234/56",
    "subjects": "Industrialization, Economic History, 19th Century"
  }
}

Scoring

F1 micro0.0%
F1 macro0.0%
Micro precision0.0%
Micro recall0.0%
Instances263
True positives0
False positives0
False negatives2,415

Speed

Model time, all inputs343.8 s
Mean per input1.31 s
Slowest input4.22 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens314.5K
Output tokens52.9K
Total tokens367.5K
Input cost$0.0472
Output cost$0.0079
Total cost$0.0551

Priced from 2026-01-23 · Input $/M $0.15 · Output $/M $0.15

Result 140 of 187

Test T0548 at 2026-01-25

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

mistral ministral-14b-2512 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision0.0%
Recall0.0%
True positives0
False positives0
False negatives3
F1 score0.0%
Total fields13

{
  "$schema": "http://json-schema.org/draft-07/schema#",
  "$defs": {
    "Author": {
      "description": "Represents an author on a library card.",
      "properties": {
        "last_name": {
          "title": "Last Name",
          "type": "string"
        },
        "first_name": {
          "title": "First Name",
          "type": "string"
        }
      },
      "required": ["last_name"],
      "title": "Author",
      "type": "object"
    },
    "LibraryReference": {
      "description": "Represents library-specific reference information.",
      "properties": {
        "shelfmark": {
          "type": "string",
          "title": "Shelfmark"
        },
        "subjects": {
          "type": "string",
          "title": "Subjects"
        }
      },
      "title": "LibraryReference",
      "type": "object"
    },
    "Publication": {
      "description": "Represents publication information from a library card.",
      "properties": {
        "title": {
          "type": "string",
          "title": "Title"
        },
        "year": {
          "type": "integer",
          "title": "Year"
        },
        "place": {
          "type": "string",
          "title": "Place"
        },
        "pages": {
          "type": "string",
          "title": "Pages"
        },
        "publisher": {
          "type": "string",
          "title": "Publisher"
        },
        "format": {
          "type": "string",
          "title": "Format"
        }
      },
      "required": ["title"],
      "title": "Publication",
      "type": "object"
    },
    "WorkType": {
      "description": "Represents the type of work referenced in a library card.",
      "properties": {
        "type": {
          "enum": ["Dissertation or thesis", "Reference"],
          "title": "Type",
          "type": "string"
        }
      },
      "required": ["type"],
      "title": "WorkType",
      "type": "object"
    }
  },
  "type": {
    "type": "Dissertation or thesis"
  },
  "author": {
    "last_name": "",
    "first_name": ""
  },
  "publication": {
    "title": "",
    "year": 0,
    "place": "",
    "pages": "",
    "publisher": "",
    "format": ""
  },
  "library_reference": {
    "shelfmark": "",
    "subjects": ""
  }
}

Scoring

F1 micro1.9%
F1 macro0.6%
Micro precision2.0%
Micro recall1.8%
Instances263
True positives43
False positives2,075
False negatives2,372

Speed

Model time, all inputs1226.4 s
Mean per input4.66 s
Slowest input11.35 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens314.5K
Output tokens161.6K
Total tokens476.1K
Input cost$0.0629
Output cost$0.0323
Total cost$0.0952

Priced from 2026-01-23 · Input $/M $0.20 · Output $/M $0.20