RISE Humanities Data Benchmark, 0.6.1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'library_cards__true' with Search Hidden 'False' returned 188 results, showing page 15 of 19.
Result 141 of 188

Test T0264 at 2026-01-25

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter qwen/qwen3-vl-8b-instruct · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision50.0%
Recall33.3%
True positives1
False positives1
False negatives2
F1 score40.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":""},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro70.1%
F1 macro61.6%
Micro precision78.2%
Micro recall63.6%
Instances263
True positives1,536
False positives429
False negatives879

Speed

Model time, all inputs26877.6 s
Mean per input102.20 s
Slowest input1085.78 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens439K
Output tokens42.6K
Total tokens481.6K
Input cost$0.0351
Output cost$0.0213
Total cost$0.0564

Priced from 2026-01-23 · Input $/M $0.08 · Output $/M $0.50

Result 142 of 188

Test T0548 at 2026-01-25

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

mistral ministral-14b-2512 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision0.0%
Recall0.0%
True positives0
False positives0
False negatives3
F1 score0.0%
Total fields13

{
  "$schema": "http://json-schema.org/draft-07/schema#",
  "$defs": {
    "Author": {
      "description": "Represents an author on a library card.",
      "properties": {
        "last_name": {
          "title": "Last Name",
          "type": "string"
        },
        "first_name": {
          "title": "First Name",
          "type": "string"
        }
      },
      "required": ["last_name"],
      "title": "Author",
      "type": "object"
    },
    "LibraryReference": {
      "description": "Represents library-specific reference information.",
      "properties": {
        "shelfmark": {
          "type": "string",
          "title": "Shelfmark"
        },
        "subjects": {
          "type": "string",
          "title": "Subjects"
        }
      },
      "title": "LibraryReference",
      "type": "object"
    },
    "Publication": {
      "description": "Represents publication information from a library card.",
      "properties": {
        "title": {
          "type": "string",
          "title": "Title"
        },
        "year": {
          "type": "integer",
          "title": "Year"
        },
        "place": {
          "type": "string",
          "title": "Place"
        },
        "pages": {
          "type": "string",
          "title": "Pages"
        },
        "publisher": {
          "type": "string",
          "title": "Publisher"
        },
        "format": {
          "type": "string",
          "title": "Format"
        }
      },
      "required": ["title"],
      "title": "Publication",
      "type": "object"
    },
    "WorkType": {
      "description": "Represents the type of work referenced in a library card.",
      "properties": {
        "type": {
          "enum": ["Dissertation or thesis", "Reference"],
          "title": "Type",
          "type": "string"
        }
      },
      "required": ["type"],
      "title": "WorkType",
      "type": "object"
    }
  },
  "type": {
    "type": "Dissertation or thesis"
  },
  "author": {
    "last_name": "",
    "first_name": ""
  },
  "publication": {
    "title": "",
    "year": 0,
    "place": "",
    "pages": "",
    "publisher": "",
    "format": ""
  },
  "library_reference": {
    "shelfmark": "",
    "subjects": ""
  }
}

Scoring

F1 micro1.9%
F1 macro0.6%
Micro precision2.0%
Micro recall1.8%
Instances263
True positives43
False positives2,075
False negatives2,372

Speed

Model time, all inputs1226.4 s
Mean per input4.66 s
Slowest input11.35 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens314.5K
Output tokens161.6K
Total tokens476.1K
Input cost$0.0629
Output cost$0.0323
Total cost$0.0952

Priced from 2026-01-23 · Input $/M $0.20 · Output $/M $0.20

Result 143 of 188

Test T0504 at 2026-01-25

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

genai gemini-3-flash-preview · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro86.6%
F1 macro85.6%
Micro precision82.9%
Micro recall90.6%
Instances263
True positives2,187
False positives450
False negatives228

Speed

Model time, all inputs11402.7 s
Mean per input43.36 s
Slowest input385.97 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens406.2K
Output tokens38.3K
Total tokens444.5K
Input cost$0.2031
Output cost$0.1148
Total cost$6.35
Reasoning cost usd$6.03
Total reasoning tokens2M

Priced from 2026-01-23 · Input $/M $0.50 · Output $/M $3.00

Result 144 of 188

Test T0162 at 2026-01-25

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-4.1-nano-2025-04-14 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision0.0%
Recall0.0%
True positives0
False positives3
False negatives3
F1 score0.0%
Total fields12

{"type":{"type":"Dissertation or thesis"},"author":{"last_name":"Montaghem","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro65.9%
F1 macro64.5%
Micro precision74.3%
Micro recall59.1%
Instances263
True positives1,428
False positives494
False negatives987

Speed

Model time, all inputs826.6 s
Mean per input3.14 s
Slowest input205.00 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens642.8K
Output tokens25.7K
Total tokens668.5K
Input cost$0.0643
Output cost$0.0103
Total cost$0.0746

Priced from 2026-01-23 · Input $/M $0.10 · Output $/M $0.40

Result 145 of 188

Test T0408 at 2026-01-25

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-5.1-2025-11-13 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro84.7%
F1 macro84.2%
Micro precision85.2%
Micro recall84.2%
Instances263
True positives2,033
False positives353
False negatives382

Speed

Model time, all inputs4935.0 s
Mean per input18.76 s
Slowest input610.15 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens421.9K
Output tokens135.1K
Total tokens557K
Input cost$0.5273
Output cost$1.35
Total cost$1.88

Priced from 2026-01-23 · Input $/M $1.25 · Output $/M $10.00

Result 146 of 188

Test T0164 at 2026-01-25

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openai gpt-4o-mini-2024-07-18 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision33.3%
Recall33.3%
True positives1
False positives2
False negatives2
F1 score33.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro2.2%
F1 macro1.5%
Micro precision54.0%
Micro recall1.1%
Instances263
True positives27
False positives23
False negatives2,388

Speed

Model time, all inputs82428.2 s
Mean per input313.42 s
Slowest input1021.38 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens6.5M
Output tokens37K
Total tokens6.6M
Input cost$0.9803
Output cost$0.0222
Total cost$1.00

Priced from 2026-01-23 · Input $/M $0.15 · Output $/M $0.60

Result 147 of 188

Test T0515 at 2026-01-24

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

anthropic claude-opus-4-5-20251101 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision20.0%
Recall33.3%
True positives1
False positives4
False negatives2
F1 score25.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"<UNKNOWN>","place":null,"pages":null,"publisher":null,"format":null,"editor":null},"library_reference":{"shelfmark":null,"subjects":null}}

Scoring

F1 micro86.3%
F1 macro85.7%
Micro precision82.9%
Micro recall90.1%
Instances263
True positives2,175
False positives449
False negatives240

Speed

Model time, all inputs1477.1 s
Mean per input5.62 s
Slowest input9.80 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens706K
Output tokens56K
Total tokens762K
Input cost$3.53
Output cost$1.40
Total cost$4.93

Priced from 2026-01-23 · Input $/M $5.00 · Output $/M $25.00

Result 148 of 188

Test T0208 at 2026-01-24

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

genai gemini-2.5-flash-lite · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"","place":"","pages":"","publisher":"","format":"","editor":null},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro69.4%
F1 macro68.3%
Micro precision73.0%
Micro recall66.2%
Instances263
True positives1,598
False positives590
False negatives817

Speed

Model time, all inputs559.5 s
Mean per input2.13 s
Slowest input150.03 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens187.3K
Output tokens175.3K
Total tokens362.5K
Input cost$0.0187
Output cost$0.0701
Total cost$0.0888

Priced from 2026-01-23 · Input $/M $0.10 · Output $/M $0.40

Result 149 of 188

Test T0537 at 2026-01-24

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

mistral mistral-large-2512 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision0.0%
Recall0.0%
True positives0
False positives0
False negatives3
F1 score0.0%
Total fields13

{
  "type": {
    "type": ""
  },
  "author": {
    "last_name": "",
    "first_name": ""
  },
  "publication": {
    "title": "",
    "year": "",
    "place": "",
    "pages": "",
    "publisher": "",
    "format": ""
  },
  "library_reference": {
    "shelfmark": "",
    "subjects": ""
  }
}

Scoring

F1 micro0.0%
F1 macro0.0%
Micro precision0.0%
Micro recall0.0%
Instances263
True positives0
False positives0
False negatives2,415

Speed

Model time, all inputs460.5 s
Mean per input1.75 s
Slowest input4.75 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens314.5K
Output tokens24.7K
Total tokens339.3K
Input cost$0.1573
Output cost$0.0371
Total cost$0.1944

Priced from 2026-01-23 · Input $/M $0.50 · Output $/M $1.50

Result 150 of 188

Test T0252 at 2026-01-24

index-card information-extraction typed, printed, handwritten 20, 19 de, fr, en, la, el, fi, sv, pl index bibliographic

openrouter meta-llama/llama-4-maverick · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Precision25.0%
Recall33.3%
True positives1
False positives3
False negatives2
F1 score29.0%
Total fields12

{"type":{"type":"Reference"},"author":{"last_name":"Montaghemi","first_name":"Nourhomadeh"},"publication":{"title":"Sonnenbichler - Montaghemi, Nourhomadeh","year":"","place":"","pages":"","publisher":"","format":"","editor":""},"library_reference":{"shelfmark":"","subjects":""}}

Scoring

F1 micro70.9%
F1 macro60.8%
Micro precision81.3%
Micro recall62.9%
Instances263
True positives1,519
False positives349
False negatives896

Speed

Model time, all inputs31033.1 s
Mean per input118.00 s
Slowest input737.02 s
Inputs timed263

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens482.5K
Output tokens176.7K
Total tokens659.2K
Input cost$0.0936
Output cost$0.1087
Total cost$0.2022

Priced from 2026-01-23 · Input $/M $0.15 · Output $/M $0.60