RISE Humanities Data Benchmark, 0.5.5

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'blacklist_cards__true' with Search Hidden 'False' returned 113 results, showing page 10 of 12.
Result 91 of 113

Test T0517 at 2026-01-24

index-card information-extraction typed, handwritten 20 de company

anthropic claude-opus-4-5-20251101 · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"company": {"transcription": "Abegg & Cie."}, "location": {"transcription": "Z\u00fcrich"}, "b_id": {"transcription": "B.51.322.GB.266"}, "date": "", "information": ""}

Scoring

Fuzzy score96.4%

Speed

Model time, all inputs191.0 s
Mean per input5.79 s
Slowest input7.86 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens93.2K
Output tokens6.4K
Total tokens99.6K
Input cost$0.4662
Output cost$0.16
Total cost$0.6262

Priced from 2026-01-23 · Input $/M $5.00 · Output $/M $25.00

Result 92 of 113

Test T0324 at 2026-01-24

index-card information-extraction typed, handwritten 20 de company

anthropic claude-opus-4-1-20250805 · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"company":{"transcription":"Abegg & Cie."},"location":{"transcription":"Zürich"},"b_id":{"transcription":"B.51.322.GB.266"},"date":null,"information":null}

Scoring

Fuzzy score91.4%

Speed

Model time, all inputs228.4 s
Mean per input6.92 s
Slowest input8.46 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens84.2K
Output tokens6.2K
Total tokens90.4K
Input cost$1.26
Output cost$0.4639
Total cost$1.73

Priced from 2026-01-23 · Input $/M $15.00 · Output $/M $75.00

Result 93 of 113

Test T0315 at 2026-01-24

index-card information-extraction typed, handwritten 20 de company

genai gemini-2.5-flash · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"company":{"transcription":"Abegg & Cie."},"location":{"transcription":"Zürich"},"b_id":{"transcription":"B.51.322.GB.266"},"date":null,"information":null}

Scoring

Fuzzy score92.6%

Speed

Model time, all inputs119.6 s
Mean per input3.62 s
Slowest input6.68 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens17.5K
Output tokens5K
Total tokens22.5K
Input cost$0.0052
Output cost$0.0126
Total cost$0.0178

Priced from 2026-01-23 · Input $/M $0.30 · Output $/M $2.50

Result 94 of 113

Test T0410 at 2025-11-25

index-card information-extraction typed, handwritten 20 de company

openai gpt-5.1-2025-11-13 · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score80.0%

{"company":{"transcription":"Abegg & Cie."},"location":{"transcription":"Zürich"},"b_id":{"transcription":"B.51.322.GB.266"},"date":"","information":[]}

Scoring

Fuzzy score80.2%

Speed

Model time, all inputs149.1 s
Mean per input4.52 s
Slowest input9.50 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens45.6K
Output tokens3.5K
Total tokens49.1K
Input cost$0.057
Output cost$0.0348
Total cost$0.0918

Priced from 2025-11-24 · Input $/M $1.25 · Output $/M $10.00

Result 95 of 113

Test T0315 at 2025-11-24

index-card information-extraction typed, handwritten 20 de company

genai gemini-2.5-flash · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score79.2%

{"company":{"transcription":"A begg & Cie."},"location":{"transcription":"Zürich"},"b_id":{"transcription":"B.51.322.GB.266"},"date":null,"information":[]}

Scoring

Fuzzy score90.2%

Speed

Model time, all inputs152.6 s
Mean per input4.62 s
Slowest input10.94 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens17.5K
Output tokens4.9K
Total tokens22.3K
Input cost$0.0052
Output cost$0.0122
Total cost$0.0175

Priced from 2025-11-24 · Input $/M $0.30 · Output $/M $2.50

Result 96 of 113

Test T0232 at 2025-10-24

index-card information-extraction typed, handwritten 20 de company

openai gpt-4.1 · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

Fuzzy score92.3%

Speed

Model time, all inputs185.6 s
Mean per input5.63 s
Slowest input12.63 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens41K
Output tokens3.8K
Total tokens44.8K
Input cost$0.082
Output cost$0.0305
Total cost$0.1125

Priced from 2025-10-01 · Input $/M $2.00 · Output $/M $8.00

Result 97 of 113

Test T0305 at 2025-10-24

index-card information-extraction typed, handwritten 20 de company

openai gpt-4o · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

Fuzzy score93.2%

Speed

Model time, all inputs186.4 s
Mean per input5.65 s
Slowest input18.25 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens41K
Output tokens3.1K
Total tokens44.1K
Input cost$0.1025
Output cost$0.0311
Total cost$0.1336

Priced from 2025-10-01 · Input $/M $2.50 · Output $/M $10.00

Result 98 of 113

Test T0324 at 2025-10-24

index-card information-extraction typed, handwritten 20 de company

anthropic claude-opus-4-1-20250805 · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

Fuzzy score88.6%

Speed

Model time, all inputs304.5 s
Mean per input9.23 s
Slowest input12.78 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens65.3K
Output tokens6.2K
Total tokens71.4K
Input cost$0.9791
Output cost$0.4625
Total cost$1.44

Priced from 2025-10-01 · Input $/M $15.00 · Output $/M $75.00

Result 99 of 113

Test T0334 at 2025-10-24

index-card information-extraction typed, handwritten 20 de company

openrouter qwen/qwen3-vl-30b-a3b-instruct · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

Fuzzy score69.5%

Speed

Model time, all inputs293.3 s
Mean per input8.89 s
Slowest input169.34 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens32.6K
Output tokens8.1K
Total tokens40.7K
Input cost$0.0065
Output cost$0.0056
Total cost$0.0122

Priced from 2025-10-20 · Input $/M $0.20 · Output $/M $0.70

Result 100 of 113

Test T0325 at 2025-10-24

index-card information-extraction typed, handwritten 20 de company

anthropic claude-sonnet-4-5-20250929 · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

Fuzzy score85.3%

Speed

Model time, all inputs194.1 s
Mean per input5.88 s
Slowest input8.33 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens74.3K
Output tokens6.6K
Total tokens80.8K
Input cost$0.2229
Output cost$0.0984
Total cost$0.3212

Priced from 2025-10-01 · Input $/M $3.00 · Output $/M $15.00