RISE Humanities Data Benchmark, 0.6.1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'blacklist_cards__true' with Search Hidden 'False' returned 180 results, showing page 4 of 18.
Result 31 of 180

Test T1412 at 2026-09-17

index-card information-extraction typed, handwritten 20 de company

anthropic claude-opus-5 · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score56.6%

{"company":{"transcription":"A b e g g  & Cie."},"location":{"transcription":"Zürich"},"b_id":{"transcription":"B.51.322.GB.266"},"date":"\"\"","information":[{"transcription":""}]}

Scoring

Fuzzy score81.2%

Speed

Model time, all inputs158.8 s
Mean per input4.81 s
Slowest input6.18 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens143.3K
Output tokens8.4K
Total tokens151.7K
Input cost$0.7165
Output cost$0.2104
Total cost$0.9268

Priced from 2026-09-16 · Input $/M $5.00 · Output $/M $25.00

Result 32 of 180

Test T0951 at 2026-09-17

index-card information-extraction typed, handwritten 20 de company

openrouter qwen/qwen3.5-397b-a17b · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score76.5%

{"company":{"transcription":"A b e g g & Cie."},"location":{"transcription":"Zürich"},"b_id":{"transcription":"B.51.322.GB.26"},"date":"","information":[]}

Scoring

Fuzzy score91.2%

Speed

Model time, all inputs4789.9 s
Mean per input145.15 s
Slowest input738.39 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens81.3K
Output tokens60.1K
Total tokens141.4K
Input cost$0.0447
Output cost$0.2104
Total cost$0.2551

Priced from 2026-09-16 · Input $/M $0.55 · Output $/M $3.50

Result 33 of 180

Test T0912 at 2026-09-17

index-card information-extraction typed, handwritten 20 de company

openrouter qwen/qwen3.5-122b-a10b · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score80.0%

{"company":{"transcription":"Abegg & Cie."},"location":{"transcription":"Zürich"},"b_id":{"transcription":"B.51.322.GB.266"},"date":"","information":[]}

Scoring

Fuzzy score84.6%

Speed

Model time, all inputs215.8 s
Mean per input6.54 s
Slowest input13.55 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens91.1K
Output tokens9.9K
Total tokens101K
Input cost$0.0237
Output cost$0.0206
Total cost$0.0443

Priced from 2026-09-16 · Input $/M $0.26 · Output $/M $2.08

Result 34 of 180

Test T0964 at 2026-09-17

index-card information-extraction typed, handwritten 20 de company

openrouter qwen/qwen3.5-plus-02-15 · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score77.1%

{"company":{"transcription":"A b e g g & Cie."},"location":{"transcription":"Zürich"},"b_id":{"transcription":"B.51.322.GB.266"},"date":"","information":[]}

Scoring

Fuzzy score92.6%

Speed

Model time, all inputs1468.1 s
Mean per input44.49 s
Slowest input171.50 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens97.6K
Output tokens72.3K
Total tokens169.8K
Input cost$0.0254
Output cost$0.1127
Total cost$0.1381

Priced from 2026-09-16 · Input $/M $0.26 · Output $/M $1.56

Result 35 of 180

Test T1135 at 2026-09-17

index-card information-extraction typed, handwritten 20 de company

openrouter stepfun/step-3.7-flash · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score80.0%

{"company":{"transcription":"Abegg & Cie."},"location":{"transcription":"Zürich"},"b_id":{"transcription":"B.51.322.GB.266"},"date":"","information":[]}

Scoring

Fuzzy score82.9%

Speed

Model time, all inputs839.1 s
Mean per input25.43 s
Slowest input64.20 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens47.4K
Output tokens129.5K
Total tokens176.9K
Input cost$0.0095
Output cost$0.1489
Total cost$0.1584

Priced from 2026-09-16 · Input $/M $0.20 · Output $/M $1.15

Result 36 of 180

Test T1197 at 2026-09-17

index-card information-extraction typed, handwritten 20 de company

anthropic claude-fable-5 · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score77.1%

{"company":{"transcription":"A b e g g & Cie."},"location":{"transcription":"Zürich"},"b_id":{"transcription":"B.51.322.GB.266"},"date":"","information":[]}

Scoring

Fuzzy score95.7%

Speed

Model time, all inputs216.6 s
Mean per input6.56 s
Slowest input7.94 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens143.4K
Output tokens8K
Total tokens151.4K
Input cost$1.43
Output cost$0.3996
Total cost$1.83

Priced from 2026-09-16 · Input $/M $10.00 · Output $/M $50.00

Result 37 of 180

Test T1682 at 2026-09-17

index-card information-extraction typed, handwritten 20 de company

anthropic claude-fable-5-1 · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score77.1%

{
  "company": {"transcription": "A b e g g & Cie."},
  "location": {"transcription": "Zürich"},
  "b_id": {"transcription": "B.51.322.GB.266"},
  "date": "",
  "information": []
}

Scoring

Fuzzy score94.8%

Speed

Model time, all inputs252.8 s
Mean per input7.66 s
Slowest input13.69 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens113.2K
Output tokens5.6K
Total tokens118.8K
Input cost$1.13
Output cost$0.2812
Total cost$1.41

Priced from 2026-09-08 · Input $/M $10.00 · Output $/M $50.00

Result 38 of 180

Test T0977 at 2026-09-17

index-card information-extraction typed, handwritten 20 de company

openrouter qwen/qwen3.5-flash-02-23 · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score80.0%

{"company":{"transcription":"Abegg & Cie."},"location":{"transcription":"Zürich"},"b_id":{"transcription":"B.51.322.GB.266"},"date":"","information":[]}

Scoring

Fuzzy score91.4%

Speed

Model time, all inputs711.2 s
Mean per input21.55 s
Slowest input52.32 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens97.7K
Output tokens59.8K
Total tokens157.5K
Input cost$0.0063
Output cost$0.0155
Total cost$0.0219

Priced from 2026-09-16 · Input $/M $0.065 · Output $/M $0.26

Result 39 of 180

Test T1697 at 2026-09-17

index-card information-extraction typed, handwritten 20 de company

openrouter meta/muse-spark-1.3 · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score80.0%

{"company":{"transcription":"Abegg & Cie."},"location":{"transcription":"Zürich"},"b_id":{"transcription":"B.51.322.GB.266"},"date":"","information":[]}

Scoring

Fuzzy score95.6%

Speed

Model time, all inputs345.9 s
Mean per input10.48 s
Slowest input17.97 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens115.5K
Output tokens43.5K
Total tokens159.1K
Input cost$0.1444
Output cost$0.185
Total cost$0.3294

Priced from 2026-09-08 · Input $/M $1.25 · Output $/M $4.25

Result 40 of 180

Test T1532 at 2026-09-17

index-card information-extraction typed, handwritten 20 de company

huggingface MiniMaxAI/MiniMax-M3 · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score80.0%

{"company":{"transcription":"Abegg & Cie."},"location":{"transcription":"Zürich"},"b_id":{"transcription":"B.51.322.GB.266"},"date":"","information":[{"transcription":""}]}

Scoring

Fuzzy score87.9%

Speed

Model time, all inputs338.4 s
Mean per input10.25 s
Slowest input15.85 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens32.3K
Output tokens3.8K
Total tokens36.1K
Input cost$0.009
Output cost$0.0041
Total cost$0.0132

Priced from 2026-09-16 · Input $/M $0.28 · Output $/M $1.10