RISE Humanities Data Benchmark, 0.5.5

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'blacklist_cards__true' with Search Hidden 'False' returned 113 results, showing page 1 of 12.
Result 1 of 113

Test T1742 at 2026-09-09

index-card information-extraction typed, handwritten 20 de company

openrouter qwen/qwen3.8-27b · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score77.1%

{"company":{"transcription":"A b e g g & Cie."},"location":{"transcription":"Zürich"},"b_id":{"transcription":"B.51.322.GB.266"},"date":"","information":[]}

Scoring

Fuzzy score94.3%

Speed

Model time, all inputs6752.6 s
Mean per input204.62 s
Slowest input680.57 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens86.7K
Output tokens20.5K
Total tokens107.2K
Input cost$0.0364
Output cost$0.0614
Total cost$0.0978

Priced from 2026-09-08 · Input $/M $0.42 · Output $/M $3.00

Result 2 of 113

Test T1727 at 2026-09-08

index-card information-extraction typed, handwritten 20 de company

openrouter qwen/qwen3.8-flash · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score80.0%

{"company":{"transcription":"Abegg & Cie."},"location":{"transcription":"Zürich"},"b_id":{"transcription":"B.51.322.GB.266"},"date":"","information":[]}

Scoring

Fuzzy score92.5%

Speed

Model time, all inputs1505.5 s
Mean per input45.62 s
Slowest input340.75 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens96.4K
Output tokens65.7K
Total tokens162.1K
Input cost$0.0118
Output cost$0.0296
Total cost$0.0452

Priced from 2026-09-08 · Input $/M $0.15 · Output $/M $0.47

Result 3 of 113

Test T1697 at 2026-09-08

index-card information-extraction typed, handwritten 20 de company

openrouter meta/muse-spark-1.3 · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score94.3%

{"company":{"transcription":"Abegg & Cie."},"location":{"transcription":"Z\u0000fcrich"},"b_id":{"transcription":"B.51.322.GB.266"},"date":"","information":null}

Scoring

Fuzzy score93.9%

Speed

Model time, all inputs390.8 s
Mean per input11.84 s
Slowest input19.51 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens115.4K
Output tokens42.4K
Total tokens157.8K
Input cost$0.1443
Output cost$0.1801
Total cost$0.3243

Priced from 2026-09-08 · Input $/M $1.25 · Output $/M $4.25

Result 4 of 113

Test T1757 at 2026-09-08

index-card information-extraction typed, handwritten 20 de company

deepseek deepseek-v4-flash-vision-exp · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score79.3%

{"company":{"transcription":"Abegg & Cie."},"location":{"transcription":"Zürich"},"b_id":{"transcription":"B.51.322.GB.26"},"date":"","information":[]}

Scoring

Fuzzy score93.0%

Speed

Model time, all inputs512.7 s
Mean per input15.54 s
Slowest input29.75 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens35.1K
Output tokens50.8K
Total tokens85.8K
Input cost$0.0154
Output cost$0.067
Total cost$0.0824

Priced from 2026-09-08 · Input $/M $0.44 · Output $/M $1.32

Result 5 of 113

Test T1712 at 2026-09-08

index-card information-extraction typed, handwritten 20 de company

openrouter z-ai/glm-5.3-flash · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score80.0%

{"company":{"transcription":"Abegg & Cie."},"location":{"transcription":"Zürich"},"b_id":{"transcription":"B.51.322.GB.266"},"date":"","information":[{"transcription":""}]}

Scoring

Fuzzy score92.5%

Speed

Model time, all inputs316.7 s
Mean per input9.60 s
Slowest input110.09 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens109.9K
Output tokens4K
Total tokens113.9K
Input cost$0.0082
Output cost$0.001
Total cost$0.0092

Priced from 2026-09-08 · Input $/M $0.075 · Output $/M $0.25

Result 6 of 113

Test T1667 at 2026-09-04

index-card information-extraction typed, handwritten 20 de company

openai gpt-6-astra · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score80.0%

{"company":{"transcription":"Abegg & Cie."},"location":{"transcription":"Zürich"},"b_id":{"transcription":"B.51.322.GB.266"},"date":"","information":[{"transcription":""}]}

Scoring

Fuzzy score95.6%

Speed

Model time, all inputs154.9 s
Mean per input4.69 s
Slowest input9.66 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens108.8K
Output tokens6.1K
Total tokens114.8K
Input cost$1.09
Output cost$0.3026
Total cost$1.39

Priced from 2026-09-03 · Input $/M $10.00 · Output $/M $50.00

Result 7 of 113

Test T1652 at 2026-09-03

index-card information-extraction typed, handwritten 20 de company

genai gemini-3.8-flash · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score77.1%

{"company":{"transcription":"A b e g g & Cie."},"location":{"transcription":"Zürich"},"b_id":{"transcription":"B.51.322.GB.266"},"date":"","information":[]}

Scoring

Fuzzy score93.1%

Speed

Model time, all inputs209.5 s
Mean per input6.35 s
Slowest input28.72 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens45K
Output tokens5.5K
Total tokens50.4K
Input cost$0.0337
Output cost$0.0205
Total cost$0.0543

Priced from 2026-09-02 · Input $/M $0.75 · Output $/M $3.75

Result 8 of 113

Test T1592 at 2026-08-18

index-card information-extraction typed, handwritten 20 de company

genai gemini-3.7-flash · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score77.1%

{"company":{"transcription":"A b e g g & Cie."},"location":{"transcription":"Zürich"},"b_id":{"transcription":"B.51.322.GB.266"},"date":"","information":[]}

Scoring

Fuzzy score94.4%

Speed

Model time, all inputs139.0 s
Mean per input4.21 s
Slowest input10.47 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens45K
Output tokens5.5K
Total tokens50.5K
Input cost$0.0337
Output cost$0.0206
Total cost$0.0544

Priced from 2026-08-18 · Input $/M $0.75 · Output $/M $3.75

Result 9 of 113

Test T1607 at 2026-08-18

index-card information-extraction typed, handwritten 20 de company

x-ai grok-4.6 · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score76.5%

{"company":{"transcription":"A b e g g & Cie."},"location":{"transcription":"Zürich"},"b_id":{"transcription":"B.51.322.GB.26"},"date":"","information":[]}

Scoring

Fuzzy score95.8%

Speed

Model time, all inputs589.6 s
Mean per input17.87 s
Slowest input39.11 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens103.7K
Output tokens3.7K
Total tokens107.5K
Input cost$0.2075
Output cost$0.0224
Total cost$0.2298

Priced from 2026-08-18 · Input $/M $2.00 · Output $/M $6.00

Result 10 of 113

Test T1637 at 2026-08-18

index-card information-extraction typed, handwritten 20 de company

openrouter z-ai/glm-5v-turbo · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score79.2%

{"company":{"transcription":"A begg & Cie."},"location":{"transcription":"Zürich"},"b_id":{"transcription":"B.51.322.GB.266"},"date":"","information":[]}

Scoring

Fuzzy score90.6%

Speed

Model time, all inputs880.7 s
Mean per input26.69 s
Slowest input41.24 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens121.5K
Output tokens24.3K
Total tokens145.9K
Input cost$0.1458
Output cost$0.0974
Total cost$0.2432

Priced from 2026-08-18 · Input $/M $1.20 · Output $/M $4.00