RISE Humanities Data Benchmark, 0.6.1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'blacklist_cards__true' with Search Hidden 'False' returned 180 results, showing page 1 of 18.
Result 1 of 180

Test T1877 at 2026-09-29

index-card information-extraction typed, handwritten 20 de company

anthropic claude-sonnet-5-5 · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score80.0%

{"company":{"transcription":"Abegg & Cie."},"location":{"transcription":"Zürich"},"b_id":{"transcription":"B.51.322.GB.266"},"date":"","information":[]}

Scoring

Fuzzy score94.6%

Speed

Model time, all inputs100.0 s
Mean per input3.03 s
Slowest input4.28 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens140.1K
Output tokens8.7K
Total tokens148.8K
Input cost$0.2803
Output cost$0.0872
Total cost$0.3674

Priced from 2026-09-28 · Input $/M $2.00 · Output $/M $10.00

Result 2 of 180

Test T1817 at 2026-09-29

index-card information-extraction typed, handwritten 20 de company

anthropic claude-opus-5-5 · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score80.0%

{"company":{"transcription":"Abegg & Cie."},"location":{"transcription":"Zürich"},"b_id":{"transcription":"B.51.322.GB.266"},"date":null,"information":[]}

Scoring

Fuzzy score90.2%

Speed

Model time, all inputs150.4 s
Mean per input4.56 s
Slowest input6.44 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens140.1K
Output tokens9.1K
Total tokens149.2K
Input cost$0.5605
Output cost$0.1815
Total cost$0.742

Priced from 2026-09-27 · Input $/M $4.00 · Output $/M $20.00

Result 3 of 180

Test T1802 at 2026-09-29

index-card information-extraction typed, handwritten 20 de company

scicore Qwen3.8-Flash-Next-FP8 · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score80.0%

{"company":{"transcription":"Abegg & Cie."},"location":{"transcription":"Zürich"},"b_id":{"transcription":"B.51.322.GB.266"},"date":"","information":[{"transcription":""}]}

Scoring

Fuzzy score93.8%

Speed

Model time, all inputs233.0 s
Mean per input7.06 s
Slowest input31.55 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens86.6K
Output tokens25.3K
Total tokens112K
Input cost$0.00
Output cost$0.00
Total cost$0.00

Priced from 2026-09-27 · Input $/M $0.00 · Output $/M $0.00

Result 4 of 180

Test T1682 at 2026-09-29

index-card information-extraction typed, handwritten 20 de company

anthropic claude-fable-5-1 · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score80.0%

{"company":{"transcription":"Abegg & Cie."},"location":{"transcription":"Zürich"},"b_id":{"transcription":"B.51.322.GB.266"},"date":"","information":[]}

Scoring

Fuzzy score94.6%

Speed

Model time, all inputs220.6 s
Mean per input6.69 s
Slowest input14.31 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens140.1K
Output tokens8.8K
Total tokens149K
Input cost$1.40
Output cost$0.4416
Total cost$1.84

Priced from 2026-09-08 · Input $/M $10.00 · Output $/M $50.00

Result 5 of 180

Test T1877 at 2026-09-28

index-card information-extraction typed, handwritten 20 de company

anthropic claude-sonnet-5-5 · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score80.0%

{
  "company": {"transcription": "Abegg & Cie."},
  "location": {"transcription": "Zürich"},
  "b_id": {"transcription": "B.51.322.GB.266"},
  "date": "",
  "information": []
}

Scoring

Fuzzy score94.9%

Speed

Model time, all inputs87.1 s
Mean per input2.64 s
Slowest input3.36 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens113.2K
Output tokens5.7K
Total tokens118.9K
Input cost$0.2264
Output cost$0.0569
Total cost$0.2833

Priced from 2026-09-28 · Input $/M $2.00 · Output $/M $10.00

Result 6 of 180

Test T1847 at 2026-09-27

index-card information-extraction typed, handwritten 20 de company

openai gpt-6-luna · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score80.0%

{"company":{"transcription":"Abegg & Cie."},"location":{"transcription":"Zürich"},"b_id":{"transcription":"B.51.322.GB.266"},"date":"","information":[]}

Scoring

Fuzzy score89.2%

Speed

Model time, all inputs155.1 s
Mean per input4.70 s
Slowest input8.84 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens108.8K
Output tokens14.6K
Total tokens123.3K
Input cost$0.0109
Output cost$0.0073
Total cost$0.0182

Priced from 2026-09-27 · Input $/M $0.10 · Output $/M $0.50

Result 7 of 180

Test T1862 at 2026-09-27

index-card information-extraction typed, handwritten 20 de company

x-ai grok-4.7 · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score80.0%

{"company":{"transcription":"Abegg & Cie."},"location":{"transcription":"Zürich"},"b_id":{"transcription":"B.51.322.GB.266"},"date":"","information":[]}

Scoring

Fuzzy score93.7%

Speed

Model time, all inputs1789.7 s
Mean per input54.23 s
Slowest input143.41 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens139.1K
Output tokens3.7K
Total tokens142.8K
Input cost$0.2782
Output cost$0.0222
Total cost$1.04
Reasoning cost usd$0.7389
Total reasoning tokens123.2K

Priced from 2026-09-27 · Input $/M $2.00 · Output $/M $6.00

Result 8 of 180

Test T1817 at 2026-09-27

index-card information-extraction typed, handwritten 20 de company

anthropic claude-opus-5-5 · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score80.0%

```json
{
  "company": {"transcription": "Abegg & Cie."},
  "location": {"transcription": "Zürich"},
  "b_id": {"transcription": "B.51.322.GB.266"},
  "date": "",
  "information": [
    {"transcription": ""}
  ]
}
```

Scoring

Fuzzy score94.6%

Speed

Model time, all inputs195.9 s
Mean per input5.94 s
Slowest input8.45 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens113.2K
Output tokens13.5K
Total tokens126.7K
Input cost$0.4528
Output cost$0.2707
Total cost$0.7235

Priced from 2026-09-27 · Input $/M $4.00 · Output $/M $20.00

Result 9 of 180

Test T1832 at 2026-09-27

index-card information-extraction typed, handwritten 20 de company

openai gpt-6-sol · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score80.0%

{"company":{"transcription":"Abegg & Cie."},"location":{"transcription":"Zürich"},"b_id":{"transcription":"B.51.322.GB.266"},"date":"","information":[]}

Scoring

Fuzzy score94.6%

Speed

Model time, all inputs168.1 s
Mean per input5.10 s
Slowest input7.83 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens108.8K
Output tokens10.1K
Total tokens118.9K
Input cost$0.2175
Output cost$0.1013
Total cost$0.3188

Priced from 2026-09-27 · Input $/M $2.00 · Output $/M $10.00

Result 10 of 180

Test T0990 at 2026-09-18

index-card information-extraction typed, handwritten 20 de company

openrouter qwen/qwen3.5-9b · temp 0.5 · dataclass Card

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score79.2%

{"company":{"transcription":"A begg & Cie."},"location":{"transcription":"Zürich"},"b_id":{"transcription":"B.51.322.GB.266"},"date":"","information":[]}

Scoring

Fuzzy score90.0%

Speed

Model time, all inputs2377.8 s
Mean per input72.05 s
Slowest input334.68 s
Inputs timed33

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens85.4K
Output tokens70.9K
Total tokens156.3K
Input cost$0.0085
Output cost$0.0106
Total cost$0.0192

Priced from 2026-09-16 · Input $/M $0.10 · Output $/M $0.15