RISE Humanities Data Benchmark, 0.6.1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'book_advert_xml__true' with Search Hidden 'False' returned 187 results, showing page 1 of 19.
Result 1 of 187

Test T1805 at 2026-09-29

newspaper-page data-correction 18 en

scicore Qwen3.8-Flash-Next-FP8 · temp 0.0 · dataclass CorrectedAdvert

Result 2 of 187

Test T1880 at 2026-09-29

newspaper-page data-correction 18 en

anthropic claude-sonnet-5-5 · temp 0.0 · dataclass CorrectedAdvert

Result 3 of 187

Test T1685 at 2026-09-29

newspaper-page data-correction 18 en

anthropic claude-fable-5-1 · temp 0.0 · dataclass CorrectedAdvert

Result 4 of 187

Test T1820 at 2026-09-29

newspaper-page data-correction 18 en

anthropic claude-opus-5-5 · temp 0.0 · dataclass CorrectedAdvert

Result 5 of 187

Test T1880 at 2026-09-28

newspaper-page data-correction 18 en

anthropic claude-sonnet-5-5 · temp 0.0 · dataclass CorrectedAdvert

Result 6 of 187

Test T1820 at 2026-09-27

newspaper-page data-correction 18 en

anthropic claude-opus-5-5 · temp 0.0 · dataclass CorrectedAdvert

Result 7 of 187

Test T1850 at 2026-09-27

newspaper-page data-correction 18 en

openai gpt-6-luna · temp 1.0 · dataclass CorrectedAdvert

Result 8 of 187

Test T1865 at 2026-09-27

newspaper-page data-correction 18 en

x-ai grok-4.7 · temp 0.0 · dataclass CorrectedAdvert

Result 9 of 187

Test T1835 at 2026-09-27

newspaper-page data-correction 18 en

openai gpt-6-sol · temp 1.0 · dataclass CorrectedAdvert

Result 10 of 187

Test T0478 at 2026-09-18

newspaper-page data-correction 18 en

openrouter meta-llama/llama-4-maverick · temp 0.0 · dataclass CorrectedAdvert