RISE Humanities Data Benchmark, 0.5.0-pre1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Search Results

Your search for Benchmark 'general_meeting_minutes__true' with Search Hidden 'False' returned 20 results, showing page 1 of 2.
Result 1 of 20

Test T0708 at 2026-03-23

{'document-type': ['minutes'], 'writing': ['typed', 'handwritten'], 'century': [20], 'language': ['it', 'fr', 'de'], 'layout': ['table'], 'entry-type': ['person', 'location'], 'task': ['information-extraction']}

Result 2 of 20

Test T0706 at 2026-03-23

{'document-type': ['minutes'], 'writing': ['typed', 'handwritten'], 'century': [20], 'language': ['it', 'fr', 'de'], 'layout': ['table'], 'entry-type': ['person', 'location'], 'task': ['information-extraction']}

Result 3 of 20

Test T0705 at 2026-03-23

{'document-type': ['minutes'], 'writing': ['typed', 'handwritten'], 'century': [20], 'language': ['it', 'fr', 'de'], 'layout': ['table'], 'entry-type': ['person', 'location'], 'task': ['information-extraction']}

Result 4 of 20

Test T0709 at 2026-03-23

{'document-type': ['minutes'], 'writing': ['typed', 'handwritten'], 'century': [20], 'language': ['it', 'fr', 'de'], 'layout': ['table'], 'entry-type': ['person', 'location'], 'task': ['information-extraction']}

Result 5 of 20

Test T0707 at 2026-03-23

{'document-type': ['minutes'], 'writing': ['typed', 'handwritten'], 'century': [20], 'language': ['it', 'fr', 'de'], 'layout': ['table'], 'entry-type': ['person', 'location'], 'task': ['information-extraction']}

Result 6 of 20

Test T0706 at 2026-03-17

{'document-type': ['minutes'], 'writing': ['typed', 'handwritten'], 'century': [20], 'language': ['it', 'fr', 'de'], 'layout': ['table'], 'entry-type': ['person', 'location'], 'task': ['information-extraction']}

Result 7 of 20

Test T0708 at 2026-03-17

{'document-type': ['minutes'], 'writing': ['typed', 'handwritten'], 'century': [20], 'language': ['it', 'fr', 'de'], 'layout': ['table'], 'entry-type': ['person', 'location'], 'task': ['information-extraction']}

Result 8 of 20

Test T0707 at 2026-03-17

{'document-type': ['minutes'], 'writing': ['typed', 'handwritten'], 'century': [20], 'language': ['it', 'fr', 'de'], 'layout': ['table'], 'entry-type': ['person', 'location'], 'task': ['information-extraction']}

Result 9 of 20

Test T0705 at 2026-03-17

{'document-type': ['minutes'], 'writing': ['typed', 'handwritten'], 'century': [20], 'language': ['it', 'fr', 'de'], 'layout': ['table'], 'entry-type': ['person', 'location'], 'task': ['information-extraction']}

Result 10 of 20

Test T0709 at 2026-03-17

{'document-type': ['minutes'], 'writing': ['typed', 'handwritten'], 'century': [20], 'language': ['it', 'fr', 'de'], 'layout': ['table'], 'entry-type': ['person', 'location'], 'task': ['information-extraction']}