RISE Humanities Data Benchmark, 0.5.0-pre1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Search Results

Your search for Benchmark 'magazine_pages__true' with Search Hidden 'False' returned 52 results, showing page 2 of 6.
Result 11 of 52

Test T0796 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Providermistral
Modelmistral-small-2506
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score6.60 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 11.7K IT + 8.7K OT = 20.4K TTCost: 0.001$0.003$0.004$
Result 12 of 52

Test T0786 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Providergenai
Modelgemini-2.5-flash-lite-preview-09-2025
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score0.00 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 13.9K IT + 5.4K OT = 19.4K TTCost: 0.001$0.002$0.004$
Result 13 of 52

Test T0805 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Provideranthropic
Modelclaude-opus-4-5-20251101
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score5.40 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 111.4K IT + 3.7K OT = 115.0K TTCost: 0.557$0.092$0.649$
Result 14 of 52

Test T0803 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Provideropenai
Modelgpt-5.2-2025-12-11
  
Temperature1.0
DataclassMagazinePage
  
Normalized Score88.50 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 90.4K IT + 2.3K OT = 92.7K TTCost: 0.158$0.032$0.190$
Result 15 of 52

Test T0783 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Providermistral
Modelmistral-large-2411
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score0.00 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 12.3K IT + 25.7K OT = 37.9K TTCost: 0.025$0.154$0.178$
Result 16 of 52

Test T0787 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Provideranthropic
Modelclaude-sonnet-4-5-20250929
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score2.20 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 111.4K IT + 4.4K OT = 115.8K TTCost: 0.334$0.066$0.400$
Result 17 of 52

Test T0788 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Provideropenrouter
Modelqwen/qwen3-vl-8b-thinking
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score0.00 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 128.5K IT + 4.8K OT = 133.3K TTCost: 0.015$0.006$0.022$
Result 18 of 52

Test T0811 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Provideranthropic
Modelclaude-opus-4-6
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score21.50 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 111.4K IT + 3.8K OT = 115.2K TTCost: 0.557$0.096$0.653$
Result 19 of 52

Test T0806 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Provideranthropic
Modelclaude-haiku-4-5-20251001
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score0.00 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 111.4K IT + 4.3K OT = 115.7K TTCost: 0.111$0.022$0.133$
Result 20 of 52

Test T0808 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Providermistral
Modelministral-14b-2512
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score0.00 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 11.7K IT + 506 OT = 12.2K TTCost: 0.002$0.000$0.002$