RISE Humanities Data Benchmark, 0.5.0-pre1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Search Results

Your search for Benchmark 'magazine_pages__true' with Search Hidden 'False' returned 52 results, showing page 3 of 6.
Result 21 of 52

Test T0771 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Providermistral
Modelpixtral-large-2411
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score7.90 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 12.3K IT + 4.8K OT = 17.1K TTCost: 0.025$0.029$0.053$
Result 22 of 52

Test T0786 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Providergenai
Modelgemini-2.5-flash-lite-preview-09-2025
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score0.00 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 13.9K IT + 5.4K OT = 19.4K TTCost: 0.001$0.002$0.004$
Result 23 of 52

Test T0811 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Provideranthropic
Modelclaude-opus-4-6
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score21.50 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 111.4K IT + 3.8K OT = 115.2K TTCost: 0.557$0.096$0.653$
Result 24 of 52

Test T0788 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Provideropenrouter
Modelqwen/qwen3-vl-8b-thinking
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score0.00 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 128.5K IT + 4.8K OT = 133.3K TTCost: 0.015$0.006$0.022$
Result 25 of 52

Test T0796 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Providermistral
Modelmistral-small-2506
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score6.60 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 11.7K IT + 8.7K OT = 20.4K TTCost: 0.001$0.003$0.004$
Result 26 of 52

Test T0774 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Provideropenai
Modelgpt-4.1-nano-2025-04-14
  
Temperature1.0
DataclassMagazinePage
  
Normalized Score2.20 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 175.9K IT + 1.4K OT = 177.2K TTCost: 0.018$0.001$0.018$
Result 27 of 52

Test T0804 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Providergenai
Modelgemini-3-flash-preview
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score84.80 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 52.3K IT + 3.2K OT = 55.5K TTCost: 0.026$0.010$0.036$
Result 28 of 52

Test T0792 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Provideropenai
Modelqwen/qwen3-vl-8b-instruct
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score0.00 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 131.2K IT + 4.4K OT = 135.6K TTCost: 0.000$0.000$0.004$
Result 29 of 52

Test T0791 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Provideropenai
Modelqwen/qwen3-vl-30b-a3b-instruct
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score0.00 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 131.2K IT + 39.4K OT = 170.6K TTCost: 0.000$0.000$0.021$
Result 30 of 52

Test T0780 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Provideropenai
Modelo3-2025-04-16
  
Temperature1.0
DataclassMagazinePage
  
Normalized Score75.40 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 54.0K IT + 31.9K OT = 85.9K TTCost: 0.108$0.255$0.363$