RISE Humanities Data Benchmark, 0.5.2-pre1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Search Results

Your search for Benchmark 'magazine_pages__true' with Search Hidden 'False' returned 80 results, showing page 7 of 8.
Result 61 of 80

Test T0773 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Provideropenai
Modelgpt-4.1-mini-2025-04-14
  
Temperature1.0
DataclassMagazinePage
  
Normalized Score19.90 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 119.0K IT + 1.8K OT = 120.7K TTCost: 0.048$0.003$0.050$
Result 62 of 80

Test T0805 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Provideranthropic
Modelclaude-opus-4-5-20251101
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score5.40 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 111.4K IT + 3.7K OT = 115.0K TTCost: 0.557$0.092$0.649$
Result 63 of 80

Test T0788 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Provideropenrouter
Modelqwen/qwen3-vl-8b-thinking
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score0.00 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 128.5K IT + 4.8K OT = 133.3K TTCost: 0.015$0.006$0.022$
Result 64 of 80

Test T0796 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Providermistral
Modelmistral-small-2506
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score6.60 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 11.7K IT + 8.7K OT = 20.4K TTCost: 0.001$0.003$0.004$
Result 65 of 80

Test T0811 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Provideranthropic
Modelclaude-opus-4-6
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score21.50 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 111.4K IT + 3.8K OT = 115.2K TTCost: 0.557$0.096$0.653$
Result 66 of 80

Test T0792 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Provideropenai
Modelqwen/qwen3-vl-8b-instruct
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score0.00 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 131.2K IT + 4.4K OT = 135.6K TTCost: 0.000$0.000$0.004$
Result 67 of 80

Test T0791 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Provideropenai
Modelqwen/qwen3-vl-30b-a3b-instruct
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score0.00 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 131.2K IT + 39.4K OT = 170.6K TTCost: 0.000$0.000$0.021$
Result 68 of 80

Test T0778 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Provideranthropic
Modelclaude-opus-4-1-20250805
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score0.00 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 98.8K IT + 4.4K OT = 103.2K TTCost: 1.482$0.328$1.810$
Result 69 of 80

Test T0804 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Providergenai
Modelgemini-3-flash-preview
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score84.80 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 52.3K IT + 3.2K OT = 55.5K TTCost: 0.026$0.010$0.036$
Result 70 of 80

Test T0776 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Provideranthropic
Modelclaude-sonnet-4-20250514
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score0.00 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 98.8K IT + 4.1K OT = 102.9K TTCost: 0.296$0.062$0.358$