RISE Humanities Data Benchmark, 0.5.2-pre1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Search Results

Your search for Benchmark 'magazine_pages__true' with Search Hidden 'False' returned 80 results, showing page 6 of 8.
Result 51 of 80

Test T0777 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Provideropenai
Modelgpt-5-nano-2025-08-07
  
Temperature1.0
DataclassMagazinePage
  
Normalized Score50.20 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 110.7K IT + 82.6K OT = 193.3K TTCost: 0.006$0.033$0.039$
Result 52 of 80

Test T0783 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Providermistral
Modelmistral-large-2411
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score0.00 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 12.3K IT + 25.7K OT = 37.9K TTCost: 0.025$0.154$0.178$
Result 53 of 80

Test T0808 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Providermistral
Modelministral-14b-2512
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score0.00 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 11.7K IT + 506 OT = 12.2K TTCost: 0.002$0.000$0.002$
Result 54 of 80

Test T0812 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Providergenai
Modelgemini-3.1-flash-lite-preview
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score0.00 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 52.3K IT + 5.1K OT = 57.4K TTCost: 0.013$0.008$0.021$
Result 55 of 80

Test T0769 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Provideropenai
Modelgpt-4o-mini-2024-07-18
  
Temperature1.0
DataclassMagazinePage
  
Normalized Score7.60 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 1.7M IT + 1.5K OT = 1.7M TTCost: 0.256$0.001$0.256$
Result 56 of 80

Test T0806 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Provideranthropic
Modelclaude-haiku-4-5-20251001
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score0.00 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 111.4K IT + 4.3K OT = 115.7K TTCost: 0.111$0.022$0.133$
Result 57 of 80

Test T0781 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Providermistral
Modelmistral-medium-2508
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score0.30 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 11.7K IT + 13.8K OT = 25.5K TTCost: 0.005$0.028$0.032$
Result 58 of 80

Test T0775 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Provideranthropic
Modelclaude-opus-4-20250514
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score2.20 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 98.8K IT + 4.5K OT = 103.3K TTCost: 1.482$0.335$1.817$
Result 59 of 80

Test T0786 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Providergenai
Modelgemini-2.5-flash-lite-preview-09-2025
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score0.00 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 13.9K IT + 5.4K OT = 19.4K TTCost: 0.001$0.002$0.004$
Result 60 of 80

Test T0785 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Providergenai
Modelgemini-2.5-flash-lite
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score0.00 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 13.9K IT + 397.9K OT = 411.8K TTCost: 0.001$0.159$0.161$