RISE Humanities Data Benchmark, 0.5.0-pre1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Search Results

Your search for Benchmark 'magazine_pages__true' with Search Hidden 'False' returned 52 results, showing page 1 of 6.
Result 1 of 52

Test T0878 at 2026-03-25

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Provideralibaba
Modelqwen3.5-397b-a17b
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score0.00 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 127.7K IT + 3.1K OT = 130.8K TTCost: 0.077$0.011$0.088$
Result 2 of 52

Test T0839 at 2026-03-25

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Provideralibaba
Modelqwen3.5-35b-a3b
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score0.00 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 127.7K IT + 5.2K OT = 132.9K TTCost: 0.032$0.010$0.042$
Result 3 of 52

Test T0865 at 2026-03-25

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Provideralibaba
Modelqwen3.5-122b-a10b
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score0.00 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 127.7K IT + 5.0K OT = 132.7K TTCost: 0.051$0.016$0.067$
Result 4 of 52

Test T0852 at 2026-03-25

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Provideralibaba
Modelqwen3.5-27b
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score0.00 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 127.7K IT + 5.7K OT = 133.4K TTCost: 0.038$0.014$0.052$
Result 5 of 52

Test T0891 at 2026-03-25

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Provideralibaba
Modelqwen3.5-flash-2026-02-23
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score0.00 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 72.2K IT + 2.4K OT = 74.5K TTCost: 0.007$0.001$0.008$
Result 6 of 52

Test T0826 at 2026-03-24

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Provideralibaba
Modelqwen3.5-plus
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score0.00 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 127.7K IT + 2.2K OT = 129.9K TTCost: 0.051$0.005$0.056$
Result 7 of 52

Test T0783 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Providermistral
Modelmistral-large-2411
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score0.00 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 12.3K IT + 25.7K OT = 37.9K TTCost: 0.025$0.154$0.178$
Result 8 of 52

Test T0778 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Provideranthropic
Modelclaude-opus-4-1-20250805
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score0.00 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 98.8K IT + 4.4K OT = 103.2K TTCost: 1.482$0.328$1.810$
Result 9 of 52

Test T0813 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Providerx-ai
Modelgrok-4.20-0309-reasoning
  
Temperature0.0
DataclassMagazinePage
  
Normalized Score49.10 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 24.4K IT + 2.7K OT = 27.1K TTCost: 0.049$0.016$0.065$
Result 10 of 52

Test T0803 at 2026-03-23

{'century': [20], 'document-type': ['newspaper-page'], 'language': ['en'], 'layout': ['prose', 'columns'], 'script': ['latin'], 'task': ['document-understanding'], 'writing': ['printed']}

Configuration
Provideropenai
Modelgpt-5.2-2025-12-11
  
Temperature1.0
DataclassMagazinePage
  
Normalized Score88.50 %
Test timeunknown seconds
Prompt

Extract all advertisements and return their bounding boxes.
The original size of the page is {width} x {height} pixels.

Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
n/a n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 90.4K IT + 2.3K OT = 92.7K TTCost: 0.158$0.032$0.190$