RISE Humanities Data Benchmark, 0.5.0-pre1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Search Results

Your search for Benchmark 'bibliographic_data__true' with Search Hidden 'False' returned 115 results, showing page 1 of 12.
Result 1 of 115

Test T0866 at 2026-03-25

{'document-type': ['book-page'], 'writing': ['printed'], 'century': [20], 'language': ['en'], 'layout': ['list'], 'entry-type': ['bibliographic'], 'task': ['information-extraction']}

Configuration
Provideralibaba
Modelqwen3.5-397b-a17b
  
Temperature0.0
DataclassDocument
  
Normalized Score67.87 %
Test timeunknown seconds
Prompt
Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
0.68 n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 17.5K IT + 10.1K OT = 27.6K TTCost: 0.011$0.036$0.047$
Result 2 of 115

Test T0840 at 2026-03-25

{'document-type': ['book-page'], 'writing': ['printed'], 'century': [20], 'language': ['en'], 'layout': ['list'], 'entry-type': ['bibliographic'], 'task': ['information-extraction']}

Configuration
Provideralibaba
Modelqwen3.5-27b
  
Temperature0.0
DataclassDocument
  
Normalized Score66.45 %
Test timeunknown seconds
Prompt
Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
0.66 n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 17.5K IT + 16.4K OT = 33.9K TTCost: 0.005$0.039$0.045$
Result 3 of 115

Test T0827 at 2026-03-25

{'document-type': ['book-page'], 'writing': ['printed'], 'century': [20], 'language': ['en'], 'layout': ['list'], 'entry-type': ['bibliographic'], 'task': ['information-extraction']}

Configuration
Provideralibaba
Modelqwen3.5-35b-a3b
  
Temperature0.0
DataclassDocument
  
Normalized Score64.87 %
Test timeunknown seconds
Prompt
Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
0.65 n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 17.5K IT + 9.7K OT = 27.2K TTCost: 0.004$0.019$0.024$
Result 4 of 115

Test T0879 at 2026-03-25

{'document-type': ['book-page'], 'writing': ['printed'], 'century': [20], 'language': ['en'], 'layout': ['list'], 'entry-type': ['bibliographic'], 'task': ['information-extraction']}

Configuration
Provideralibaba
Modelqwen3.5-flash-2026-02-23
  
Temperature0.0
DataclassDocument
  
Normalized Score67.34 %
Test timeunknown seconds
Prompt
Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
0.67 n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 17.5K IT + 10.4K OT = 27.9K TTCost: 0.002$0.004$0.006$
Result 5 of 115

Test T0853 at 2026-03-25

{'document-type': ['book-page'], 'writing': ['printed'], 'century': [20], 'language': ['en'], 'layout': ['list'], 'entry-type': ['bibliographic'], 'task': ['information-extraction']}

Configuration
Provideralibaba
Modelqwen3.5-122b-a10b
  
Temperature0.0
DataclassDocument
  
Normalized Score65.34 %
Test timeunknown seconds
Prompt
Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
0.65 n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 17.5K IT + 10.1K OT = 27.6K TTCost: 0.007$0.032$0.039$
Result 6 of 115

Test T0814 at 2026-03-24

{'document-type': ['book-page'], 'writing': ['printed'], 'century': [20], 'language': ['en'], 'layout': ['list'], 'entry-type': ['bibliographic'], 'task': ['information-extraction']}

Configuration
Provideralibaba
Modelqwen3.5-plus
  
Temperature0.0
DataclassDocument
  
Normalized Score69.18 %
Test timeunknown seconds
Prompt
Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
0.69 n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 17.5K IT + 10.3K OT = 27.8K TTCost: 0.007$0.025$0.032$
Result 7 of 115

Test T0718 at 2026-03-23

{'document-type': ['book-page'], 'writing': ['printed'], 'century': [20], 'language': ['en'], 'layout': ['list'], 'entry-type': ['bibliographic'], 'task': ['information-extraction']}

Configuration
Providerx-ai
Modelgrok-4.20-0309-reasoning
  
Temperature0.0
DataclassDocument
  
Normalized Score44.60 %
Test timeunknown seconds
Prompt
Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
0.45 n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 4.9K IT + 14.6K OT = 19.5K TTCost: 0.010$0.088$0.098$
Result 8 of 115

Test T0693 at 2026-03-23

{'document-type': ['book-page'], 'writing': ['printed'], 'century': [20], 'language': ['en'], 'layout': ['list'], 'entry-type': ['bibliographic'], 'task': ['information-extraction']}

Configuration
Provideropenai
Modelgpt-5.3-codex
  
Temperature0.0
DataclassDocument
  
Normalized Score68.17 %
Test timeunknown seconds
Prompt
Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
0.68 n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 13.8K IT + 7.1K OT = 20.8K TTCost: 0.024$0.099$0.123$
Result 9 of 115

Test T0488 at 2026-03-23

{'document-type': ['book-page'], 'writing': ['printed'], 'century': [20], 'language': ['en'], 'layout': ['list'], 'entry-type': ['bibliographic'], 'task': ['information-extraction']}

Configuration
Provideropenai
Modelgpt-5.2-2025-12-11
  
Temperature0.0
DataclassDocument
  
Normalized Score66.07 %
Test timeunknown seconds
Prompt
Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
0.66 n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 12.3K IT + 10.3K OT = 22.6K TTCost: 0.022$0.144$0.165$
Result 10 of 115

Test T0693 at 2026-03-17

{'document-type': ['book-page'], 'writing': ['printed'], 'century': [20], 'language': ['en'], 'layout': ['list'], 'entry-type': ['bibliographic'], 'task': ['information-extraction']}

Configuration
Provideropenai
Modelgpt-5.3-codex
  
Temperature0.0
DataclassDocument
  
Normalized Score67.49 %
Test timeunknown seconds
Prompt
Results

no valid result

Scoring
Fuzzy Score F1 micro / macro Micro precision/recall Tue/False Positives
0.67 n/a n/a n/a n/a n/a n/a n/a n/a
      Micro Precision Micro Recall Instances TP FP FN
Costs / Pricing
Pricing Date: n/an/aTokens: 13.8K IT + 5.9K OT = 19.6K TTCost: 0.024$0.082$0.106$