RISE Humanities Data Benchmark, 0.6.0-pre1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'book_advert_xml__true' with Search Hidden 'False' returned 179 results, showing page 5 of 18.
Result 41 of 179

Test T0509 at 2026-09-16

newspaper-page data-correction 18 en

genai gemini-3-flash-preview · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Fixed the malformed opening tag <TITLE by adding the missing closing angle bracket and separating it from the text content."}

Scoring

Fuzzy score95.6%
Items46

Speed

Model time, all inputs4412.9 s
Mean per input88.26 s
Slowest input142.10 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens31.5K
Output tokens27.1K
Total tokens58.6K
Input cost$0.0158
Output cost$0.0812
Total cost$3.75
Reasoning cost usd$3.66
Total reasoning tokens1.2M

Priced from 2026-09-16 · Input $/M $0.50 · Output $/M $3.00

Result 42 of 179

Test T0728 at 2026-09-16

newspaper-page data-correction 18 en

x-ai grok-4.20-0309-reasoning · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Added missing '>' character after the opening '<TITLE' tag to make the XML well-formed."}

Scoring

Fuzzy score97.7%
Items50

Speed

Model time, all inputs2344.1 s
Mean per input46.88 s
Slowest input85.21 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens43.5K
Output tokens27K
Total tokens70.5K
Input cost$0.0543
Output cost$0.0676
Total cost$0.723
Reasoning cost usd$0.601
Total reasoning tokens240.4K

Priced from 2026-09-16 · Input $/M $1.25 · Output $/M $2.50

Result 43 of 179

Test T0662 at 2026-09-16

newspaper-page data-correction 18 en

openai gpt-5.4-2026-03-05 · temp 1.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Inserted the missing opening angle bracket in the TITLE start tag, changing '<TITLEEine Tossanische Bibel</TITLE>' to '<TITLE>Eine Tossanische Bibel</TITLE>'."}

Scoring

Fuzzy score97.0%
Items50

Speed

Model time, all inputs299.4 s
Mean per input5.99 s
Slowest input10.68 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens33.6K
Output tokens26.5K
Total tokens60.1K
Input cost$0.0841
Output cost$0.3968
Total cost$0.4809

Priced from 2026-09-16 · Input $/M $2.50 · Output $/M $15.00

Result 44 of 179

Test T0485 at 2026-09-16

newspaper-page data-correction 18 en

cohere command-r7b-12-2024 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score18.9%

{
  "fixed_xml": "<advertisement><book title='The Great Adventure' author='John Doe' year='2023' genre='Fantasy'/></advertisement>",
  "number_of_fixes": 2,
  "explanation": "Added missing closing tags for the book element and added attributes for title, author, year, and genre."
}

Scoring

Fuzzy score18.1%
Items50

Speed

Model time, all inputs70.5 s
Mean per input1.41 s
Slowest input3.22 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens38.4K
Output tokens5.5K
Total tokens43.9K
Input cost$0.0014
Output cost$0.0008
Total cost$0.0023

Priced from 2026-09-16 · Input $/M $0.0375 · Output $/M $0.15

Result 45 of 179

Test T1655 at 2026-09-16

newspaper-page data-correction 18 en

genai gemini-3.8-flash · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Fixed malformed opening tag <TITLEEine to <TITLE>Eine."}

Scoring

Fuzzy score95.9%
Items50

Speed

Model time, all inputs1420.4 s
Mean per input28.41 s
Slowest input102.61 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens31.5K
Output tokens29.2K
Total tokens60.8K
Input cost$0.0237
Output cost$0.1096
Total cost$1.67
Reasoning cost usd$1.53
Total reasoning tokens409K

Priced from 2026-09-02 · Input $/M $0.75 · Output $/M $3.75

Result 46 of 179

Test T1036 at 2026-09-16

newspaper-page data-correction 18 en

deepseek deepseek-v4-pro · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"The opening tag <TITLEEine was missing the > delimiter, so it was corrected to <TITLE>."}

Scoring

Fuzzy score98.1%
Items49

Speed

Model time, all inputs6674.6 s
Mean per input133.49 s
Slowest input302.73 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens43.4K
Output tokens477.9K
Total tokens521.2K
Input cost$0.0572
Output cost$1.89
Total cost$1.95

Priced from 2026-09-16 · Input $/M $1.32 · Output $/M $3.96

Result 47 of 179

Test T1370 at 2026-09-16

newspaper-page data-correction 18 en

openai gpt-5.6-terra · temp 1.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Added the missing closing angle bracket to the opening TITLE tag."}

Scoring

Fuzzy score98.2%
Items50

Speed

Model time, all inputs487.1 s
Mean per input9.74 s
Slowest input28.94 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens33.6K
Output tokens37.4K
Total tokens71.1K
Input cost$0.0672
Output cost$0.4493
Total cost$0.5165

Priced from 2026-09-16 · Input $/M $2.00 · Output $/M $12.00

Result 48 of 179

Test T0486 at 2026-09-16

newspaper-page data-correction 18 en

cohere command-a-03-2025 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score15.8%

{"fixed_xml":"<advertisement><book><title>The Great Adventure</title><author>John Doe</author><price>19.99</price><description>An exciting journey through unknown lands.</description></book><book><title>Mystery of the Lost City</title><author>Jane Smith</author><price>24.95</price><description>A thrilling mystery set in an ancient city.</description></book></advertisement>","number_of_corrections":4,"explanation":"Corrected the root element to <advertisement>, added missing closing tags for <book>, and ensured all elements are properly nested and closed. Also, added a missing <description> tag for the second book."}

Scoring

Fuzzy score22.4%
Items50

Speed

Model time, all inputs227.9 s
Mean per input4.56 s
Slowest input15.47 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens38.4K
Output tokens7.2K
Total tokens45.6K
Input cost$0.096
Output cost$0.0722
Total cost$0.1682

Priced from 2026-09-16 · Input $/M $2.50 · Output $/M $10.00

Result 49 of 179

Test T1520 at 2026-09-16

newspaper-page data-correction 18 en

huggingface Qwen/Qwen3-VL-235B-A22B-Instruct · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Added missing closing angle bracket '>' after '<TITLE' to properly close the opening tag, making it '<TITLE>'. The original XML had '<TITLEEine Tossanische Bibel</TITLE>', which is malformed because the opening tag was not properly terminated."}

Scoring

Fuzzy score97.0%
Items50

Speed

Model time, all inputs1202.5 s
Mean per input24.05 s
Slowest input56.29 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens32.1K
Output tokens30.6K
Total tokens62.7K
Input cost$0.0064
Output cost$0.0269
Total cost$0.0333

Priced from 2026-09-16 · Input $/M $0.20 · Output $/M $0.88

Result 50 of 179

Test T1595 at 2026-09-16

newspaper-page data-correction 18 en

genai gemini-3.7-flash · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Fixed malformed opening tag '<TITLEEine' to properly formed '<TITLE>Eine'."}

Scoring

Fuzzy score97.9%
Items50

Speed

Model time, all inputs432.8 s
Mean per input8.66 s
Slowest input19.64 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens31.5K
Output tokens30K
Total tokens61.5K
Input cost$0.0237
Output cost$0.1124
Total cost$0.4767
Reasoning cost usd$0.3407
Total reasoning tokens90.8K

Priced from 2026-08-18 · Input $/M $0.75 · Output $/M $3.75