RISE Humanities Data Benchmark, 0.6.1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'book_advert_xml__true' with Search Hidden 'False' returned 187 results, showing page 11 of 19.
Result 101 of 187

Test T1187 at 2026-07-01

newspaper-page data-correction 18 en

anthropic claude-sonnet-5 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"The opening tag <TITLE> was malformed as \"<TITLEEine Tossanische Bibel</TITLE>\", missing the closing angle bracket \">\" after \"TITLE\". Corrected it to \"<TITLE>Eine Tossanische Bibel</TITLE>\" so the tag properly opens and closes around the text content."}

Scoring

Fuzzy score96.6%
Items50

Speed

Model time, all inputs593.5 s
Mean per input11.87 s
Slowest input26.88 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens84.6K
Output tokens61.3K
Total tokens145.9K
Input cost$0.1692
Output cost$0.6127
Total cost$0.7819

Priced from 2026-07-01 · Input $/M $2.00 · Output $/M $10.00

Result 102 of 187

Test T0483 at 2026-06-30

newspaper-page data-correction 18 en

cohere command-r-08-2024 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score16.9%

{
    "fixed_xml": "<?xml version=\"1.0\" encoding=\"UTF-8\"?>
<advert>
    <title>The Great Adventure</title>
    <author>John Smith</author>
    <price>19.99</price>
    <description>Embark on an epic journey with this thrilling novel. Follow the hero's quest and uncover the secrets of the ancient world.</description>
    <publisher>Adventure Press</publisher>
    <isbn>978-3-423-42342-3</isbn>
</advert>",
    "number_of_fixes": 1,
    "explanation": "The original XML was missing the root element. I added the <advert> tag to encapsulate all the content."
}

Scoring

Fuzzy score21.4%
Items50

Speed

Model time, all inputs298.6 s
Mean per input5.97 s
Slowest input112.09 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens18.9K
Output tokens7.7K
Total tokens26.6K
Input cost$0.0007
Output cost$0.0012
Total cost$0.0019

Priced from 2026-06-11 · Input $/M $0.15 · Output $/M $0.60

Result 103 of 187

Test T0485 at 2026-06-30

newspaper-page data-correction 18 en

cohere command-r7b-12-2024 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score18.9%

{
  "fixed_xml": "<advertisement><book title='The Great Adventure' author='John Doe' year='2023' genre='Fantasy'/></advertisement>",
  "number_of_fixes": 2,
  "explanation": "Added missing closing tags for the book element and ensured proper nesting. Also, added missing attributes for the book element."
}

Scoring

Fuzzy score19.5%
Items50

Speed

Model time, all inputs137.1 s
Mean per input2.74 s
Slowest input6.13 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens38.4K
Output tokens6.2K
Total tokens44.6K
Input cost$0.0014
Output cost$0.0009
Total cost$0.0024

Priced from 2026-03-02 · Input $/M $0.0375 · Output $/M $0.15

Result 104 of 187

Test T0484 at 2026-06-30

newspaper-page data-correction 18 en

cohere command-r-plus-08-2024 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score17.2%

{"fixed_xml":"<?xml version=\"1.0\" encoding=\"UTF-8\"?>
<book>
  <title>The Adventures of Huckleberry Finn</title>
  <author>Mark Twain</author>
  <genre>Adventure</genre>
  <price currency=\"USD\">12.99</price>
  <description>A classic tale of a young boy's journey down the Mississippi River.</description>
</book>","number_of_corrections":1,"explanation":"Added a missing XML declaration at the beginning of the document."}

Scoring

Fuzzy score8.2%
Items50

Speed

Model time, all inputs259.9 s
Mean per input5.20 s
Slowest input13.12 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens18.9K
Output tokens8.3K
Total tokens27.2K
Input cost$0.0473
Output cost$0.083
Total cost$0.1302

Priced from 2026-06-23 · Input $/M $2.50 · Output $/M $10.00

Result 105 of 187

Test T0486 at 2026-06-30

newspaper-page data-correction 18 en

cohere command-a-03-2025 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score15.8%

{"fixed_xml":"<advertisement><book><title>The Great Adventure</title><author>John Doe</author><price>19.99</price><description>An exciting journey through unknown lands.</description></book><book><title>Mystery of the Lost City</title><author>Jane Smith</author><price>24.95</price><description>A thrilling mystery set in an ancient city.</description></book></advertisement>","number_of_corrections":3,"explanation":"1. Added a root element <advertisement> to encapsulate all content. 2. Added missing closing tags for <book> elements to ensure proper nesting. 3. Ensured all elements are properly closed and nested within the XML structure."}

Scoring

Fuzzy score23.6%
Items50

Speed

Model time, all inputs319.4 s
Mean per input6.39 s
Slowest input30.03 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens38.4K
Output tokens8K
Total tokens46.4K
Input cost$0.096
Output cost$0.0802
Total cost$0.1762

Priced from 2026-03-02 · Input $/M $2.50 · Output $/M $10.00

Result 106 of 187

Test T1174 at 2026-06-29

newspaper-page data-correction 18 en

genai gemini-3.1-flash-lite · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Added a missing closing bracket to the TITLE tag."}

Scoring

Fuzzy score96.4%
Items50

Speed

Model time, all inputs103.6 s
Mean per input2.07 s
Slowest input3.92 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens31.5K
Output tokens24.7K
Total tokens56.2K
Input cost$0.0079
Output cost$0.037
Total cost$0.0449

Priced from 2026-06-29 · Input $/M $0.25 · Output $/M $1.50

Result 107 of 187

Test T1112 at 2026-06-08

newspaper-page data-correction 18 en

x-ai grok-4.3 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Corrected malformed opening <TITLE> tag (missing '>')."}

Scoring

Fuzzy score98.6%
Items50

Speed

Model time, all inputs667.1 s
Mean per input13.34 s
Slowest input24.75 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens40.7K
Output tokens26.3K
Total tokens66.9K
Input cost$0.0508
Output cost$0.0656
Total cost$0.2387
Reasoning cost usd$0.1223
Total reasoning tokens48.9K

Priced from 2026-06-05 · Input $/M $1.25 · Output $/M $2.50

Result 108 of 187

Test T1125 at 2026-06-05

newspaper-page data-correction 18 en

openrouter meta-llama/llama-4-scout-17b-16e-instruct · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Added closing tag for TITLE element and also added a space between 'TITLE' and 'Eine' for proper XML syntax."}

Scoring

Fuzzy score56.6%
Items50

Speed

Model time, all inputs405.5 s
Mean per input8.11 s
Slowest input21.38 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens29.1K
Output tokens17.5K
Total tokens46.6K
Input cost$0.0023
Output cost$0.0052
Total cost$0.0076

Priced from 2026-06-05 · Input $/M $0.08 · Output $/M $0.30

Result 109 of 187

Test T1138 at 2026-06-05

newspaper-page data-correction 18 en

openrouter stepfun/step-3.7-flash-20260528 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"The only fault was a malformed opening <TITLE> tag missing its closing angle bracket, which created a tag name mismatch with the corresponding closing </TITLE> tag. The fix adds the missing > to correctly open the TITLE element, restoring XML well-formedness. No additional tags were added as the existing text content (the '37.' prefix) is valid as child content of the <ITEM> element."}

Scoring

Fuzzy score86.3%
Items50

Speed

Model time, all inputs37230.2 s
Mean per input744.60 s
Slowest input27555.08 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens31.7K
Output tokens1.1M
Total tokens1.1M
Input cost$0.0063
Output cost$1.24
Total cost$1.25

Priced from 2026-06-05 · Input $/M $0.20 · Output $/M $1.15

Result 110 of 187

Test T0542 at 2026-06-04

newspaper-page data-correction 18 en

mistral mistral-large-2512 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score80.5%

{"fixed_xml":"<file name=\"line_1.json\"><ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM></file>","number_of_corrections":1,"explanation":"Added missing closing angle bracket for the opening TITLE tag. The original XML had '<TITLEEine Tossanische Bibel</TITLE>', which was corrected to '<TITLE>Eine Tossanische Bibel</TITLE>'."}

Scoring

Fuzzy score57.0%
Items50

Speed

Model time, all inputs509.9 s
Mean per input10.20 s
Slowest input24.94 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens24.3K
Output tokens21.9K
Total tokens46.2K
Input cost$0.0122
Output cost$0.0328
Total cost$0.045

Priced from 2026-03-02 · Input $/M $0.50 · Output $/M $1.50