RISE Humanities Data Benchmark, 0.6.0-pre1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'book_advert_xml__true' with Search Hidden 'False' returned 179 results, showing page 4 of 18.
Result 31 of 179

Test T1047 at 2026-09-16

newspaper-page data-correction 18 en

openai gpt-5.5-2026-04-23 · temp 1.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Corrected the malformed opening TITLE tag from <TITLEEine to <TITLE>."}

Scoring

Fuzzy score98.2%
Items50

Speed

Model time, all inputs748.0 s
Mean per input14.96 s
Slowest input43.18 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens33.6K
Output tokens65.2K
Total tokens98.8K
Input cost$0.1681
Output cost$1.96
Total cost$2.12

Priced from 2026-09-16 · Input $/M $5.00 · Output $/M $30.00

Result 32 of 179

Test T1670 at 2026-09-16

newspaper-page data-correction 18 en

openai gpt-6-astra · temp 1.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Inserted the missing '>' in the opening <TITLE> tag."}

Scoring

Fuzzy score97.4%
Items50

Speed

Model time, all inputs929.6 s
Mean per input18.59 s
Slowest input29.91 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens33.6K
Output tokens35.7K
Total tokens69.3K
Input cost$0.3362
Output cost$1.78
Total cost$2.12

Priced from 2026-09-03 · Input $/M $10.00 · Output $/M $50.00

Result 33 of 179

Test T0650 at 2026-09-16

newspaper-page data-correction 18 en

anthropic claude-sonnet-4-6 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"The opening tag `<TITLEEine` was malformed — it was missing the closing `>` after `TITLE`. It has been corrected to `<TITLE>`, properly separating the tag from the text content \"Eine Tossanische Bibel\"."}

Scoring

Fuzzy score96.7%
Items50

Speed

Model time, all inputs734.3 s
Mean per input14.69 s
Slowest input26.71 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens74.1K
Output tokens46.1K
Total tokens120.2K
Input cost$0.2224
Output cost$0.6909
Total cost$0.9133

Priced from 2026-09-16 · Input $/M $3.00 · Output $/M $15.00

Result 34 of 179

Test T1032 at 2026-09-16

newspaper-page data-correction 18 en

anthropic claude-opus-4-7 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Added the missing closing bracket for the opening <TITLE> tag, which was written as <TITLEEine instead of <TITLE>Eine."}

Scoring

Fuzzy score97.8%
Items50

Speed

Model time, all inputs594.1 s
Mean per input11.88 s
Slowest input27.57 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens101.4K
Output tokens50.5K
Total tokens151.8K
Input cost$0.5069
Output cost$1.26
Total cost$1.77

Priced from 2026-09-16 · Input $/M $5.00 · Output $/M $25.00

Result 35 of 179

Test T1385 at 2026-09-16

newspaper-page data-correction 18 en

openai gpt-5.6-luna · temp 1.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Added the missing closing angle bracket in the opening TITLE tag."}

Scoring

Fuzzy score97.6%
Items50

Speed

Model time, all inputs818.8 s
Mean per input16.37 s
Slowest input130.79 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens33.6K
Output tokens58.9K
Total tokens92.5K
Input cost$0.0067
Output cost$0.0707
Total cost$0.0774

Priced from 2026-09-16 · Input $/M $0.20 · Output $/M $1.20

Result 36 of 179

Test T0483 at 2026-09-16

newspaper-page data-correction 18 en

cohere command-r-08-2024 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score17.9%

{
    "fixed_xml": "<?xml version=\"1.0\" encoding=\"UTF-8\"?>
<book_advert>
    <title>The Great Adventure</title>
    <author>John Smith</author>
    <publisher>Adventure Press</publisher>
    <price>29.99</price>
    <description>Embark on an epic journey with this thrilling novel.</description>
</book_advert>",
    "number_of_fixes": 1,
    "explanation": "The original XML was missing the root element tag. I added <book_advert> as the root element to encapsulate the content."
}

Scoring

Fuzzy score21.2%
Items50

Speed

Model time, all inputs161.1 s
Mean per input3.22 s
Slowest input4.30 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens18.9K
Output tokens7.6K
Total tokens26.5K
Input cost$0.0028
Output cost$0.0046
Total cost$0.0074

Priced from 2026-09-16 · Input $/M $0.15 · Output $/M $0.60

Result 37 of 179

Test T0485 at 2026-09-16

newspaper-page data-correction 18 en

cohere command-r7b-12-2024 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score18.9%

{
  "fixed_xml": "<advertisement><book title='The Great Adventure' author='John Doe' year='2023' genre='Fantasy'/></advertisement>",
  "number_of_fixes": 2,
  "explanation": "Added missing closing tags for the book element and added attributes for title, author, year, and genre."
}

Scoring

Fuzzy score18.1%
Items50

Speed

Model time, all inputs70.5 s
Mean per input1.41 s
Slowest input3.22 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens38.4K
Output tokens5.5K
Total tokens43.9K
Input cost$0.0014
Output cost$0.0008
Total cost$0.0023

Priced from 2026-09-16 · Input $/M $0.0375 · Output $/M $0.15

Result 38 of 179

Test T1430 at 2026-09-16

newspaper-page data-correction 18 en

genai gemini-3.6-flash · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Added missing closing angle bracket to the <TITLE> opening tag."}

Scoring

Fuzzy score94.7%
Items47

Speed

Model time, all inputs1273.7 s
Mean per input25.47 s
Slowest input175.45 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens31.5K
Output tokens114K
Total tokens145.5K
Input cost$0.0237
Output cost$0.4274
Total cost$0.9737
Reasoning cost usd$0.5226
Total reasoning tokens139.4K

Priced from 2026-09-16 · Input $/M $0.75 · Output $/M $3.75

Result 39 of 179

Test T1060 at 2026-09-16

newspaper-page data-correction 18 en

genai gemini-3.5-flash · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Fixed the malformed opening tag <TITLEEine Tossanische Bibel</TITLE> by adding the missing '>' to make it <TITLE>Eine Tossanische Bibel</TITLE>."}

Scoring

Fuzzy score94.5%
Items44

Speed

Model time, all inputs1813.8 s
Mean per input36.28 s
Slowest input240.30 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens31.5K
Output tokens221K
Total tokens252.5K
Input cost$0.0473
Output cost$1.99
Total cost$4.45
Reasoning cost usd$2.42
Total reasoning tokens268.6K

Priced from 2026-09-16 · Input $/M $1.50 · Output $/M $9.00

Result 40 of 179

Test T0520 at 2026-09-16

newspaper-page data-correction 18 en

anthropic claude-opus-4-5-20251101 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Fixed the opening TITLE tag which was missing the closing angle bracket. Changed '<TITLEEine' to '<TITLE>Eine'."}

Scoring

Fuzzy score97.0%
Items50

Speed

Model time, all inputs433.3 s
Mean per input8.67 s
Slowest input18.40 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens74.1K
Output tokens35.4K
Total tokens109.5K
Input cost$0.3705
Output cost$0.8855
Total cost$1.26

Priced from 2026-09-16 · Input $/M $5.00 · Output $/M $25.00