RISE Humanities Data Benchmark, 0.6.0-pre1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'book_advert_xml__true' with Search Hidden 'False' returned 179 results, showing page 14 of 18.
Result 131 of 179

Test T0638 at 2026-03-16

newspaper-page data-correction 18 en

anthropic claude-opus-4-6 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"The opening TITLE tag was missing its closing angle bracket. Changed \"&lt;TITLEEine\" to \"&lt;TITLE&gt;Eine\" to properly form the opening &lt;TITLE&gt; tag."}

Scoring

Fuzzy score97.5%
Items50

Speed

Model time, all inputs648.6 s
Mean per input12.97 s
Slowest input21.58 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens74.1K
Output tokens41.3K
Total tokens115.4K
Input cost$0.3707
Output cost$1.03
Total cost$1.40

Priced from 2026-03-16 · Input $/M $5.00 · Output $/M $25.00

Result 132 of 179

Test T0691 at 2026-03-16

newspaper-page data-correction 18 en

genai gemini-3.1-pro-preview · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Added missing closing angle bracket '>' to the opening TITLE tag."}

Scoring

Fuzzy score94.3%
Items50

Speed

Model time, all inputs2734.8 s
Mean per input54.70 s
Slowest input480.37 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens31.5K
Output tokens31.9K
Total tokens63.5K
Input cost$0.0631
Output cost$0.3833
Total cost$4.31
Reasoning cost usd$3.86
Total reasoning tokens321.9K

Priced from 2026-03-16 · Input $/M $2.00 · Output $/M $12.00

Result 133 of 179

Test T0662 at 2026-03-16

newspaper-page data-correction 18 en

openai gpt-5.4-2026-03-05 · temp 1.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Corrected the malformed opening TITLE tag from '<TITLEEine' to '<TITLE>' so the XML is well-formed."}

Scoring

Fuzzy score96.8%
Items50

Speed

Model time, all inputs354.4 s
Mean per input7.09 s
Slowest input14.44 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens33.6K
Output tokens26.2K
Total tokens59.8K
Input cost$0.0841
Output cost$0.3932
Total cost$0.4773

Priced from 2026-03-16 · Input $/M $2.50 · Output $/M $15.00

Result 134 of 179

Test T0650 at 2026-03-16

newspaper-page data-correction 18 en

anthropic claude-sonnet-4-6 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"The opening tag `<TITLEEine` was malformed — it was missing the closing `>` after `TITLE`. It has been corrected to `<TITLE>`, properly separating the tag from the text content \"Eine Tossanische Bibel\"."}

Scoring

Fuzzy score96.5%
Items50

Speed

Model time, all inputs622.8 s
Mean per input12.46 s
Slowest input20.30 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens74.1K
Output tokens46.3K
Total tokens120.4K
Input cost$0.2224
Output cost$0.6945
Total cost$0.9169

Priced from 2026-03-16 · Input $/M $3.00 · Output $/M $15.00

Result 135 of 179

Test T0498 at 2026-02-10

newspaper-page data-correction 18 en

openai gpt-5.2-2025-12-11 · temp 1.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Added missing opening/closing tags around the TITLE text: changed <TITLEEine Tossanische Bibel</TITLE> to <TITLE>Eine Tossanische Bibel</TITLE>. No other structural issues detected."}

Scoring

Fuzzy score94.9%
Items50

Speed

Model time, all inputs444.4 s
Mean per input8.89 s
Slowest input16.35 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens33.7K
Output tokens28.4K
Total tokens62K
Input cost$0.0589
Output cost$0.3972
Total cost$0.4561

Priced from 2026-01-23 · Input $/M $1.75 · Output $/M $14.00

Result 136 of 179

Test T0454 at 2026-01-26

newspaper-page data-correction 18 en

openai gpt-5.1-2025-11-13 · temp 1.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score93.4%

{"fixed_xml":"<ITEM>
  <NR>37.</NR>
  <BIBL>
    <TITLE>Eine Tossanische Bibel</TITLE>
  </BIBL>
</ITEM>","number_of_corrections":3,"explanation":"1) Wrapped the leading number and dot (\"37.\") in a <NR> element to separate record numbering from the bibliographic block. 2) Inserted the missing closing angle bracket in <TITLEEine Tossanische Bibel</TITLE> so that it becomes a valid <TITLE> element. 3) Reformatted the XML with line breaks and indentation for clarity; structural content remains the same aside from the added <NR> element."}

Scoring

Fuzzy score95.1%
Items50

Speed

Model time, all inputs542.0 s
Mean per input10.84 s
Slowest input24.69 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens33.6K
Output tokens37.8K
Total tokens71.5K
Input cost$0.042
Output cost$0.3785
Total cost$0.4205

Priced from 2026-01-23 · Input $/M $1.25 · Output $/M $10.00

Result 137 of 179

Test T0453 at 2026-01-26

newspaper-page data-correction 18 en

openai gpt-5-nano-2025-08-07 · temp 1.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Fixed missing closing angle bracket after TITLE tag in <TITLE> Eine Tossanische Bibel</TITLE>. The corrected XML is well-formed with properly nested elements."}

Scoring

Fuzzy score93.6%
Items50

Speed

Model time, all inputs4363.6 s
Mean per input87.27 s
Slowest input197.80 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens33.6K
Output tokens573.2K
Total tokens606.9K
Input cost$0.0017
Output cost$0.2293
Total cost$0.231

Priced from 2026-01-23 · Input $/M $0.05 · Output $/M $0.40

Result 138 of 179

Test T0455 at 2026-01-26

newspaper-page data-correction 18 en

openai o3-2025-04-16 · temp 1.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score92.1%

{"fixed_xml":"<ITEM><NUM>37.</NUM><BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":2,"explanation":"1. Inserted missing closing angle bracket in the opening TITLE tag (<TITLE>).
2. Wrapped the ordinal \"37.\" inside a <NUM> element to separate metadata from bibliographic content, improving structural clarity."}

Scoring

Fuzzy score92.4%
Items50

Speed

Model time, all inputs860.0 s
Mean per input17.20 s
Slowest input33.15 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens33.6K
Output tokens142.2K
Total tokens175.8K
Input cost$0.0672
Output cost$1.14
Total cost$1.21

Priced from 2026-01-23 · Input $/M $2.00 · Output $/M $8.00

Result 139 of 179

Test T0451 at 2026-01-25

newspaper-page data-correction 18 en

openai gpt-5-2025-08-07 · temp 1.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Added the missing closing angle bracket in the TITLE start tag to form <TITLE>."}

Scoring

Fuzzy score95.0%
Items50

Speed

Model time, all inputs3962.2 s
Mean per input79.24 s
Slowest input127.34 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens33.6K
Output tokens270.6K
Total tokens304.2K
Input cost$0.042
Output cost$2.71
Total cost$2.75

Priced from 2026-01-23 · Input $/M $1.25 · Output $/M $10.00

Result 140 of 179

Test T0448 at 2026-01-25

newspaper-page data-correction 18 en

openai gpt-4.1-2025-04-14 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":2,"explanation":"1. Added missing '>' after <TITLE to form <TITLE>.
2. Inserted a space between <TITLE> and the book title for readability, though this is optional in XML."}

Scoring

Fuzzy score94.7%
Items50

Speed

Model time, all inputs1299.8 s
Mean per input26.00 s
Slowest input520.12 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens33.9K
Output tokens61.7K
Total tokens95.6K
Input cost$0.0678
Output cost$0.4934
Total cost$0.5611

Priced from 2026-01-23 · Input $/M $2.00 · Output $/M $8.00