RISE Humanities Data Benchmark, 0.6.1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'book_advert_xml__true' with Search Hidden 'False' returned 187 results, showing page 18 of 19.
Result 171 of 187

Test T0468 at 2025-12-08

newspaper-page data-correction 18 en

anthropic claude-sonnet-4-5-20250929 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Fixed missing closing angle bracket in the opening TITLE tag. Changed \"<TITLEEine\" to \"<TITLE>Eine\" to properly separate the tag from the content."}

Scoring

Fuzzy score97.5%
Items50

Speed

Model time, all inputs544.0 s
Mean per input10.88 s
Slowest input18.50 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens74.1K
Output tokens36.8K
Total tokens110.9K
Input cost$0.2223
Output cost$0.5514
Total cost$0.7736

Priced from 2025-11-24 · Input $/M $3.00 · Output $/M $15.00

Result 172 of 187

Test T0450 at 2025-12-08

newspaper-page data-correction 18 en

openai gpt-4.1-nano-2025-04-14 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Added missing '>' after <TITLE> tag to properly close the tag."}

Scoring

Fuzzy score93.4%
Items50

Speed

Model time, all inputs315.4 s
Mean per input6.31 s
Slowest input15.09 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens33.8K
Output tokens25.4K
Total tokens59.1K
Input cost$0.0034
Output cost$0.0101
Total cost$0.0135

Priced from 2025-11-24 · Input $/M $0.10 · Output $/M $0.40

Result 173 of 187

Test T0452 at 2025-12-08

newspaper-page data-correction 18 en

openai gpt-5-mini-2025-08-07 · temp 1.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Inserted the missing '>' to correctly open the <TITLE> tag (<TITLE> instead of <TITLE). The XML is now well-formed."}

Scoring

Fuzzy score91.1%
Items50

Speed

Model time, all inputs2740.6 s
Mean per input54.81 s
Slowest input126.56 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens33.7K
Output tokens176.7K
Total tokens210.4K
Input cost$0.0084
Output cost$0.3534
Total cost$0.3618

Priced from 2025-11-24 · Input $/M $0.25 · Output $/M $2.00

Result 174 of 187

Test T0453 at 2025-12-08

newspaper-page data-correction 18 en

openai gpt-5-nano-2025-08-07 · temp 1.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Fixed malformed TITLE tag by inserting the missing '>' after TITLE to close the opening tag. The corrected XML is well-formed: <ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>."}

Scoring

Fuzzy score94.2%
Items50

Speed

Model time, all inputs5419.8 s
Mean per input108.40 s
Slowest input205.33 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens33.6K
Output tokens600.3K
Total tokens633.9K
Input cost$0.0017
Output cost$0.2401
Total cost$0.2418

Priced from 2025-11-24 · Input $/M $0.05 · Output $/M $0.40

Result 175 of 187

Test T0458 at 2025-12-08

newspaper-page data-correction 18 en

genai gemini-2.5-pro · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"The opening tag for TITLE was malformed. The text content 'Eine Tossanische Bibel' was incorrectly included within the tag name. The tag has been corrected to <TITLE> and the text placed as its content."}

Scoring

Fuzzy score91.5%
Items50

Speed

Model time, all inputs2918.1 s
Mean per input58.36 s
Slowest input112.24 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens31.5K
Output tokens28.6K
Total tokens60.1K
Input cost$0.0394
Output cost$0.2857
Total cost$3.99
Reasoning cost usd$3.66
Total reasoning tokens366.2K

Priced from 2025-11-24 · Input $/M $1.25 · Output $/M $10.00

Result 176 of 187

Test T0448 at 2025-12-08

newspaper-page data-correction 18 en

openai gpt-4.1-2025-04-14 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":2,"explanation":"1. Added missing '>' after <TITLE to properly open the TITLE tag. 2. Inserted a space between <TITLE and the text content. These changes ensure the XML is well-formed and the TITLE element is correctly structured."}

Scoring

Fuzzy score96.5%
Items50

Speed

Model time, all inputs735.1 s
Mean per input14.70 s
Slowest input290.47 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens33.8K
Output tokens28.6K
Total tokens62.4K
Input cost$0.0676
Output cost$0.2288
Total cost$0.2965

Priced from 2025-11-24 · Input $/M $2.00 · Output $/M $8.00

Result 177 of 187

Test T0447 at 2025-12-08

newspaper-page data-correction 18 en

openai gpt-4o-mini-2024-07-18 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"The original XML had a malformed TITLE tag. I added the missing closing angle bracket for the TITLE tag."}

Scoring

Fuzzy score91.8%
Items50

Speed

Model time, all inputs2114.4 s
Mean per input42.29 s
Slowest input672.97 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens33.9K
Output tokens56.2K
Total tokens90K
Input cost$0.0051
Output cost$0.0337
Total cost$0.0388

Priced from 2025-11-24 · Input $/M $0.15 · Output $/M $0.60

Result 178 of 187

Test T0460 at 2025-12-08

newspaper-page data-correction 18 en

genai gemini-2.5-flash-lite · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"The closing tag for the TITLE element was missing and has been added."}

Scoring

Fuzzy score63.6%
Items50

Speed

Model time, all inputs4236.6 s
Mean per input84.73 s
Slowest input342.23 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens31.5K
Output tokens1M
Total tokens1M
Input cost$0.0032
Output cost$0.4014
Total cost$0.4045

Priced from 2025-11-24 · Input $/M $0.10 · Output $/M $0.40

Result 179 of 187

Test T0478 at 2025-12-08

newspaper-page data-correction 18 en

openrouter meta-llama/llama-4-maverick · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score61.8%

{"fixed_xml":"<file name=\"line_1.json\"><content><![CDATA[{
  \"xml_string\": \"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>\"
}]]></content></file>","number_of_corrections":3,"explanation":"The original XML was not well-formed because it contained JSON data directly inside the <file> element. To fix this, I added a <content> element to wrap the JSON data and used CDATA to escape the JSON content so it doesn't interfere with the XML parsing. Additionally, I properly closed all the tags and ensured the XML is well-formed."}

Scoring

Fuzzy score67.2%
Items50

Speed

Model time, all inputs561.3 s
Mean per input11.23 s
Slowest input41.98 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens29.7K
Output tokens24.8K
Total tokens54.5K
Input cost$0.0049
Output cost$0.0153
Total cost$0.0201

Priced from 2025-11-24 · Input $/M $0.15 · Output $/M $0.60

Result 180 of 187

Test T0477 at 2025-12-08

newspaper-page data-correction 18 en

openrouter qwen/qwen3-vl-8b-thinking · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Score0.0%

Scoring

Fuzzy score1.9%
Items50

Speed

Model time, all inputs457.3 s
Mean per input9.15 s
Slowest input21.86 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens21K
Output tokens23.2K
Total tokens44.2K
Input cost$0.0038
Output cost$0.0488
Total cost$0.0525

Priced from 2025-11-24 · Input $/M $0.18 · Output $/M $2.10