RISE Humanities Data Benchmark, 0.6.0-pre1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'book_advert_xml__true' with Search Hidden 'False' returned 179 results, showing page 1 of 18.
Result 1 of 179

Test T1073 at 2026-09-18

newspaper-page data-correction 18 en

openrouter qwen/qwen3.7-plus · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"The opening tag for the title element was malformed and missing a closing angle bracket ('<TITLEEine' instead of '<TITLE>'). This has been corrected to ensure the XML is well-formed. As a historian, I also note that 'Tossanische' is likely an OCR error or archaic spelling for 'Toskanische' (Tuscan), but the structural XML fault has been resolved."}

Scoring

Fuzzy score97.1%
Items45

Speed

Model time, all inputs6646.4 s
Mean per input141.41 s
Slowest input414.43 s
Inputs timed47

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens30.7K
Output tokens339.7K
Total tokens370.4K
Input cost$0.0066
Output cost$0.3323
Total cost$0.4446

Priced from 2026-09-16 · Input $/M $0.32 · Output $/M $1.28

Result 2 of 179

Test T1790 at 2026-09-18

newspaper-page data-correction 18 en

openrouter qwen/qwen3.8-max-0902 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"The opening tag <TITLE> was malformed because its closing angle bracket was missing, causing the element name to run into the text content. I corrected it to <TITLE> so the TITLE element is properly opened and closed within BIBL and ITEM."}

Scoring

Fuzzy score97.4%
Items48

Speed

Model time, all inputs13892.9 s
Mean per input277.86 s
Slowest input748.54 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens34.5K
Output tokens475.4K
Total tokens509.9K
Input cost$0.0555
Output cost$2.38
Total cost$2.92

Priced from 2026-09-16 · Input $/M $2.00 · Output $/M $6.00

Result 3 of 179

Test T1520 at 2026-09-18

newspaper-page data-correction 18 en

huggingface Qwen/Qwen3-VL-235B-A22B-Instruct · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Added missing closing angle bracket '>' after 'TITLE' to properly close the opening tag. The original tag '<TITLEEine Tossanische Bibel' was malformed and should be '<TITLE>Eine Tossanische Bibel</TITLE>'."}

Scoring

Fuzzy score97.8%
Items50

Speed

Model time, all inputs2530.0 s
Mean per input50.60 s
Slowest input158.34 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens32.1K
Output tokens31.3K
Total tokens63.4K
Input cost$0.0064
Output cost$0.0276
Total cost$0.034

Priced from 2026-09-16 · Input $/M $0.20 · Output $/M $0.88

Result 4 of 179

Test T0478 at 2026-09-18

newspaper-page data-correction 18 en

openrouter meta-llama/llama-4-maverick · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score82.1%

{"fixed_xml":"<ITEMS><ITEM><TEXT>37.</TEXT><BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM></ITEMS>","number_of_corrections":3,"explanation":"The provided XML string is mostly well-formed but lacks a proper root element and has some formatting issues. The 'TITLE' element contains text that starts with a capital letter but is not properly separated from the preceding element. I've added a root element 'ITEMS' to encapsulate the 'ITEM' element. I've also added a 'TEXT' element around the text '37.' to structure it properly within the 'ITEM' element. The 'TITLE' element's content has been adjusted to be properly formatted."}

Scoring

Fuzzy score92.0%
Items50

Speed

Model time, all inputs1775.4 s
Mean per input35.51 s
Slowest input83.89 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens28.2K
Output tokens27.9K
Total tokens56.2K
Input cost$0.0053
Output cost$0.0182
Total cost$0.0235

Priced from 2026-09-16 · Input $/M $0.1875 · Output $/M $0.6525

Result 5 of 179

Test T1715 at 2026-09-18

newspaper-page data-correction 18 en

openrouter z-ai/glm-5.3-flash · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"The opening <TITLE> tag was malformed: it was missing the closing angle bracket (>), so the tag name and its text content were merged into '<TITLEEine Tossanische Bibel'. Inserting '>' after 'TITLE' produces a well-formed element: <TITLE>Eine Tossanische Bibel</TITLE>. All other tags (<ITEM>, <BIBL>) were correctly formed and properly nested, so no further changes were necessary."}

Scoring

Fuzzy score97.5%
Items33

Speed

Model time, all inputs10713.8 s
Mean per input214.28 s
Slowest input797.16 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens35.9K
Output tokens446.6K
Total tokens482.5K
Input cost$0.0011
Output cost$0.0272
Total cost$0.1956

Priced from 2026-09-08 · Input $/M $0.075 · Output $/M $0.25

Result 6 of 179

Test T1730 at 2026-09-18

newspaper-page data-correction 18 en

openrouter qwen/qwen3.8-flash · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Added the missing closing angle bracket '>' to the opening <TITLE> tag so that the XML is well-formed."}

Scoring

Fuzzy score97.5%
Items41

Speed

Model time, all inputs12282.6 s
Mean per input250.66 s
Slowest input537.33 s
Inputs timed49

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens34.9K
Output tokens513.2K
Total tokens548.1K
Input cost$0.003
Output cost$0.1346
Total cost$0.2452

Priced from 2026-09-18 · Input $/M $0.15 · Output $/M $0.47

Result 7 of 179

Test T1745 at 2026-09-18

newspaper-page data-correction 18 en

openrouter qwen/qwen3.8-27b · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"The start tag was malformed as <TITLEEine Tossanische Bibel>. A space was added after the tag name TITLE so that the title text is properly enclosed in <TITLE>...</TITLE>."}

Scoring

Fuzzy score98.4%
Items44

Speed

Model time, all inputs16958.0 s
Mean per input360.81 s
Slowest input1059.26 s
Inputs timed47

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens35K
Output tokens495.1K
Total tokens530.1K
Input cost$0.0059
Output cost$0.9776
Total cost$1.23

Priced from 2026-09-18 · Input $/M $0.214 · Output $/M $2.55

Result 8 of 179

Test T1019 at 2026-09-18

newspaper-page data-correction 18 en

openrouter google/gemma-4-31b-it · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"The opening tag for the title was malformed; it lacked the closing angle bracket ('>'), causing the text 'Eine Tossanische Bibel' to be treated as part of the tag definition rather than the element content. I closed the <TITLE> tag to properly encapsulate the bibliographic data."}

Scoring

Fuzzy score97.3%
Items50

Speed

Model time, all inputs1121.2 s
Mean per input22.42 s
Slowest input50.01 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens32.3K
Output tokens33.7K
Total tokens66K
Input cost$0.0029
Output cost$0.0115
Total cost$0.0144

Priced from 2026-09-16 · Input $/M $0.09 · Output $/M $0.34

Result 9 of 179

Test T1640 at 2026-09-18

newspaper-page data-correction 18 en

openrouter z-ai/glm-5v-turbo · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"The opening tag for the element 'TITLE' was malformed, missing the closing angle bracket (>). I inserted the bracket to properly delimit the tag from the text content."}

Scoring

Fuzzy score98.6%
Items48

Speed

Model time, all inputs7555.3 s
Mean per input151.11 s
Slowest input327.42 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens40.2K
Output tokens89.2K
Total tokens129.5K
Input cost$0.00
Output cost$0.00
Total cost$0.4052

Priced from 2026-08-18 · Input $/M $1.20 · Output $/M $4.00

Result 10 of 179

Test T0941 at 2026-09-17

newspaper-page data-correction 18 en

openrouter qwen/qwen3.5-35b-a3b · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

Fuzzy score89.0%
Items30

Speed

Model time, all inputs13679.7 s
Mean per input273.59 s
Slowest input719.23 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens39.1K
Output tokens916.9K
Total tokens956K
Input cost$0.0062
Output cost$1.15
Total cost$0.8486

Priced from 2026-09-16 · Input $/M $0.1625 · Output $/M $1.30