RISE Humanities Data Benchmark, 0.6.0-pre1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'book_advert_xml__true' with Search Hidden 'False' returned 179 results, showing page 9 of 18.
Result 81 of 179

Test T1490 at 2026-08-15

newspaper-page data-correction 18 en

huggingface swiss-ai/apertus-v1.5-70b · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score90.8%

{"fixed_xml":"<ROOT><ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM></ROOT>","number_of_corrections":1,"explanation":"The original XML string was not well-formed. It lacked a root element and proper XML syntax. I added a root element `<ROOT>` to make it a valid XML document."}

Scoring

Fuzzy score83.3%
Items50

Speed

Model time, all inputs615.2 s
Mean per input12.30 s
Slowest input191.69 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens33.3K
Output tokens27.4K
Total tokens60.7K
Input cost$0.0273
Output cost$0.0799
Total cost$0.1073

Priced from 2026-08-10 · Input $/M $0.82 · Output $/M $2.92

Result 82 of 179

Test T1505 at 2026-08-14

newspaper-page data-correction 18 en

huggingface swiss-ai/apertus-v1.5-8b · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score99.2%

{"fixed_xml":"<ITEM>37. <BIBL><TITLEEine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"The provided XML string is not well-formed. The root element is missing, and the tags are not properly closed. The 'ITEM' tag should be the root element, and the 'BIBL' tag should be closed."}

Scoring

Fuzzy score90.7%
Items50

Speed

Model time, all inputs619.0 s
Mean per input12.38 s
Slowest input69.45 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens33.2K
Output tokens27.5K
Total tokens60.7K
Input cost$0.0033
Output cost$0.0055
Total cost$0.0088

Priced from 2026-08-10 · Input $/M $0.10 · Output $/M $0.20

Result 83 of 179

Test T1460 at 2026-08-04

newspaper-page data-correction 18 en

x-ai grok-4.5 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Added missing '>' to close the opening TITLE tag, which was malformed as <TITLEEine..."}

Scoring

Fuzzy score97.7%
Items50

Speed

Model time, all inputs2784.0 s
Mean per input55.68 s
Slowest input100.61 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens47.3K
Output tokens27.8K
Total tokens75.1K
Input cost$0.0946
Output cost$0.167
Total cost$1.19
Reasoning cost usd$0.9309
Total reasoning tokens155.1K

Priced from 2026-08-04 · Input $/M $2.00 · Output $/M $6.00

Result 84 of 179

Test T1400 at 2026-07-25

newspaper-page data-correction 18 en

openrouter moonshotai/kimi-k3 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"The opening <TITLE> tag was malformed: the closing angle bracket was missing, so the tag name and the title text were merged into '<TITLEEine Tossanische Bibel'. Added the missing '>' to properly open the TITLE element, yielding '<TITLE>Eine Tossanische Bibel</TITLE>'. The rest of the structure (ITEM, BIBL, closing TITLE) was already well-formed."}

Scoring

Fuzzy score98.7%
Items50

Speed

Model time, all inputs6677.4 s
Mean per input133.55 s
Slowest input604.13 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens43.2K
Output tokens244.5K
Total tokens287.6K
Input cost$0.1262
Output cost$3.49
Total cost$3.79

Priced from 2026-07-25 · Input $/M $3.00 · Output $/M $15.00

Result 85 of 179

Test T1385 at 2026-07-25

newspaper-page data-correction 18 en

openai gpt-5.6-luna · temp 1.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Added the missing opening angle bracket in the TITLE tag."}

Scoring

Fuzzy score97.7%
Items50

Speed

Model time, all inputs462.2 s
Mean per input9.24 s
Slowest input20.65 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens33.6K
Output tokens62.5K
Total tokens96.2K
Input cost$0.0336
Output cost$0.3752
Total cost$0.4088

Priced from 2026-07-25 · Input $/M $1.00 · Output $/M $6.00

Result 86 of 179

Test T1415 at 2026-07-25

newspaper-page data-correction 18 en

anthropic claude-opus-5 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"The opening <TITLE> tag was malformed: it was missing its closing angle bracket, so the tag name ran into the title text (\"<TITLEEine Tossanische Bibel\"). I inserted the missing \">\" so the element reads <TITLE>Eine Tossanische Bibel</TITLE>. All other tags (<ITEM>, <BIBL>) were already properly opened, nested and closed, and the item number \"37.\" remains as text directly inside <ITEM>, as is conventional for Avisblatt book advert entries."}

Scoring

Fuzzy score97.5%
Items50

Speed

Model time, all inputs694.6 s
Mean per input13.89 s
Slowest input26.88 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens81.2K
Output tokens58.2K
Total tokens139.4K
Input cost$0.4061
Output cost$1.46
Total cost$1.86

Priced from 2026-07-25 · Input $/M $5.00 · Output $/M $25.00

Result 87 of 179

Test T1370 at 2026-07-25

newspaper-page data-correction 18 en

openai gpt-5.6-terra · temp 1.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score98.4%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Toskanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":2,"explanation":"Corrected the malformed opening TITLE tag by adding the missing closing angle bracket, and normalized the apparent typo “Tossanische” to “Toskanische.”"}

Scoring

Fuzzy score98.1%
Items50

Speed

Model time, all inputs332.6 s
Mean per input6.65 s
Slowest input15.66 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens33.6K
Output tokens37.1K
Total tokens70.8K
Input cost$0.0841
Output cost$0.5572
Total cost$0.6412

Priced from 2026-07-25 · Input $/M $2.50 · Output $/M $15.00

Result 88 of 179

Test T1445 at 2026-07-25

newspaper-page data-correction 18 en

genai gemini-3.5-flash-lite · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Added the missing closing '>' for the <TITLE> tag."}

Scoring

Fuzzy score96.6%
Items50

Speed

Model time, all inputs159.0 s
Mean per input3.18 s
Slowest input12.55 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens31.5K
Output tokens25.8K
Total tokens57.4K
Input cost$0.0095
Output cost$0.0646
Total cost$0.074

Priced from 2026-07-25 · Input $/M $0.30 · Output $/M $2.50

Result 89 of 179

Test T1430 at 2026-07-25

newspaper-page data-correction 18 en

genai gemini-3.6-flash · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Fixed malformed open tag '<TITLE' by changing it to '<TITLE>'."}

Scoring

Fuzzy score90.7%
Items50

Speed

Model time, all inputs1349.3 s
Mean per input26.99 s
Slowest input224.16 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens31.5K
Output tokens215.6K
Total tokens247.2K
Input cost$0.0473
Output cost$1.62
Total cost$2.73
Reasoning cost usd$1.06
Total reasoning tokens141.9K

Priced from 2026-07-25 · Input $/M $1.50 · Output $/M $7.50

Result 90 of 179

Test T1355 at 2026-07-25

newspaper-page data-correction 18 en

openai gpt-5.6-sol · temp 1.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Added the missing closing angle bracket to the opening TITLE tag."}

Scoring

Fuzzy score98.4%
Items50

Speed

Model time, all inputs522.5 s
Mean per input10.45 s
Slowest input38.38 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens33.6K
Output tokens39.6K
Total tokens73.2K
Input cost$0.1681
Output cost$1.19
Total cost$1.36

Priced from 2026-07-25 · Input $/M $5.00 · Output $/M $30.00