RISE Humanities Data Benchmark, 0.6.1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'book_advert_xml__true' with Search Hidden 'False' returned 186 results, showing page 1 of 19.
Result 1 of 186

Test T1820 at 2026-09-29

newspaper-page data-correction 18 en

anthropic claude-opus-5-5 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"The opening <TITLE> tag was missing its closing '>' (it read '<TITLEEine'). I added the bracket so the tag is well-formed and the title text 'Eine Tossanische Bibel' sits inside it. No other tags were needed because the existing ITEM/BIBL/TITLE structure is correct."}

Scoring

Fuzzy score96.5%
Items50

Speed

Model time, all inputs970.5 s
Mean per input19.41 s
Slowest input34.94 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens76.4K
Output tokens96.3K
Total tokens172.8K
Input cost$0.3057
Output cost$1.93
Total cost$2.23

Priced from 2026-09-27 · Input $/M $4.00 · Output $/M $20.00

Result 2 of 186

Test T1880 at 2026-09-29

newspaper-page data-correction 18 en

anthropic claude-sonnet-5-5 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"The opening tag <TITLE was missing its closing angle bracket. I changed it to <TITLE>. The rest of the structure (ITEM, BIBL, TITLE) was already well-formed. I left the text \"Tossanische\" as it appears in the original, although it may be a transcription of \"Toscanische\"."}

Scoring

Fuzzy score98.3%
Items50

Speed

Model time, all inputs560.3 s
Mean per input11.21 s
Slowest input24.05 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens76.4K
Output tokens79.7K
Total tokens156.2K
Input cost$0.1528
Output cost$0.7974
Total cost$0.9503

Priced from 2026-09-28 · Input $/M $2.00 · Output $/M $10.00

Result 3 of 186

Test T1685 at 2026-09-29

newspaper-page data-correction 18 en

anthropic claude-fable-5-1 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"The opening TITLE tag was malformed: \"<TITLEEine Tossanische Bibel</TITLE>\" lacked the closing \">\" after the tag name, so the text was absorbed into the element name and the closing </TITLE> had no matching opener. Inserted the missing \">\" to yield \"<TITLE>Eine Tossanische Bibel</TITLE>\". The surrounding ITEM and BIBL structure was already well-formed and left unchanged."}

Scoring

Fuzzy score97.5%
Items50

Speed

Model time, all inputs1521.9 s
Mean per input30.44 s
Slowest input61.40 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens76.4K
Output tokens122.8K
Total tokens199.3K
Input cost$0.7642
Output cost$6.14
Total cost$6.91

Priced from 2026-09-08 · Input $/M $10.00 · Output $/M $50.00

Result 4 of 186

Test T1880 at 2026-09-28

newspaper-page data-correction 18 en

anthropic claude-sonnet-5-5 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

```json
{
  "fixed_xml": "<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>",
  "number_of_fixes": 1,
  "explanation": "The opening tag <TITLE was missing its closing angle bracket, so the text 'Eine' was fused into the tag name (<TITLEEine ...>) and the element was not well-formed. I changed it to <TITLE>Eine Tossanische Bibel</TITLE>. The remaining structure (ITEM > BIBL > TITLE) was already correctly nested and closed, so no other tags were added or changed. I left the wording 'Tossanische' as it is, although it may be a historical spelling or a typo for 'Toskanische' (Tuscan). Please check it against the source."
}
```

Scoring

Fuzzy score97.3%
Items50

Speed

Model time, all inputs785.5 s
Mean per input15.71 s
Slowest input27.45 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens47.5K
Output tokens115.4K
Total tokens162.9K
Input cost$0.0949
Output cost$1.15
Total cost$1.25

Priced from 2026-09-28 · Input $/M $2.00 · Output $/M $10.00

Result 5 of 186

Test T1820 at 2026-09-27

newspaper-page data-correction 18 en

anthropic claude-opus-5-5 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

```json
{
  "fixed_xml": "<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>",
  "number_of_fixes": 1,
  "explanation": "The opening <TITLE> tag was missing its closing angle bracket ('>'). As written, '<TITLEEine' would be read as a malformed element name, so the document was not well-formed. Adding the '>' separates the tag name from the text 'Eine Tossanische Bibel'. The rest of the structure is correct: ITEM contains the item number '37.' followed by a BIBL element, which wraps the TITLE, and every element is properly nested and closed. No further tags were added. The entry is a short title, 'A Tossanus Bible', which likely refers to the annotated Luther Bible associated with the Reformed theologian Paul Tossanus (early 17th century). It contains no separate author, date or place of publication that could be tagged, so extra markup would not be justified."
}
```

Scoring

Fuzzy score98.6%
Items50

Speed

Model time, all inputs916.5 s
Mean per input18.33 s
Slowest input44.90 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens47.5K
Output tokens101.7K
Total tokens149.1K
Input cost$0.1899
Output cost$2.03
Total cost$2.22

Priced from 2026-09-27 · Input $/M $4.00 · Output $/M $20.00

Result 6 of 186

Test T1850 at 2026-09-27

newspaper-page data-correction 18 en

openai gpt-6-luna · temp 1.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Added the missing opening angle bracket delimiter in the TITLE tag."}

Scoring

Fuzzy score98.3%
Items50

Speed

Model time, all inputs395.4 s
Mean per input7.91 s
Slowest input17.16 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens33.6K
Output tokens50.8K
Total tokens84.4K
Input cost$0.0034
Output cost$0.0254
Total cost$0.0288

Priced from 2026-09-27 · Input $/M $0.10 · Output $/M $0.50

Result 7 of 186

Test T1865 at 2026-09-27

newspaper-page data-correction 18 en

x-ai grok-4.7 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"The opening TITLE tag was missing its closing angle bracket, so '<TITLEEine' was corrected to '<TITLE>Eine'."}

Scoring

Fuzzy score96.7%
Items50

Speed

Model time, all inputs12986.1 s
Mean per input259.72 s
Slowest input590.51 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens99.1K
Output tokens27.7K
Total tokens126.8K
Input cost$0.1983
Output cost$0.166
Total cost$5.92
Reasoning cost usd$5.56
Total reasoning tokens926.7K

Priced from 2026-09-27 · Input $/M $2.00 · Output $/M $6.00

Result 8 of 186

Test T1835 at 2026-09-27

newspaper-page data-correction 18 en

openai gpt-6-sol · temp 1.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Added the missing > to the opening TITLE tag."}

Scoring

Fuzzy score91.6%
Items50

Speed

Model time, all inputs451.9 s
Mean per input9.04 s
Slowest input16.66 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens33.6K
Output tokens37.6K
Total tokens71.2K
Input cost$0.0672
Output cost$0.376
Total cost$0.4432

Priced from 2026-09-27 · Input $/M $2.00 · Output $/M $10.00

Result 9 of 186

Test T0478 at 2026-09-18

newspaper-page data-correction 18 en

openrouter meta-llama/llama-4-maverick · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score82.1%

{"fixed_xml":"<ITEMS><ITEM><TEXT>37.</TEXT><BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM></ITEMS>","number_of_corrections":3,"explanation":"The provided XML string is mostly well-formed but lacks a proper root element and has some formatting issues. The 'TITLE' element contains text that starts with a capital letter but is not properly separated from the preceding element. I've added a root element 'ITEMS' to encapsulate the 'ITEM' element. I've also added a 'TEXT' element around the text '37.' to structure it properly within the 'ITEM' element. The 'TITLE' element's content has been adjusted to be properly formatted."}

Scoring

Fuzzy score92.0%
Items50

Speed

Model time, all inputs1775.4 s
Mean per input35.51 s
Slowest input83.89 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens28.2K
Output tokens27.9K
Total tokens56.2K
Input cost$0.0053
Output cost$0.0182
Total cost$0.0235

Priced from 2026-09-16 · Input $/M $0.1875 · Output $/M $0.6525

Result 10 of 186

Test T1730 at 2026-09-18

newspaper-page data-correction 18 en

openrouter qwen/qwen3.8-flash · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Added the missing closing angle bracket '>' to the opening <TITLE> tag so that the XML is well-formed."}

Scoring

Fuzzy score97.5%
Items41

Speed

Model time, all inputs12282.6 s
Mean per input250.66 s
Slowest input537.33 s
Inputs timed49

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens34.9K
Output tokens513.2K
Total tokens548.1K
Input cost$0.004
Output cost$0.2412
Total cost$0.2452

Priced from 2026-09-18 · Input $/M $0.15 · Output $/M $0.47