A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.
A test run includes:
Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.
newspaper-page data-correction 18 en
anthropic claude-opus-5-5 · temp 0.0 · dataclass CorrectedAdvert
One model response from this run. The per-input comparison needs JavaScript.
{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"The opening <TITLE> tag was missing its closing '>' (it read '<TITLEEine'). I added the bracket so the tag is well-formed and the title text 'Eine Tossanische Bibel' sits inside it. No other tags were needed because the existing ITEM/BIBL/TITLE structure is correct."}
newspaper-page data-correction 18 en
anthropic claude-sonnet-5-5 · temp 0.0 · dataclass CorrectedAdvert
One model response from this run. The per-input comparison needs JavaScript.
{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"The opening tag <TITLE was missing its closing angle bracket. I changed it to <TITLE>. The rest of the structure (ITEM, BIBL, TITLE) was already well-formed. I left the text \"Tossanische\" as it appears in the original, although it may be a transcription of \"Toscanische\"."}
newspaper-page data-correction 18 en
anthropic claude-fable-5-1 · temp 0.0 · dataclass CorrectedAdvert
One model response from this run. The per-input comparison needs JavaScript.
{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"The opening TITLE tag was malformed: \"<TITLEEine Tossanische Bibel</TITLE>\" lacked the closing \">\" after the tag name, so the text was absorbed into the element name and the closing </TITLE> had no matching opener. Inserted the missing \">\" to yield \"<TITLE>Eine Tossanische Bibel</TITLE>\". The surrounding ITEM and BIBL structure was already well-formed and left unchanged."}
newspaper-page data-correction 18 en
anthropic claude-sonnet-5-5 · temp 0.0 · dataclass CorrectedAdvert
One model response from this run. The per-input comparison needs JavaScript.
```json
{
"fixed_xml": "<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>",
"number_of_fixes": 1,
"explanation": "The opening tag <TITLE was missing its closing angle bracket, so the text 'Eine' was fused into the tag name (<TITLEEine ...>) and the element was not well-formed. I changed it to <TITLE>Eine Tossanische Bibel</TITLE>. The remaining structure (ITEM > BIBL > TITLE) was already correctly nested and closed, so no other tags were added or changed. I left the wording 'Tossanische' as it is, although it may be a historical spelling or a typo for 'Toskanische' (Tuscan). Please check it against the source."
}
```
newspaper-page data-correction 18 en
anthropic claude-opus-5-5 · temp 0.0 · dataclass CorrectedAdvert
One model response from this run. The per-input comparison needs JavaScript.
```json
{
"fixed_xml": "<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>",
"number_of_fixes": 1,
"explanation": "The opening <TITLE> tag was missing its closing angle bracket ('>'). As written, '<TITLEEine' would be read as a malformed element name, so the document was not well-formed. Adding the '>' separates the tag name from the text 'Eine Tossanische Bibel'. The rest of the structure is correct: ITEM contains the item number '37.' followed by a BIBL element, which wraps the TITLE, and every element is properly nested and closed. No further tags were added. The entry is a short title, 'A Tossanus Bible', which likely refers to the annotated Luther Bible associated with the Reformed theologian Paul Tossanus (early 17th century). It contains no separate author, date or place of publication that could be tagged, so extra markup would not be justified."
}
```
newspaper-page data-correction 18 en
openai gpt-6-luna · temp 1.0 · dataclass CorrectedAdvert
One model response from this run. The per-input comparison needs JavaScript.
{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Added the missing opening angle bracket delimiter in the TITLE tag."}
newspaper-page data-correction 18 en
x-ai grok-4.7 · temp 0.0 · dataclass CorrectedAdvert
One model response from this run. The per-input comparison needs JavaScript.
{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"The opening TITLE tag was missing its closing angle bracket, so '<TITLEEine' was corrected to '<TITLE>Eine'."}
newspaper-page data-correction 18 en
openai gpt-6-sol · temp 1.0 · dataclass CorrectedAdvert
One model response from this run. The per-input comparison needs JavaScript.
{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Added the missing > to the opening TITLE tag."}
newspaper-page data-correction 18 en
openrouter meta-llama/llama-4-maverick · temp 0.0 · dataclass CorrectedAdvert
One model response from this run. The per-input comparison needs JavaScript.
{"fixed_xml":"<ITEMS><ITEM><TEXT>37.</TEXT><BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM></ITEMS>","number_of_corrections":3,"explanation":"The provided XML string is mostly well-formed but lacks a proper root element and has some formatting issues. The 'TITLE' element contains text that starts with a capital letter but is not properly separated from the preceding element. I've added a root element 'ITEMS' to encapsulate the 'ITEM' element. I've also added a 'TEXT' element around the text '37.' to structure it properly within the 'ITEM' element. The 'TITLE' element's content has been adjusted to be properly formatted."}
newspaper-page data-correction 18 en
openrouter qwen/qwen3.8-flash · temp 0.0 · dataclass CorrectedAdvert
One model response from this run. The per-input comparison needs JavaScript.
{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Added the missing closing angle bracket '>' to the opening <TITLE> tag so that the XML is well-formed."}