RISE Humanities Data Benchmark, 0.6.0-pre1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'book_advert_xml__true' with Search Hidden 'False' returned 179 results, showing page 8 of 18.
Result 71 of 179

Test T1655 at 2026-09-03

newspaper-page data-correction 18 en

genai gemini-3.8-flash · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Added the missing closing angle bracket to the opening <TITLE> tag."}

Scoring

Fuzzy score98.0%
Items50

Speed

Model time, all inputs1462.8 s
Mean per input29.26 s
Slowest input121.88 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens31.5K
Output tokens31K
Total tokens62.5K
Input cost$0.0237
Output cost$0.1161
Total cost$1.69
Reasoning cost usd$1.55
Total reasoning tokens412.7K

Priced from 2026-09-02 · Input $/M $0.75 · Output $/M $3.75

Result 72 of 179

Test T1610 at 2026-08-18

newspaper-page data-correction 18 en

x-ai grok-4.6 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Inserted the missing '>' to properly close the opening TITLE tag."}

Scoring

Fuzzy score98.4%
Items50

Speed

Model time, all inputs5389.7 s
Mean per input107.79 s
Slowest input201.89 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens47.3K
Output tokens27.2K
Total tokens74.5K
Input cost$0.0946
Output cost$0.1631
Total cost$1.95
Reasoning cost usd$1.69
Total reasoning tokens281.3K

Priced from 2026-08-18 · Input $/M $2.00 · Output $/M $6.00

Result 73 of 179

Test T1595 at 2026-08-18

newspaper-page data-correction 18 en

genai gemini-3.7-flash · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Fixed the malformed opening XML tag '<TITLEEine Tossanische Bibel' to '<TITLE>Eine Tossanische Bibel'."}

Scoring

Fuzzy score97.7%
Items50

Speed

Model time, all inputs379.3 s
Mean per input7.59 s
Slowest input16.47 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens31.5K
Output tokens30.2K
Total tokens61.8K
Input cost$0.0237
Output cost$0.1134
Total cost$0.4777
Reasoning cost usd$0.3407
Total reasoning tokens90.9K

Priced from 2026-08-18 · Input $/M $0.75 · Output $/M $3.75

Result 74 of 179

Test T1625 at 2026-08-18

newspaper-page data-correction 18 en

openrouter meta/muse-spark-1.2 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Fixed the malformed opening TITLE tag: '<TITLEEine' was missing the closing '>' and a space. Corrected to '<TITLE>' to make the XML well-formed and properly nested as ITEM > BIBL > TITLE. The text content '37.' was left as mixed content inside ITEM as it is valid XML."}

Scoring

Fuzzy score96.2%
Items50

Speed

Model time, all inputs1161.7 s
Mean per input23.23 s
Slowest input41.78 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens28.1K
Output tokens184.1K
Total tokens212.1K
Input cost$0.0351
Output cost$0.7822
Total cost$0.8174

Priced from 2026-08-18 · Input $/M $1.25 · Output $/M $4.25

Result 75 of 179

Test T1640 at 2026-08-18

newspaper-page data-correction 18 en

openrouter z-ai/glm-5v-turbo · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Fixed malformed opening tag <TITLE> which was missing the closing angle bracket (>). The original had <TITLEEine Tossanische Bibel instead of <TITLE>Eine Tossanische Bibel."}

Scoring

Fuzzy score98.4%
Items48

Speed

Model time, all inputs4219.6 s
Mean per input84.39 s
Slowest input175.24 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens40.2K
Output tokens71.5K
Total tokens111.7K
Input cost$0.0483
Output cost$0.286
Total cost$0.3342

Priced from 2026-08-18 · Input $/M $1.20 · Output $/M $4.00

Result 76 of 179

Test T1550 at 2026-08-16

newspaper-page data-correction 18 en

huggingface thinkingmachines/Inkling-Small · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"The opening <TITLE> tag was malformed because the closing angle bracket (>) was missing after the tag name, causing the tag to merge with the title text. Inserted the missing > to properly delimit the element. No additional structural tags were needed as the item number and bibliographic wrapper are already correctly placed within <ITEM> and <BIBL>."}

Scoring

Fuzzy score93.4%
Items50

Speed

Model time, all inputs3271.2 s
Mean per input65.42 s
Slowest input206.36 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens28.3K
Output tokens360.6K
Total tokens388.9K
Input cost$0.0127
Output cost$0.4327
Total cost$0.4455

Priced from 2026-08-15 · Input $/M $0.45 · Output $/M $1.20

Result 77 of 179

Test T1580 at 2026-08-16

newspaper-page data-correction 18 en

huggingface meta-models/Muse-Glimmer-30B · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"The opening TITLE tag was malformed as <TITLEEine ...>. The closing '>' for the opening tag was missing after TITLE, causing the tag name and content to merge. Fixed to <TITLE>Eine Tossanische Bibel</TITLE>."}

Scoring

Fuzzy score95.8%
Items50

Speed

Model time, all inputs3301.3 s
Mean per input66.03 s
Slowest input230.66 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens28.9K
Output tokens301.3K
Total tokens330.2K
Input cost$0.0101
Output cost$0.4519
Total cost$0.462

Priced from 2026-08-15 · Input $/M $0.35 · Output $/M $1.50

Result 78 of 179

Test T1520 at 2026-08-15

newspaper-page data-correction 18 en

huggingface Qwen/Qwen3-VL-235B-A22B-Instruct · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Added missing closing angle bracket '>' after 'TITLE' to properly close the opening tag. The original tag '<TITLEEine Tossanische Bibel' was malformed and needed to be corrected to '<TITLE>Eine Tossanische Bibel'."}

Scoring

Fuzzy score97.7%
Items50

Speed

Model time, all inputs1406.5 s
Mean per input28.13 s
Slowest input57.39 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens32.1K
Output tokens31.5K
Total tokens63.6K
Input cost$0.0064
Output cost$0.0277
Total cost$0.0341

Priced from 2026-08-15 · Input $/M $0.20 · Output $/M $0.88

Result 79 of 179

Test T1490 at 2026-08-15

newspaper-page data-correction 18 en

huggingface swiss-ai/apertus-v1.5-70b · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score90.8%

{"fixed_xml":"<ROOT><ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM></ROOT>","number_of_corrections":1,"explanation":"The original XML string was not well-formed. It lacked a root element and proper XML syntax. I added a root element `<ROOT>` to make it a valid XML document."}

Scoring

Fuzzy score83.3%
Items50

Speed

Model time, all inputs615.2 s
Mean per input12.30 s
Slowest input191.69 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens33.3K
Output tokens27.4K
Total tokens60.7K
Input cost$0.0273
Output cost$0.0799
Total cost$0.1073

Priced from 2026-08-10 · Input $/M $0.82 · Output $/M $2.92

Result 80 of 179

Test T1565 at 2026-08-15

newspaper-page data-correction 18 en

huggingface thinkingmachines/Inkling · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Score0.0%

The user wants me to fix an XML string. Let's look at the input:

`<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>`

This looks like valid XML already. The tags are properly nested and closed:
- `<ITEM>` opens and closes
- `<BIBL>` opens and closes
- `<TITLE>` opens and closes

Wait, is there anything faulty? Maybe missing XML declaration? Or maybe the content should have some structure? The prompt says "Add xml-tags if faulty where it makes sense."

Let's check if there's any implicit issue. The string is:
`<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>`

This is well-formed XML. However, maybe the number "37." should be in a specific tag? Or maybe there's a missing root element? But the prompt asks to fix faulty XML and add tags where it makes sense.

Actually, looking at the schema requested in the response: `number_of_corrections` (not `number_of_fixes` as mentioned in the prompt text, but the JSON schema says `number_of_corrections`). Wait, the user said: "Use the keys 'fixed_xml', 'number_of_fixes', 'explanation'." But then they provide a schema that requires `fixed_xml`, `number_of_corrections`, and `explanation`. I should follow the schema provided in the JSON example, which uses `number_of_corrections`. Actually, let me read carefully.

The user says:
"Format your response as JSON. Use the keys 'fixed_xml', 'number_of_fixes', 'explanation'."

Then they provide:
"You MUST respond with valid JSON matching this exact schema: {"description": "Main output structure...", "properties": {"fixed_xml": ..., "number_of_corrections": ..., "explanation": ...}, "required": ["fixed_xml", "number_of_corrections"], ...}"

There's a conflict: user asks for `number_of_fixes` but schema requires `number_of_corrections`. I should probably use `number_of_corrections` since they said "matching this exact schema" and it's required. Or maybe include both? No, the schema is explicit. I'll use `number_of_corrections` but maybe I can mention? Actually, to be safe and match the exact schema, I must use `number_of_corrections`. The prompt text might have a typo.

Now, about the XML fix. Is there actually anything to fix? Let's examine the XML content:
`<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>`

It seems correct. But maybe the issue is that there's no XML declaration (`<?xml version="1.0"?>`)? Or maybe the text "37." should be wrapped in a tag like `<NO>` or `<NUMBER>`? Without context, it's hard to tell.

Wait, maybe the original XML is missing closing tags? No, they are all there.

Let's reconsider. The file is named `line_1.json` and contains an `xml_string`. The content is an advert for a book. Maybe the structure should be:
`<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>`

Perhaps the fault is that `<BIBL>` should contain more, or `<ITEM>` should have an attribute? Or maybe "37." is not part of the XML and should be outside? No, it's inside.

Actually, maybe the XML is missing a root element if this is a fragment? But it's presented as the xml_string.

Given the instruction "Add xml-tags if faulty where it makes sense", and since the XML is technically well-formed, maybe I should just return it as is with 0 corrections? But that seems unlikely; usually these tasks have a hidden error.

Let's look very closely:
`<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>`

Wait, is there a missing space or something? No.

Could it be that `<BIBL>` is misspelled? Maybe it should be `<BIBLIO>`? Or `<BIBL>` is correct?

Maybe the issue is that the text inside `<TITLE>` contains special characters that need escaping? "Eine Tossanische Bibel" - no special XML chars.

Perhaps the number "37." should be in its own element, e.g., `<NUMBER>37.</NUMBER>`? The prompt says "Add xml-tags if faulty where it makes sense." If the original has "37." as bare text inside `<ITEM>`, maybe it should be tagged.

But I don't have a schema for this specific domain. However, looking at the context "Book Advert XML files from Avisblatt" (from the schema description), maybe there's a standard structure.

Actually, maybe the XML is missing the closing `</ITEM>`? No, it's there.

Wait, I see: `<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>` - this is perfectly valid.

But maybe the problem is that there is no root element wrapping the `<ITEM>`? If this is the entire file, it's okay for a fragment, but usually XML documents need a single root. However, the string is named `xml_string` inside a JSON file, so it might just be a fragment.

Given that I need to provide a fix, and there might be 0 fixes, but that feels like failing the task. Let's think about common XML errors in these types of prompts. Often they have unclosed tags, mismatched tags, missing tags, or text outside tags.

Here, everything is closed. Could there be an invisible error? Let's check character by character:
`<` `I` `T` `E` `M` `>` `3` `7` `.` ` ` `<` `B` `I` `B` `L` `>` `<` `T` `I` `T` `L` `E` `>` `E` `i` `n` `e` ` ` `T` `o` `s` `s` `a` `n` `i` `s` `c` `h` `e` ` ` `B` `i` `b` `e` `l` `<` `/` `T` `I` `T` `L` `E` `>` `<` `/` `B` `I` `B` `L` `>` `<` `/` `I` `T` `E` `M` `>`

All good.

Wait, maybe the issue is that `<BIBL>` should be `<BIBLIOGRAPHY>`? Or maybe `<ITEM>` should be `<ITEM NO="37">`?

Actually, in many historical bibliographic XML formats, the number is an attribute: `<ITEM NO="37">`. That would make sense: "37." is the item number, and it should be an attribute rather than text content. So the fix could be:
`<ITEM NO="37"> <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL> </ITEM>`

But is that "fixing faulty XML"? The original is valid XML, but semantically the number might be misplaced. The prompt says "Add xml-tags if faulty where it makes sense." This could mean restructuring.

Alternatively, maybe the bare text "37. " should be wrapped in a `<NO>` tag: `<ITEM><NO>37.</NO> <BIBL>...</BIBL></ITEM>`.

Given the ambiguity, I think the safest answer is that the XML is well-formed and requires no corrections, so `number_of_corrections` is 0, and `fixed_xml` is the same as input. But I worry that the task expects some change.

Let's re-read the prompt: "Fix this xml. Add xml-tags if faulty where it makes sense."

If it's not faulty, then no fixes are needed. I will state that it's already well-formed.

But wait, there's a subtle thing: in the JSON input, the value is:
`"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>"`

Could there be a missing closing quote or something in the JSON? No, that's the input to me.

Actually, I just noticed: the user might have intended for there to be an error like `<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL>` (missing `</ITEM>`). But in the text provided, `</ITEM>` is present.

Let me copy exactly what's in the prompt:
`<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>`

Yes, it's there.

Given all this, I will return the XML unchanged, with 0 corrections, and explain that it is already well-formed.

However, I need to match the schema exactly. The schema requires:
- `fixed_xml` (string)
- `number_of_corrections` (integer)
- `explanation` (optional, but I'll include it)

The user asked for keys `fixed_xml`, `number_of_fixes`, `explanation`. But the schema says `number_of_corrections`. Since they said "You MUST respond with valid JSON matching this exact schema", I will use `number_of_corrections`. I will not include `number_of_fixes` because that would violate the schema (extra properties are usually okay, but they said

Scoring

Fuzzy score31.3%
Items50

Speed

Model time, all inputs1702.8 s
Mean per input34.06 s
Slowest input45.41 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens34.7K
Output tokens93.9K
Total tokens128.6K
Input cost$0.0347
Output cost$0.3804
Total cost$0.415

Priced from 2026-08-15 · Input $/M $1.00 · Output $/M $4.05