RISE Humanities Data Benchmark, 0.6.0-pre1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'book_advert_xml__true' with Search Hidden 'False' returned 179 results, showing page 6 of 18.
Result 51 of 179

Test T1595 at 2026-09-16

newspaper-page data-correction 18 en

genai gemini-3.7-flash · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Fixed malformed opening tag '<TITLEEine' to properly formed '<TITLE>Eine'."}

Scoring

Fuzzy score97.9%
Items50

Speed

Model time, all inputs432.8 s
Mean per input8.66 s
Slowest input19.64 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens31.5K
Output tokens30K
Total tokens61.5K
Input cost$0.0237
Output cost$0.1124
Total cost$0.4767
Reasoning cost usd$0.3407
Total reasoning tokens90.8K

Priced from 2026-08-18 · Input $/M $0.75 · Output $/M $3.75

Result 52 of 179

Test T1047 at 2026-09-16

newspaper-page data-correction 18 en

openai gpt-5.5-2026-04-23 · temp 1.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Corrected the malformed opening TITLE tag from <TITLEEine to <TITLE>."}

Scoring

Fuzzy score98.2%
Items50

Speed

Model time, all inputs748.0 s
Mean per input14.96 s
Slowest input43.18 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens33.6K
Output tokens65.2K
Total tokens98.8K
Input cost$0.1681
Output cost$1.96
Total cost$2.12

Priced from 2026-09-16 · Input $/M $5.00 · Output $/M $30.00

Result 53 of 179

Test T1036 at 2026-09-16

newspaper-page data-correction 18 en

deepseek deepseek-v4-pro · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"The opening tag <TITLEEine was missing the > delimiter, so it was corrected to <TITLE>."}

Scoring

Fuzzy score98.1%
Items49

Speed

Model time, all inputs6674.6 s
Mean per input133.49 s
Slowest input302.73 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens43.4K
Output tokens477.9K
Total tokens521.2K
Input cost$0.0572
Output cost$1.89
Total cost$1.95

Priced from 2026-09-16 · Input $/M $1.32 · Output $/M $3.96

Result 54 of 179

Test T0728 at 2026-09-16

newspaper-page data-correction 18 en

x-ai grok-4.20-0309-reasoning · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Added missing '>' character after the opening '<TITLE' tag to make the XML well-formed."}

Scoring

Fuzzy score97.7%
Items50

Speed

Model time, all inputs2344.1 s
Mean per input46.88 s
Slowest input85.21 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens43.5K
Output tokens27K
Total tokens70.5K
Input cost$0.0543
Output cost$0.0676
Total cost$0.723
Reasoning cost usd$0.601
Total reasoning tokens240.4K

Priced from 2026-09-16 · Input $/M $1.25 · Output $/M $2.50

Result 55 of 179

Test T1445 at 2026-09-16

newspaper-page data-correction 18 en

genai gemini-3.5-flash-lite · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score88.3%

{"fixed_xml":"<ITEM><NUMBER>37.</NUMBER> <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":2,"explanation":"Added closing tag for TITLE and introduced a NUMBER tag to properly structure the item number."}

Scoring

Fuzzy score95.9%
Items50

Speed

Model time, all inputs101.3 s
Mean per input2.03 s
Slowest input3.40 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens31.5K
Output tokens26.2K
Total tokens57.7K
Input cost$0.0095
Output cost$0.0654
Total cost$0.0749

Priced from 2026-09-16 · Input $/M $0.30 · Output $/M $2.50

Result 56 of 179

Test T0520 at 2026-09-16

newspaper-page data-correction 18 en

anthropic claude-opus-4-5-20251101 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Fixed the opening TITLE tag which was missing the closing angle bracket. Changed '<TITLEEine' to '<TITLE>Eine'."}

Scoring

Fuzzy score97.0%
Items50

Speed

Model time, all inputs433.3 s
Mean per input8.67 s
Slowest input18.40 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens74.1K
Output tokens35.4K
Total tokens109.5K
Input cost$0.3705
Output cost$0.8855
Total cost$1.26

Priced from 2026-09-16 · Input $/M $5.00 · Output $/M $25.00

Result 57 of 179

Test T1099 at 2026-09-16

newspaper-page data-correction 18 en

mistral mistral-medium-3.5 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Added missing closing angle bracket (>) after the opening <TITLE> tag."}

Scoring

Fuzzy score96.1%
Items50

Speed

Model time, all inputs271.0 s
Mean per input5.42 s
Slowest input10.85 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens40.3K
Output tokens31.6K
Total tokens72K
Input cost$0.0605
Output cost$0.2373
Total cost$0.2977

Priced from 2026-09-16 · Input $/M $1.50 · Output $/M $7.50

Result 58 of 179

Test T1460 at 2026-09-16

newspaper-page data-correction 18 en

x-ai grok-4.5 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Inserted missing '>' to close the opening TITLE tag."}

Scoring

Fuzzy score96.7%
Items50

Speed

Model time, all inputs2552.0 s
Mean per input51.04 s
Slowest input89.24 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens61.7K
Output tokens26.5K
Total tokens88.2K
Input cost$0.1234
Output cost$0.1591
Total cost$1.14
Reasoning cost usd$0.8612
Total reasoning tokens143.5K

Priced from 2026-09-16 · Input $/M $2.00 · Output $/M $6.00

Result 59 of 179

Test T0498 at 2026-09-16

newspaper-page data-correction 18 en

openai gpt-5.2-2025-12-11 · temp 1.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Added missing opening tag delimiter for the TITLE element: changed <TITLEEine Tossanische Bibel</TITLE> to <TITLE>Eine Tossanische Bibel</TITLE>."}

Scoring

Fuzzy score96.6%
Items50

Speed

Model time, all inputs727.5 s
Mean per input14.55 s
Slowest input332.83 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens33.7K
Output tokens29K
Total tokens62.7K
Input cost$0.0589
Output cost$0.4066
Total cost$0.4655

Priced from 2026-09-16 · Input $/M $1.75 · Output $/M $14.00

Result 60 of 179

Test T0638 at 2026-09-16

newspaper-page data-correction 18 en

anthropic claude-opus-4-6 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"The opening TITLE tag was missing the closing angle bracket: `<TITLEEine` was corrected to `<TITLE>Eine`. This was a malformed opening tag where the tag name ran into the text content."}

Scoring

Fuzzy score97.4%
Items50

Speed

Model time, all inputs621.4 s
Mean per input12.43 s
Slowest input25.75 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens74.1K
Output tokens40.5K
Total tokens114.6K
Input cost$0.3707
Output cost$1.01
Total cost$1.38

Priced from 2026-09-16 · Input $/M $5.00 · Output $/M $25.00