RISE Humanities Data Benchmark, 0.6.0-pre1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'book_advert_xml__true' with Search Hidden 'False' returned 179 results, showing page 16 of 18.
Result 151 of 179

Test T0542 at 2026-01-25

newspaper-page data-correction 18 en

mistral mistral-large-2512 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score80.5%

{"fixed_xml":"<file name=\"line_1.json\"><ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM></file>","number_of_corrections":1,"explanation":"Added missing closing angle bracket '>' after the opening TITLE tag in the BIBL element."}

Scoring

Fuzzy score93.3%
Items50

Speed

Model time, all inputs660.8 s
Mean per input13.22 s
Slowest input26.00 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens40.3K
Output tokens35.5K
Total tokens75.8K
Input cost$0.0202
Output cost$0.0533
Total cost$0.0734

Priced from 2026-01-23 · Input $/M $0.50 · Output $/M $1.50

Result 152 of 179

Test T0480 at 2026-01-25

newspaper-page data-correction 18 en

openrouter qwen/qwen3-vl-8b-instruct · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":2,"explanation":"Added missing XML tags to properly structure the title and correct the malformed tag sequence."}

Scoring

Fuzzy score72.4%
Items50

Speed

Model time, all inputs384.9 s
Mean per input7.70 s
Slowest input25.96 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens35.3K
Output tokens23.7K
Total tokens59.1K
Input cost$0.0028
Output cost$0.0119
Total cost$0.0147

Priced from 2026-01-23 · Input $/M $0.08 · Output $/M $0.50

Result 153 of 179

Test T0452 at 2026-01-25

newspaper-page data-correction 18 en

openai gpt-5-mini-2025-08-07 · temp 1.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Fixed a malformed opening TITLE tag (was '<TITLEEine ...'); added the missing '>' so it reads '<TITLE>Eine Tossanische Bibel</TITLE>'. No other changes made."}

Scoring

Fuzzy score91.8%
Items50

Speed

Model time, all inputs2637.1 s
Mean per input52.74 s
Slowest input122.36 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens33.6K
Output tokens189.5K
Total tokens223.1K
Input cost$0.0084
Output cost$0.379
Total cost$0.3874

Priced from 2026-01-23 · Input $/M $0.25 · Output $/M $2.00

Result 154 of 179

Test T0520 at 2026-01-24

newspaper-page data-correction 18 en

anthropic claude-opus-4-5-20251101 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Fixed the opening TITLE tag which was missing the closing angle bracket. Changed '<TITLEEine' to '<TITLE>Eine'."}

Scoring

Fuzzy score97.0%
Items50

Speed

Model time, all inputs465.5 s
Mean per input9.31 s
Slowest input16.87 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens74.1K
Output tokens35.8K
Total tokens109.9K
Input cost$0.3705
Output cost$0.895
Total cost$1.27

Priced from 2026-01-23 · Input $/M $5.00 · Output $/M $25.00

Result 155 of 179

Test T0460 at 2026-01-24

newspaper-page data-correction 18 en

genai gemini-2.5-flash-lite · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score97.6%

{"fixed_xml":"<ITEM><BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":2,"explanation":"Removed the number '37.' from the beginning of the ITEM tag as it appears to be extraneous data. Also, enclosed the title 'Eine Tossanische Bibel' within the TITLE tags correctly."}

Scoring

Fuzzy score61.8%
Items50

Speed

Model time, all inputs2045.5 s
Mean per input40.91 s
Slowest input124.09 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens31.5K
Output tokens1.1M
Total tokens1.2M
Input cost$0.0032
Output cost$0.4592
Total cost$0.4624

Priced from 2026-01-23 · Input $/M $0.10 · Output $/M $0.40

Result 156 of 179

Test T0468 at 2026-01-24

newspaper-page data-correction 18 en

anthropic claude-sonnet-4-5-20250929 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Fixed missing closing angle bracket in the opening TITLE tag. Changed \"<TITLEEine\" to \"<TITLE>Eine\" to properly separate the tag from the content."}

Scoring

Fuzzy score97.4%
Items50

Speed

Model time, all inputs495.1 s
Mean per input9.90 s
Slowest input17.83 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens74.1K
Output tokens36.9K
Total tokens111K
Input cost$0.2223
Output cost$0.5535
Total cost$0.7758

Priced from 2026-01-23 · Input $/M $3.00 · Output $/M $15.00

Result 157 of 179

Test T0531 at 2026-01-24

newspaper-page data-correction 18 en

anthropic claude-haiku-4-5-20251001 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Fixed malformed TITLE tag. The opening tag was missing the '>' character, appearing as '<TITLEEine' instead of '<TITLE>Eine'. This has been corrected to properly open and close the TITLE element."}

Scoring

Fuzzy score95.3%
Items50

Speed

Model time, all inputs265.8 s
Mean per input5.32 s
Slowest input11.14 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens74.1K
Output tokens39.1K
Total tokens113.2K
Input cost$0.0741
Output cost$0.1955
Total cost$0.2696

Priced from 2026-01-23 · Input $/M $1.00 · Output $/M $5.00

Result 158 of 179

Test T0459 at 2026-01-24

newspaper-page data-correction 18 en

genai gemini-2.5-flash · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"The opening <TITLE> tag was malformed, missing the closing angle bracket after the tag name."}

Scoring

Fuzzy score72.4%
Items50

Speed

Model time, all inputs3484.4 s
Mean per input69.69 s
Slowest input246.00 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens31.5K
Output tokens727.4K
Total tokens758.9K
Input cost$0.0095
Output cost$1.82
Total cost$2.43
Reasoning cost usd$0.5979
Total reasoning tokens239.2K

Priced from 2026-01-23 · Input $/M $0.30 · Output $/M $2.50

Result 159 of 179

Test T0458 at 2026-01-24

newspaper-page data-correction 18 en

genai gemini-2.5-pro · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"The opening tag '<TITLEEine Tossanische Bibel>' was malformed. The tag name was merged with its text content. This was corrected by separating the tag into '<TITLE>' and the content 'Eine Tossanische Bibel'."}

Scoring

Fuzzy score92.2%
Items50

Speed

Model time, all inputs2201.5 s
Mean per input44.03 s
Slowest input105.13 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens31.5K
Output tokens29.2K
Total tokens60.8K
Input cost$0.0394
Output cost$0.2923
Total cost$3.24
Reasoning cost usd$2.91
Total reasoning tokens290.6K

Priced from 2026-01-23 · Input $/M $1.25 · Output $/M $10.00

Result 160 of 179

Test T0455 at 2025-12-09

newspaper-page data-correction 18 en

openai o3-2025-04-16 · temp 1.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Inserted the missing closing angle bracket (>) in the opening TITLE tag so that it reads <TITLE>."}

Scoring

Fuzzy score94.5%
Items50

Speed

Model time, all inputs1274.7 s
Mean per input25.49 s
Slowest input44.77 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens33.6K
Output tokens147.2K
Total tokens180.8K
Input cost$0.0672
Output cost$1.18
Total cost$1.24

Priced from 2025-11-24 · Input $/M $2.00 · Output $/M $8.00