RISE Humanities Data Benchmark, 0.6.1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'book_advert_xml__true' with Search Hidden 'False' returned 187 results, showing page 19 of 19.
Result 181 of 187

Test T0477 at 2025-12-08

newspaper-page data-correction 18 en

openrouter qwen/qwen3-vl-8b-thinking · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Score0.0%

Scoring

Fuzzy score1.9%
Items50

Speed

Model time, all inputs457.3 s
Mean per input9.15 s
Slowest input21.86 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens21K
Output tokens23.2K
Total tokens44.2K
Input cost$0.0038
Output cost$0.0488
Total cost$0.0525

Priced from 2025-11-24 · Input $/M $0.18 · Output $/M $2.10

Result 182 of 187

Test T0449 at 2025-12-08

newspaper-page data-correction 18 en

openai gpt-4.1-mini-2025-04-14 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"The closing angle bracket '>' was missing after the opening <TITLE> tag. Added it to correct the XML syntax."}

Scoring

Fuzzy score90.1%
Items50

Speed

Model time, all inputs1570.0 s
Mean per input31.40 s
Slowest input628.08 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens33.9K
Output tokens61.1K
Total tokens95K
Input cost$0.0136
Output cost$0.0977
Total cost$0.1113

Priced from 2025-11-24 · Input $/M $0.40 · Output $/M $1.60

Result 183 of 187

Test T0455 at 2025-12-08

newspaper-page data-correction 18 en

openai o3 · temp 1.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Score0.0%

Scoring

Fuzzy score0.0%
Items50

Speed

Model time, all inputs454.7 s
Mean per input9.09 s
Slowest input12.09 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens0
Output tokens0
Total tokens0
Input cost$0.00
Output cost$0.00
Total cost$0.00

Priced from 2025-11-24 · Input $/M $2.00 · Output $/M $8.00

Result 184 of 187

Test T0459 at 2025-12-08

newspaper-page data-correction 18 en

genai gemini-2.5-flash · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"The opening <TITLE tag was missing its closing angle bracket '>'. This has been added to form a well-formed tag."}

Scoring

Fuzzy score90.8%
Items50

Speed

Model time, all inputs1536.3 s
Mean per input30.73 s
Slowest input267.82 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens31.5K
Output tokens74.8K
Total tokens106.3K
Input cost$0.0095
Output cost$0.1869
Total cost$0.7558
Reasoning cost usd$0.5595
Total reasoning tokens223.8K

Priced from 2025-11-24 · Input $/M $0.30 · Output $/M $2.50

Result 185 of 187

Test T0479 at 2025-12-08

newspaper-page data-correction 18 en

openrouter qwen/qwen3-vl-30b-a3b-instruct · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":2,"explanation":"The original XML was missing a space after the opening <TITLE> tag and had a missing closing quote after 'TITLE'. The corrected version adds the space and ensures proper tag closure."}

Scoring

Fuzzy score96.4%
Items50

Speed

Model time, all inputs626.1 s
Mean per input12.52 s
Slowest input71.20 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens34.5K
Output tokens37.8K
Total tokens72.3K
Input cost$0.0069
Output cost$0.0227
Total cost$0.0296

Priced from 2025-11-24 · Input $/M $0.20 · Output $/M $0.60

Result 186 of 187

Test T0450 at 2025-12-08

newspaper-page data-correction 18 en

openai gpt-4.1-nano-2025-04-14 · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"Added missing '>' after <TITLE> tag to properly close the tag."}

Scoring

Fuzzy score93.4%
Items50

Speed

Model time, all inputs315.4 s
Mean per input6.31 s
Slowest input15.09 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens33.8K
Output tokens25.4K
Total tokens59.1K
Input cost$0.0034
Output cost$0.0101
Total cost$0.0135

Priced from 2025-11-24 · Input $/M $0.10 · Output $/M $0.40

Result 187 of 187

Test T0458 at 2025-12-08

newspaper-page data-correction 18 en

genai gemini-2.5-pro · temp 0.0 · dataclass CorrectedAdvert

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score100.0%

{"fixed_xml":"<ITEM>37. <BIBL><TITLE>Eine Tossanische Bibel</TITLE></BIBL></ITEM>","number_of_corrections":1,"explanation":"The opening tag for TITLE was malformed. The text content 'Eine Tossanische Bibel' was incorrectly included within the tag name. The tag has been corrected to <TITLE> and the text placed as its content."}

Scoring

Fuzzy score91.5%
Items50

Speed

Model time, all inputs2918.1 s
Mean per input58.36 s
Slowest input112.24 s
Inputs timed50

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens31.5K
Output tokens28.6K
Total tokens60.1K
Input cost$0.0394
Output cost$0.2857
Total cost$3.99
Reasoning cost usd$3.66
Total reasoning tokens366.2K

Priced from 2025-11-24 · Input $/M $1.25 · Output $/M $10.00