RISE Humanities Data Benchmark, 0.6.1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'medieval_manuscripts__true' with Search Hidden 'False' returned 177 results, showing page 7 of 18.
Result 61 of 177

Test T1516 at 2026-09-16

manuscript transcription handwritten 15 de prose

huggingface Qwen/Qwen3-VL-235B-A22B-Instruct · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score93.7%
CER ↓6.2%

{"folios":[{"folio":"3","text":"Vnd ein pferit die wir vnd
minen knechten vber Rulfend
Gen So was meryn ein Weg
Genne den wir machten
vnd vielen die knechte drut
vnd vil in vntz an den arb
vnd die pferit vnz an die
pettel vnd was ze mot ein grosser
nebel & wir kum gepaffen
vnd als mit grosser arbeit kame
wir ze mitten tag zu sant
kristoffel vff den berg So
So fach ich die bucher So gar
vil herrin wopen in stone
die ir stur & gyn geben hand
So fand mineb vatterb flugen
wopen sich in dem smey","addition1":"","addition2":"","addition3":""}]}

Scoring

Fuzzy score72.0%
CER ↓28.8%

Speed

Model time, all inputs220.9 s
Mean per input18.41 s
Slowest input24.93 s
Inputs timed12

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens39.3K
Output tokens2.8K
Total tokens42K
Input cost$0.0079
Output cost$0.0024
Total cost$0.0103

Priced from 2026-09-16 · Input $/M $0.20 · Output $/M $0.88

Result 62 of 177

Test T1366 at 2026-09-16

manuscript transcription handwritten 15 de prose

openai gpt-5.6-terra · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score90.6%
CER ↓10.3%

{"folios":[{"folio":"3","text":"Vnd ein pfferit die mir vnd
minen knechten vber hulffend
Dan So was menchem hengweg
zůmū den mir machtend
vnd vrlends die knechte do
vnd vil in vns angieng anß
vnd die pfferit vns an die
stet vnd vns ze mol ein grosse
nebel in kam gespānd
vnd als mit grossen arbeit kamē
wir ze mitten tag zu Sant-
Ruffel vff den berg do
do hub ich die bücher do gar
vil herinn wapen zu hand
die in iren sachen gelegen hand
do stund mines vatters slagen
wapen ab in dem einen","addition1":"","addition2":"","addition3":""}]}

Scoring

Fuzzy score69.6%
CER ↓33.3%

Speed

Model time, all inputs280.1 s
Mean per input23.34 s
Slowest input44.34 s
Inputs timed12

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens47.5K
Output tokens15.7K
Total tokens63.2K
Input cost$0.0951
Output cost$0.1878
Total cost$0.2829

Priced from 2026-09-16 · Input $/M $2.00 · Output $/M $12.00

Result 63 of 177

Test T1108 at 2026-09-16

manuscript transcription handwritten 15 de prose

x-ai grok-4.3 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score95.8%
CER ↓3.9%

{"folios":[{"folio":"3","text":"Und ein pferit die mir vnd
minen knechten vber Rulfend
sen Do was menan kein weg
denne den wir machtend
vnd vriechend die knechte vnd
vnd vil in vntz an den arb
vnd die pferit vntz an die
pettel vnd was ze mol ein grosser
nebel dz wir kum gepassend
vnd als mit grosser arbeit kome
wir ze mittem tag zu sant
kirchoffel vff den berg Do
do sach ich die bucher Do gar
vil herren wapen in ston
die ir stin do hin geleyg hand
Do stund mines vatters Rulgen
wopen och in dem einen","addition1":"","addition2":"","addition3":""}]}

Scoring

Fuzzy score68.5%
CER ↓34.2%

Speed

Model time, all inputs142.8 s
Mean per input11.90 s
Slowest input17.05 s
Inputs timed12

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens39.2K
Output tokens2.7K
Total tokens41.9K
Input cost$0.049
Output cost$0.0067
Total cost$0.084
Reasoning cost usd$0.0283
Total reasoning tokens11.3K

Priced from 2026-09-16 · Input $/M $1.25 · Output $/M $2.50

Result 64 of 177

Test T0658 at 2026-09-16

manuscript transcription handwritten 15 de prose

openai gpt-5.4-2026-03-05 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score93.1%
CER ↓7.1%

{"folios":[{"folio":"3","text":"Vnd ein pferit die mir mit
minen knechten vber hulffend
bay do was meinen kẽy weg
Sünne. den wir machtenß
vnd irslend die knecht dick
vnd vil in vntz an den aͧb
vnd die pferit vntz an die
petel vnd was ze mol ein grosser
nebel dz wir kaum geꝑaͤchend
vnd alß mit groffer arbeit kam
wir ze mittentag zů sant
krisoffel vff den berg do
do hub ich die bücher do gar
vil hering wapen zu stont
Die ir für dꝛhin gelien hand
So sind mines vatters flügen
wappen och in dem einen","addition1":"","addition2":"","addition3":""}]}

Scoring

Fuzzy score77.7%
CER ↓24.0%

Speed

Model time, all inputs75.0 s
Mean per input6.25 s
Slowest input10.37 s
Inputs timed12

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens42.5K
Output tokens2.5K
Total tokens45K
Input cost$0.1062
Output cost$0.0378
Total cost$0.144

Priced from 2026-09-16 · Input $/M $2.50 · Output $/M $15.00

Result 65 of 177

Test T1170 at 2026-09-16

manuscript transcription handwritten 15 de prose

genai gemini-3.1-flash-lite · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score95.2%
CER ↓4.2%

{"folios":[{"folio":"3","text":"Vnd ein pferit die mir vnd
minen knechten ibex hulfend
sen do was menen kein weg
denne den wir machtend
und viehend die knecht dut
vnd vil in vns an den arb
vnd die pferit vns an die
rettel vnd was ze mol ein grosse
nebel dz wir kum gepassend
vnd als mit grosser arbeit kame
wir ze mittem tag zu sant
Kristoffel vff den berg do
do sach ich die bucher do gar
vil herren wopen in stont
die in stur selbs geben hand
do stund mines vatters rigen
wopen ach in dem einen","addition1":"","addition2":"","addition3":""}]}

Scoring

Fuzzy score79.4%
CER ↓23.4%

Speed

Model time, all inputs36.1 s
Mean per input3.01 s
Slowest input4.55 s
Inputs timed12

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens19.1K
Output tokens2.9K
Total tokens22K
Input cost$0.0048
Output cost$0.0043
Total cost$0.0091

Priced from 2026-09-16 · Input $/M $0.25 · Output $/M $1.50

Result 66 of 177

Test T1351 at 2026-09-16

manuscript transcription handwritten 15 de prose

openai gpt-5.6-sol · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score94.1%
CER ↓4.9%

{"folios":[{"folio":"3","text":"vnd ein pfferit Die mir vnd
minen knechten wider hulffend
Dan Do was nemlich kein weg
Denne Daz wir muestend
vnd wielend Die knecht druf
vnd vil in vns an den arß
vnd Die pfferit was an Die
kettel vnd als ze mol ein grosser
nebel dz wir kum gesachend
vnd als mit grosser arbeit kamen
wir ze mittem tag zu sant
Cristoffel uff den berg Do
Do sach ich Die buecher Do gar
vil herren woppen jn stund
Die ir für do hin geben hand
Do stund mines vatters seligen
woppen och in dem einen","addition1":"","addition2":"","addition3":""}]}

Scoring

Fuzzy score76.4%
CER ↓24.8%

Speed

Model time, all inputs1119.5 s
Mean per input93.29 s
Slowest input127.21 s
Inputs timed12

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens47.5K
Output tokens48.1K
Total tokens95.7K
Input cost$0.1902
Output cost$0.9628
Total cost$1.15

Priced from 2026-09-16 · Input $/M $4.00 · Output $/M $20.00

Result 67 of 177

Test T1456 at 2026-09-16

manuscript transcription handwritten 15 de prose

x-ai grok-4.5 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score96.3%
CER ↓2.9%

{"folios":[{"folio":"3","text":"Vnd ein pferit die mir vnd
minen knechten vber hulfend
zen do was nienen kein weg
denne den wir machtend
vnd v̈beriend die knechte dick
vnd vil in vntz an den arb
vnd die pferit vntz an die
pettel vnd was ze mol ein grosser
nebel dz wir kum gesechend
vnd als mit grosser arbeit kame
wir ze mittem tag zü sant
cristoffel vff den berg do
do sach ich die bücher do gar
vil herren wapen jn stond
die ir stift do hin geben hand
do stund mines vatters schilt
wapen och in dem einen","addition1":"","addition2":"","addition3":""}]}

Scoring

Fuzzy score83.4%
CER ↓16.5%

Speed

Model time, all inputs583.1 s
Mean per input48.59 s
Slowest input109.29 s
Inputs timed12

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens43.6K
Output tokens2.8K
Total tokens46.4K
Input cost$0.0871
Output cost$0.0167
Total cost$0.2755
Reasoning cost usd$0.1717
Total reasoning tokens28.6K

Priced from 2026-09-16 · Input $/M $2.00 · Output $/M $6.00

Result 68 of 177

Test T0505 at 2026-09-16

manuscript transcription handwritten 15 de prose

genai gemini-3-flash-preview · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score95.9%
CER ↓3.5%

{"folios":[{"folio":"3","text":"Und ein pferit die mir und
minen knechten vber hulfend
den do was nenen kein weg
denne den wir machend
und vichend die knecht dik
und vil in vntz an den ars
und die pferit vntz an die
settel und was ze mol ein grosseꝛ
nebel dz wir kum gepassend
und als mit grosseꝛ arbeit kōme
wir ze mittem tag zu sant
kristoffel vff den berg do
do sach ich die bücher do gar
vil herren wopen in stund
die ir stür darzu geben hand
do stund mines vatteꝛs seligen
wopen ouch in dem einen","addition1":"","addition2":"","addition3":""}]}

Scoring

Fuzzy score76.0%
CER ↓24.0%

Speed

Model time, all inputs1142.3 s
Mean per input95.19 s
Slowest input111.57 s
Inputs timed12

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens19.1K
Output tokens2.8K
Total tokens22K
Input cost$0.0096
Output cost$0.0085
Total cost$0.9934
Reasoning cost usd$0.9753
Total reasoning tokens325.1K

Priced from 2026-09-16 · Input $/M $0.50 · Output $/M $3.00

Result 69 of 177

Test T1771 at 2026-09-16

manuscript transcription handwritten 15 de prose

deepseek deepseek-flash · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score92.1%
CER ↓6.6%

{"folios":[{"folio":"3","text":"Vnd ein pferit Die mir vnd
miner knechten vber halfend
Das Do was menan fuen weg
Samc Das wir machtend
vnd viehend Die knechte Die
vnd vil in vns an Den and
vnd Die pferit vns an Die
perg vnd was ze mol ein großer
nebel Dz wir kum gesahend
vnd als mit grosser arbeit kamen
wir ze mittern tag zue sant
kristoffel vff Den berg Do
Do sach ich Die buecher Do gar
vil Herren wapen In stane
Die ir piter vnd gin geheien hand
Do stand mines vatters pflegen
wapen och in dem emer","addition1":"","addition2":"","addition3":""}]}

Scoring

Fuzzy score70.4%
CER ↓32.8%

Speed

Model time, all inputs920.5 s
Mean per input76.71 s
Slowest input118.19 s
Inputs timed12

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens21.4K
Output tokens206.3K
Total tokens227.7K
Input cost$0.0064
Output cost$0.2475
Total cost$0.254

Priced from 2026-09-16 · Input $/M $0.30 · Output $/M $1.20

Result 70 of 177

Test T1741 at 2026-09-09

manuscript transcription handwritten 15 de prose

openrouter qwen/qwen3.8-27b · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score92.4%
CER ↓7.9%

{"folios":[{"folio":"3","text":"Vnd ein pferit die mir vnd
minen knechten vber Rulpend
ten do was menen keyn weg
deme den wir machend
vnd vichend die knechte daz
vnd vil in vntgangen arb
vnd die pferit vns an die
reut vnd was ze mol ein groffe
nebel daz wir kum gepastend
vnd als mit groffer arbeit kome
wir ze mittem tag zu pant
kropffel vff den berg do
do fah ich die bucher do gar
vil ferrin wopen zu pone
die ir stir edeln geben gand
do fand minnes vattero plegen
wopen setz in dem sinen","addition1":"","addition2":"","addition3":""}]}

Scoring

Fuzzy score72.9%
CER ↓30.8%

Speed

Model time, all inputs1429.8 s
Mean per input119.15 s
Slowest input452.31 s
Inputs timed12

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens40.2K
Output tokens68K
Total tokens108.1K
Input cost$0.0169
Output cost$0.2039
Total cost$0.2207

Priced from 2026-09-08 · Input $/M $0.42 · Output $/M $3.00