RISE Humanities Data Benchmark, 0.6.1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'medieval_manuscripts__true' with Search Hidden 'False' returned 177 results, showing page 3 of 18.
Result 21 of 177

Test T0963 at 2026-09-17

manuscript transcription handwritten 15 de prose

openrouter qwen/qwen3.5-plus-02-15 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score94.7%
CER ↓5.1%

{"folios":[{"folio":"3","text":"Vnd ein pferit die mir vnd
minnen knechten vber Aulfend
Sey Do was menen lang weg
denne das wir machtend
vnd vielend die knechte daz
vnd vil in vntz an den arb
vnd die pferit vns an die
pettel vnd was ze mol ein grosse
nebel dz wir kum gepassend
vnd als mit grosser arbeit kame
wir ze mittem tag zu sant
kristoffel vff den berg do
do sach ich die bucher do gar
vil herren wapey in stone
die ir stur helm gebery hand
do stund mines vatters plugen
wapey och in dem einen","addition1":"","addition2":"","addition3":""}]}

Scoring

Fuzzy score74.5%
CER ↓30.0%

Speed

Model time, all inputs540.8 s
Mean per input45.07 s
Slowest input173.98 s
Inputs timed12

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens39.7K
Output tokens28.7K
Total tokens68.4K
Input cost$0.0103
Output cost$0.0448
Total cost$0.0551

Priced from 2026-09-16 · Input $/M $0.26 · Output $/M $1.56

Result 22 of 177

Test T1636 at 2026-09-17

manuscript transcription handwritten 15 de prose

openrouter z-ai/glm-5v-turbo · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

{
  "folios": [
    {
      "folio": "3",
      "text": "Vnd ein pfeſcit Sie mir vn̄
minen Enckten wider Auſſend
Son ſo waß minenſon ſein weg
Seme Son mir maſten
vn̄ vichtenſo Sie Encte ſuſ
vn̄ vil in vnganſen anb
vno Sie pfeſcit vng an Sie
Pertel vn̄ waß ze mot ein groſŝe
nebet eſ wir kün yepaſſenb
vn̄ aſſo mit groſŝer awſſeit bome
wir ze mittentag zū Punt
Eniſtoſſel vfſſen berg ſo
ſo paß iſt die bichter ſo gar
vil gerüſch woſchen zū fone
Sie in für ſein geſichy gand
ſo paind mineſ wattero Plügen
woſchen aſſ in ſem einſey",
      "addition1": "",
      "addition2": "",
      "addition3": ""
    }
  ]
}

Scoring

Fuzzy score0.0%
CER ↓100.0%

Speed

Model time, all inputs1988.0 s
Mean per input165.67 s
Slowest input233.43 s
Inputs timed12

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens51.1K
Output tokens21.1K
Total tokens72.2K
Input cost$0.0448
Output cost$0.0843
Total cost$0.1291

Priced from 2026-08-18 · Input $/M $1.20 · Output $/M $4.00

Result 23 of 177

Test T0937 at 2026-09-17

manuscript transcription handwritten 15 de prose

openrouter qwen/qwen3.5-35b-a3b · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score92.4%
CER ↓7.9%

{"folios":[{"folio":"3","text":"vnd ein pferit die wir vnd
minen knechten vber gulfend
gen do was minen künig weg
deme das wir machend
vnd vichend die knecht vnd
vnd vil in vnz an den and
vnd sie pferit vnz an die
pottel vnd was ze mol im grofse
nebel dz wir kum gelaufend
vnd als mir grofser arbeit kome
wir ze mittem tag zu pant
krapfel vff den berg do
do sach ich die bücher so gar
vil herren wapen zu pons
die in pür vnd hin gelaufen
do stund mines vatters pflegen
wapen och in dem zwen","addition1":"","addition2":"","addition3":""}]}

Scoring

Fuzzy score64.3%
CER ↓37.6%

Speed

Model time, all inputs1471.6 s
Mean per input122.64 s
Slowest input513.86 s
Inputs timed12

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens25.3K
Output tokens44.5K
Total tokens69.8K
Input cost$0.0038
Output cost$0.0576
Total cost$0.0613

Priced from 2026-09-16 · Input $/M $0.1625 · Output $/M $1.30

Result 24 of 177

Test T1196 at 2026-09-17

manuscript transcription handwritten 15 de prose

anthropic claude-fable-5 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score98.0%
CER ↓0.9%

{"folios":[{"folio":"3","text":"Vnd ein pferit die mir vnd
minen knechten vber hulfend
den do was menan kein weg
denne den wir machtend
vnd vielend die knecht dick
vnd vil in vntz an den ars
vnd die pferit vntz an die
settel vnd was ze mol ein grosser
nebel dz wir kum gesachend
vnd also mit grosser arbeit kame
wir ze mittem tag zuo sant
kristoffel vff den berg do
do sach ich die buecher do gar
vil herren wopen in stein
die ir stur de hin geben hand
do stuond mines vatters seligen
wopen och in dem einen","addition1":"","addition2":"","addition3":""}]}

Scoring

Fuzzy score87.8%
CER ↓11.6%

Speed

Model time, all inputs137.4 s
Mean per input11.45 s
Slowest input16.82 s
Inputs timed12

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens61.9K
Output tokens4.2K
Total tokens66.1K
Input cost$0.6186
Output cost$0.2106
Total cost$0.8291

Priced from 2026-09-16 · Input $/M $10.00 · Output $/M $50.00

Result 25 of 177

Test T1696 at 2026-09-17

manuscript transcription handwritten 15 de prose

openrouter meta/muse-spark-1.3 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score95.8%
CER ↓3.4%

{"folios":[{"folio":"3","text":"Vnd ein pferit die wir vnd
minen knechten vber hulfend
den do was menen kein weg
denne den wir machtend
vnd vichend die knecht dik
vnd vil in vntz an den arb
vnd die pferit vntz an die
pettel vnd was ze mal ein grosser
nebel dz wir kum gesahend
vnd also mit grosser arbeit kame
wir ze mittem tag zuo sant
kristoffel vff den berg do
Do sach ich die bucher do gar
vil herren wopen zuo rome
Die in stir Ochm geben hand
Do stund mines vatters schigen
wopen och in dem emen","addition1":"","addition2":"","addition3":""}]}

Scoring

Fuzzy score82.4%
CER ↓19.1%

Speed

Model time, all inputs268.7 s
Mean per input22.40 s
Slowest input35.31 s
Inputs timed12

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens51.5K
Output tokens30.5K
Total tokens82K
Input cost$0.0643
Output cost$0.1296
Total cost$0.1939

Priced from 2026-09-08 · Input $/M $1.25 · Output $/M $4.25

Result 26 of 177

Test T1546 at 2026-09-17

manuscript transcription handwritten 15 de prose

huggingface thinkingmachines/Inkling-Small · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score24.0%
CER ↓78.2%

{"folios":[{"folio":"1","text":"Und ein pferit die mir und
minen knechten oder gulsend
Ich was menig tag weg
Inne den murfend
und vilend sie brucht
vnd vil in vng an den ang
vnd die pferit vng in die
part und was ze mol im grossen
nebel vnd als mit grosser arbeit
kam vnd ze mittem tag zu sant
Fursch vff die bucher
So frech ich die bucher
zu faren wogen in sone
die in stur schin gretin hand
So stur wogen vch in dem","addition1":"a","addition2":"b","addition3":"c"}]}

Scoring

Fuzzy score44.0%
CER ↓62.0%

Speed

Model time, all inputs824.6 s
Mean per input68.72 s
Slowest input103.01 s
Inputs timed12

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens28K
Output tokens102.8K
Total tokens130.8K
Input cost$0.0126
Output cost$0.1234
Total cost$0.136

Priced from 2026-09-16 · Input $/M $0.45 · Output $/M $1.20

Result 27 of 177

Test T0976 at 2026-09-17

manuscript transcription handwritten 15 de prose

openrouter qwen/qwen3.5-flash-02-23 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score93.5%
CER ↓6.7%

{"folios":[{"folio":"3","text":"vnd ein pferit die wir vnd
minen knechten wider rufend
sen do was meinen kün weg
deme den wir machten vnd
vielen die knecht vnd
vil in vnz an den and
vnd die pferit vnz an die
pettel vnd was ze mol ein grofse
nebel dz wir kum gepaßend
vnd als mit großer arbeit kam
wir ze mittem tag zu pant
krapffel vff den berg do
do sach ich die bücher so gar
vil herren wappen zu ston
die ir stir ob im geben hand
do stund meines vatters pflegen
wappen och in dem einen","addition1":"","addition2":"","addition3":""}]}

Scoring

Fuzzy score67.9%
CER ↓34.5%

Speed

Model time, all inputs745.7 s
Mean per input62.14 s
Slowest input137.71 s
Inputs timed12

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens39.7K
Output tokens76.5K
Total tokens116.2K
Input cost$0.0026
Output cost$0.0199
Total cost$0.0225

Priced from 2026-09-16 · Input $/M $0.065 · Output $/M $0.26

Result 28 of 177

Test T1681 at 2026-09-17

manuscript transcription handwritten 15 de prose

anthropic claude-fable-5-1 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score98.2%
CER ↓0.6%

{
  "folios": [
    {
      "folio": "3",
      "text": "vnd ein pferit die mir vnd
minen knechten vber hulfend
den do was nienan kein weg
denne den wir machtend
vnd vielend die knecht dick
vnd vil in vntz an den ars
vnd die pferit vntz an die
settel vnd was ze mol ein grosser
nebel dz wir kum gesachend
vnd als mit grosser arbeit kame
wir ze mittem tag zuo sant
kristoffel vff den berg do
do sach ich die buecher do gar
vil herren wopen in stone
die ir stur de hin geben hand
do stuond mines vatters seligen
wopen och in dem einen",
      "addition1": "",
      "addition2": "",
      "addition3": ""
    }
  ]
}

Scoring

Fuzzy score85.1%
CER ↓15.1%

Speed

Model time, all inputs238.8 s
Mean per input19.90 s
Slowest input31.50 s
Inputs timed12

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens52.6K
Output tokens12.3K
Total tokens65K
Input cost$0.5261
Output cost$0.6175
Total cost$1.14

Priced from 2026-09-08 · Input $/M $10.00 · Output $/M $50.00

Result 29 of 177

Test T1531 at 2026-09-17

manuscript transcription handwritten 15 de prose

huggingface MiniMaxAI/MiniMax-M3 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score91.9%
CER ↓8.5%

{"folios":[{"folio":"3","text":"vnd ein pherit die wir vnd
minen knehten oder hulffind
den So was minen kneht weg-
denne den wir machtend
vnd richtend die kneht sich
mit vil in vng an den vnd
vnd die pherit vng an die
ketel vnd was ze mol in grosser
not dz wir kum geachtend
vnd als mit grosser arbeite kam-
en wir ze mittem tag zu sant
fridhoff vff den berge do
do sach ich die brucken do gar
vil herren wopen in hand
die in fůr schin geberg hand
do stund minen vatterd klaget
wopen och in dem einig","addition1":"","addition2":"","addition3":""}]}

Scoring

Fuzzy score65.3%
CER ↓39.4%

Speed

Model time, all inputs415.5 s
Mean per input34.63 s
Slowest input263.96 s
Inputs timed12

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens14.5K
Output tokens2.7K
Total tokens17.2K
Input cost$0.0041
Output cost$0.0029
Total cost$0.007

Priced from 2026-09-16 · Input $/M $0.28 · Output $/M $1.10

Result 30 of 177

Test T0302 at 2026-09-17

manuscript transcription handwritten 15 de prose

openrouter qwen/qwen3-vl-30b-a3b-instruct · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score91.6%
CER ↓10.7%

{"folios":[{"folio":"3","text":"Und ein pferit die mir und
 minen knechten vber hulfend
 So was menen kein weg
 deme den wir machend
 und viel in vntz an den and
 end sic pferit vmb an die
 pettel und was ze mot ein groffe
 nebel & win kum gecpaffend
 und als mit groffer arbeit komt
 wir ze mittem tag zu Punt
 enproffel vff den berg so
 So sach ich die bucher So gar
 vil derrey wopen im ston
 Die in stur & gun gebey hand
 So sind minen vatter b ligen
 wopen sich in dem sinn","addition1":"","addition2":"","addition3":""}]}

Scoring

Fuzzy score64.4%
CER ↓38.4%

Speed

Model time, all inputs88.1 s
Mean per input7.34 s
Slowest input18.81 s
Inputs timed12

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens36.8K
Output tokens2.7K
Total tokens39.5K
Input cost$0.0055
Output cost$0.0016
Total cost$0.0071

Priced from 2026-09-16 · Input $/M $0.15 · Output $/M $0.60