RISE Humanities Data Benchmark, 0.6.1

Search Test Runs

 

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

  • Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
  • Model configuration – provider, model version, temperature, and other generation parameters.
  • Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
  • Usage and cost data – token counts and calculated API costs.
  • Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Loading module...

Search Results

Your search for Benchmark 'medieval_manuscripts__true' with Search Hidden 'False' returned 177 results, showing page 4 of 18.
Result 31 of 177

Test T1396 at 2026-09-17

manuscript transcription handwritten 15 de prose

openrouter moonshotai/kimi-k3 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Scoring

Fuzzy score69.8%
CER ↓32.3%

Speed

Model time, all inputs2919.9 s
Mean per input243.33 s
Slowest input389.53 s
Inputs timed12

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens54K
Output tokens125.6K
Total tokens179.6K
Input cost$0.0742
Output cost$1.84
Total cost$1.91

Priced from 2026-09-16 · Input $/M $2.65 · Output $/M $13.28

Result 32 of 177

Test T1681 at 2026-09-17

manuscript transcription handwritten 15 de prose

anthropic claude-fable-5-1 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score98.2%
CER ↓0.6%

{
  "folios": [
    {
      "folio": "3",
      "text": "vnd ein pferit die mir vnd
minen knechten vber hulfend
den do was nienan kein weg
denne den wir machtend
vnd vielend die knecht dick
vnd vil in vntz an den ars
vnd die pferit vntz an die
settel vnd was ze mol ein grosser
nebel dz wir kum gesachend
vnd als mit grosser arbeit kame
wir ze mittem tag zuo sant
kristoffel vff den berg do
do sach ich die buecher do gar
vil herren wopen in stone
die ir stur de hin geben hand
do stuond mines vatters seligen
wopen och in dem einen",
      "addition1": "",
      "addition2": "",
      "addition3": ""
    }
  ]
}

Scoring

Fuzzy score85.1%
CER ↓15.1%

Speed

Model time, all inputs238.8 s
Mean per input19.90 s
Slowest input31.50 s
Inputs timed12

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens52.6K
Output tokens12.3K
Total tokens65K
Input cost$0.5261
Output cost$0.6175
Total cost$1.14

Priced from 2026-09-08 · Input $/M $10.00 · Output $/M $50.00

Result 33 of 177

Test T1621 at 2026-09-17

manuscript transcription handwritten 15 de prose

openrouter meta/muse-spark-1.2 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score95.4%
CER ↓3.9%

{"folios":[{"folio":"3","text":"Vnd ein pferit die wir mit
minen knechten vber hulffend
den So was mench kuon weg
denne den wir machtend
vnd vichend die knecht duk
vnd vil in vntz an den arb
vnd die pferit vntz an die
petel vnd was ze mal ein grosser
nebel dz wir kum gesahend
vnd also mit grosser arbeit kame
wir ze mittem tag zuo sant
kristoffel vff den berg do
do sach ich die buocher do gar
vil herren wopen zu stone
die in stur edlen geben hand
do stuond mines vatters seligen
wopen och in dem cimey","addition1":"","addition2":"","addition3":""}]}

Scoring

Fuzzy score82.9%
CER ↓17.7%

Speed

Model time, all inputs148.5 s
Mean per input12.38 s
Slowest input19.72 s
Inputs timed12

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens49.6K
Output tokens36.3K
Total tokens86K
Input cost$0.0621
Output cost$0.1544
Total cost$0.2165

Priced from 2026-08-18 · Input $/M $1.25 · Output $/M $4.25

Result 34 of 177

Test T0976 at 2026-09-17

manuscript transcription handwritten 15 de prose

openrouter qwen/qwen3.5-flash-02-23 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score93.5%
CER ↓6.7%

{"folios":[{"folio":"3","text":"vnd ein pferit die wir vnd
minen knechten wider rufend
sen do was meinen kün weg
deme den wir machten vnd
vielen die knecht vnd
vil in vnz an den and
vnd die pferit vnz an die
pettel vnd was ze mol ein grofse
nebel dz wir kum gepaßend
vnd als mit großer arbeit kam
wir ze mittem tag zu pant
krapffel vff den berg do
do sach ich die bücher so gar
vil herren wappen zu ston
die ir stir ob im geben hand
do stund meines vatters pflegen
wappen och in dem einen","addition1":"","addition2":"","addition3":""}]}

Scoring

Fuzzy score67.9%
CER ↓34.5%

Speed

Model time, all inputs745.7 s
Mean per input62.14 s
Slowest input137.71 s
Inputs timed12

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens39.7K
Output tokens76.5K
Total tokens116.2K
Input cost$0.0026
Output cost$0.0199
Total cost$0.0225

Priced from 2026-09-16 · Input $/M $0.065 · Output $/M $0.26

Result 35 of 177

Test T0303 at 2026-09-17

manuscript transcription handwritten 15 de prose

openrouter qwen/qwen3-vl-8b-instruct · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score91.7%
CER ↓7.9%

{"folios":[{"folio":"3","text":"Und ein pferit Die mir und
minen knechten vber Rulfend
sag Do was menen zyn weg
some den wir machtend
und vielend die knchte dud
und vil in vntz angen arb
und die pferit vntz an die
pettel und was ze mol ein groffer
nebel dz wir zum ogepaffend
und als mit groffer arbeit kome
wir ze mitttem tag zu Pant
Einpoffel vff den berg do
Do sach ich die buchter do gar
vil gerrin Woppen in stone
Die ir stur Edlhm gebeey hand
Do pound mince vatterb Plugen
Woppen ach in dem ziney","addition1":"","addition2":"","addition3":""}]}

Scoring

Fuzzy score64.0%
CER ↓38.4%

Speed

Model time, all inputs43.2 s
Mean per input3.60 s
Slowest input4.97 s
Inputs timed12

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens36.3K
Output tokens2.7K
Total tokens39K
Input cost$0.0042
Output cost$0.0012
Total cost$0.0055

Priced from 2026-09-16 · Input $/M $0.117 · Output $/M $0.455

Result 36 of 177

Test T1196 at 2026-09-17

manuscript transcription handwritten 15 de prose

anthropic claude-fable-5 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score98.0%
CER ↓0.9%

{"folios":[{"folio":"3","text":"Vnd ein pferit die mir vnd
minen knechten vber hulfend
den do was menan kein weg
denne den wir machtend
vnd vielend die knecht dick
vnd vil in vntz an den ars
vnd die pferit vntz an die
settel vnd was ze mol ein grosser
nebel dz wir kum gesachend
vnd also mit grosser arbeit kame
wir ze mittem tag zuo sant
kristoffel vff den berg do
do sach ich die buecher do gar
vil herren wopen in stein
die ir stur de hin geben hand
do stuond mines vatters seligen
wopen och in dem einen","addition1":"","addition2":"","addition3":""}]}

Scoring

Fuzzy score87.8%
CER ↓11.6%

Speed

Model time, all inputs137.4 s
Mean per input11.45 s
Slowest input16.82 s
Inputs timed12

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens61.9K
Output tokens4.2K
Total tokens66.1K
Input cost$0.6186
Output cost$0.2106
Total cost$0.8291

Priced from 2026-09-16 · Input $/M $10.00 · Output $/M $50.00

Result 37 of 177

Test T0924 at 2026-09-17

manuscript transcription handwritten 15 de prose

openrouter qwen/qwen3.5-27b · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score93.9%
CER ↓5.8%

{"folios":[{"folio":"3","text":"vnd ein pferit die mir vnd
minen knechten wider gultend
san so was menan kanz weg
deme dan wir machgend
vnd vichend die knechte durb
vnd vil in vntz an den arb
vnd die pferit vntz an die
pettel vnd was ze mol ein grosser
nebel dz wir kum gepatsend
vnd als mit grosser arbeit kame
wir ze mittren tag zu Pant
kirstoffel uff den berg do
do sach ich die burger do gar
vil herren wopen im ston
sie ir stur dz sin gebey hand
do stand minneb vatterb slugen
wopen och in dem einen","addition1":"","addition2":"","addition3":""}]}

Scoring

Fuzzy score73.9%
CER ↓29.6%

Speed

Model time, all inputs474.9 s
Mean per input39.58 s
Slowest input183.67 s
Inputs timed12

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens37.9K
Output tokens16.7K
Total tokens54.5K
Input cost$0.0074
Output cost$0.026
Total cost$0.0334

Priced from 2026-09-16 · Input $/M $0.195 · Output $/M $1.56

Result 38 of 177

Test T1531 at 2026-09-17

manuscript transcription handwritten 15 de prose

huggingface MiniMaxAI/MiniMax-M3 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score91.9%
CER ↓8.5%

{"folios":[{"folio":"3","text":"vnd ein pherit die wir vnd
minen knehten oder hulffind
den So was minen kneht weg-
denne den wir machtend
vnd richtend die kneht sich
mit vil in vng an den vnd
vnd die pherit vng an die
ketel vnd was ze mol in grosser
not dz wir kum geachtend
vnd als mit grosser arbeite kam-
en wir ze mittem tag zu sant
fridhoff vff den berge do
do sach ich die brucken do gar
vil herren wopen in hand
die in fůr schin geberg hand
do stund minen vatterd klaget
wopen och in dem einig","addition1":"","addition2":"","addition3":""}]}

Scoring

Fuzzy score65.3%
CER ↓39.4%

Speed

Model time, all inputs415.5 s
Mean per input34.63 s
Slowest input263.96 s
Inputs timed12

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens14.5K
Output tokens2.7K
Total tokens17.2K
Input cost$0.0041
Output cost$0.0029
Total cost$0.007

Priced from 2026-09-16 · Input $/M $0.28 · Output $/M $1.10

Result 39 of 177

Test T0950 at 2026-09-17

manuscript transcription handwritten 15 de prose

openrouter qwen/qwen3.5-397b-a17b · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score93.8%
CER ↓5.9%

{"folios":[{"folio":"3","text":"Vnd ein pferit die mir vnd
minen knechten vber Ruland
ten do was menen ganꝛ weg
deme den wir machtend
vnd vieliend die knehte daz
vnd vil in vntz an den arb
vnd die pferit vntz an die
sattel vnd was ze mol ein groſſer
nebel dz wir kum gepaſſend
vnd alſ mit groſſer arbeit kame
wir ze mittem tag zu Sant
Cristoffel vff den berg do
do ſach ich die buͤchſen do gar
vil herren wapen zu ſtand
die ir stirn ſchin gebey hand
do ſtund mines vatters pluͤgen
wapen och in dem einen","addition1":"","addition2":"","addition3":""}]}

Scoring

Fuzzy score70.4%
CER ↓30.9%

Speed

Model time, all inputs5296.0 s
Mean per input441.33 s
Slowest input1691.02 s
Inputs timed12

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens32.3K
Output tokens64.3K
Total tokens96.7K
Input cost$0.0178
Output cost$0.2252
Total cost$0.243

Priced from 2026-09-16 · Input $/M $0.55 · Output $/M $3.50

Result 40 of 177

Test T1082 at 2026-09-16

manuscript transcription handwritten 15 de prose

anthropic claude-opus-4-8 · temp 0.0 · dataclass Document

Response vs ground truth

One model response from this run. The per-input comparison needs JavaScript.

Fuzzy score98.2%
CER ↓1.8%

{"folios":[{"folio":"3","text":"vnd ein pferit die mir vnd
 minen knechten iber hulffend
 den do was menen kein weg
 denne den wir machtend
 vnd ziehend die knecht dick
 vnd vil in vntz an den ars
 vnd die pferit vntz an die
 settel vnd was ze mol ein grosser
 nebel dz wir kum gesahend
 vnd als mit grosser arbeit kome
 wir ze mittem tag zu sant
 kristoffel vff den berg do
 do sach ich die bücher do gar
 vil herren wopen zu stone
 die ir stür do hin geben hand
 do stund mines vatters seligen
 wopen och in dem einen","addition1":null,"addition2":null,"addition3":null}]}

Scoring

Fuzzy score81.0%
CER ↓19.3%

Speed

Model time, all inputs119.4 s
Mean per input9.95 s
Slowest input17.13 s
Inputs timed12

Summed model time, not elapsed: requests run in parallel and the worker count is not recorded, so the run finished sooner than this.

Costs / Pricing

Input tokens61.9K
Output tokens3.9K
Total tokens65.8K
Input cost$0.3093
Output cost$0.0978
Total cost$0.4072

Priced from 2026-09-16 · Input $/M $5.00 · Output $/M $25.00