RISE Humanities Data Benchmark, 0.6.1

Benchmark Results

Library Cards

A comprehensive benchmark focused on catalog card analysis and information extraction from historical library catalog systems. This benchmark evaluates models on structured data extraction from digitized catalog cards, testing their ability to parse complex bibliographic information, author names, dates, and hierarchical catalog structures from historical Swiss library records.

Dataset Description Result Overview Test Runs

This benchmark has been run 246 times. It uses f1_macro metric.

Overview

Tested providers: alibaba, anthropic, cohere, deepseek, genai, huggingface, mistral, openai, openrouter, scicore, x-ai

Tested models: GLM-4.5V-FP8, MiniMaxAI/MiniMax-M3:deepinfra, Qwen/Qwen3-VL-235B-A22B-Instruct:deepinfra, Qwen3.8-Flash-Next-FP8, claude-3-5-sonnet-20241022, claude-3-7-sonnet-20250219, claude-3-opus-20240229, claude-fable-5, claude-fable-5-1, claude-haiku-4-5-20251001, claude-opus-4-1-20250805, claude-opus-4-20250514, claude-opus-4-5-20251101, claude-opus-4-6, claude-opus-4-7, claude-opus-4-8, claude-opus-5, claude-opus-5-5, claude-sonnet-4-20250514, claude-sonnet-4-5-20250929, claude-sonnet-4-6, claude-sonnet-5, claude-sonnet-5-5, command-a-vision-07-2025, deepseek-flash, deepseek-v4-flash-vision-exp, gemini-2.0-flash, gemini-2.0-flash-lite, gemini-2.5-flash, gemini-2.5-flash-lite, gemini-2.5-flash-lite-preview-09-2025, gemini-2.5-flash-preview-09-2025, gemini-2.5-pro, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-flash-lite, gemini-3.1-flash-lite-preview, gemini-3.1-pro-preview, gemini-3.5-flash, gemini-3.5-flash-lite, gemini-3.6-flash, gemini-3.7-flash, gemini-3.8-flash, google/gemma-4-26b-a4b-it, google/gemma-4-31b-it, gpt-4.1, gpt-4.1-mini, gpt-4.1-nano, gpt-4o, gpt-4o-mini, gpt-5, gpt-5-mini, gpt-5-nano, gpt-5.1-2025-11-13, gpt-5.2-2025-12-11, gpt-5.3-codex, gpt-5.4-2026-03-05, gpt-5.5-2026-04-23, gpt-5.6-luna, gpt-5.6-sol, gpt-5.6-terra, gpt-6-astra, gpt-6-luna, gpt-6-sol, grok-4.20-0309-reasoning, grok-4.3, grok-4.5, grok-4.6, grok-4.7, magistral-medium-2509, magistral-small-2509, meta-llama/llama-4-maverick, meta-llama/llama-4-scout, meta-models/Muse-Glimmer-30B:together, meta/muse-spark-1.2, meta/muse-spark-1.3, ministral-14b-2512, ministral-8b-2512, mistral-large-2411, mistral-large-2512, mistral-medium-2505, mistral-medium-2508, mistral-medium-3.5, mistral-small-2506, moonshotai/kimi-k3, o3, pixtral-12b, pixtral-large-2411, qwen/qwen3-vl-30b-a3b-instruct, qwen/qwen3-vl-8b-instruct, qwen/qwen3-vl-8b-thinking, qwen/qwen3.5-122b-a10b, qwen/qwen3.5-27b, qwen/qwen3.5-35b-a3b, qwen/qwen3.5-397b-a17b, qwen/qwen3.5-9b, qwen/qwen3.5-flash-02-23, qwen/qwen3.5-plus-02-15, qwen/qwen3.6-plus, qwen/qwen3.7-plus, qwen/qwen3.8-27b, qwen/qwen3.8-flash, qwen/qwen3.8-max, qwen/qwen3.8-max-0902, qwen3.5-122b-a10b, qwen3.5-27b, qwen3.5-35b-a3b, qwen3.5-397b-a17b, qwen3.5-flash-2026-02-23, qwen3.5-plus-2026-02-15, qwen35-397b-a17b-fp8, stepfun/step-3.7-flash, swiss-ai/Apertus-v1.5-70B:publicai, swiss-ai/Apertus-v1.5-8B:publicai, thinkingmachines/Inkling-Small:deepinfra, thinkingmachines/Inkling:together, x-ai/grok-4, z-ai/glm-5.3-flash, z-ai/glm-5v-turbo

Last 5 Runs

ScoreDateProviderModel
90.811 week agoanthropicclaude-opus-5-5
79.331 week agoscicoreQwen3.8-Flash-Next-FP8
91.031 week agoanthropicclaude-sonnet-5-5
89.081 week agoanthropicclaude-fable-5-1
91.441 week agoanthropicclaude-sonnet-5-5

All test runs

Contributors

RoleContributors
Domain expertGabriel Müller
Data curatorGabriel Müller
AnnotatorMaximilian Hindermann, Gabriel Müller
AnalystMaximilian Hindermann
EngineerMaximilian Hindermann

Tags
  • Type(s): index-card
  • Benchmark task(s):  information-extraction
  • Writing: typed, printed, handwritten
  • Source creation (century): 20, 19
  • Source Layout: index
  • Language(s): de, fr, en, la, el, fi, sv, pl