RISE Humanities Data Benchmark

A test run is a single execution of a benchmark test using a defined model configuration.
Each run represents how a particular large language model (LLM) — such as GPT-4, Claude-3, or Gemini — performed on a given task at a specific time, with specific settings.

A test run includes:

Prompt and role definition – what the model was asked to do and from what perspective (e.g. “as a historian”).
Model configuration – provider, model version, temperature, and other generation parameters.
Results – the model’s actual response and its evaluation (scores such as F1 or accuracy).
Usage and cost data – token counts and calculated API costs.
Metadata – information like the test date, benchmark name, and person who executed it.

Together, test runs make it possible to compare models, providers, and configurations across benchmarks in a transparent and reproducible way.

Result 81 of 119

Test T0234 at 2025-10-17

{'document-type': ['book-page'], 'writing': ['printed'], 'century': [20], 'language': ['en'], 'layout': ['list'], 'entry-type': ['bibliographic'], 'task': ['information-extraction']}

Configuration

Provider	openrouter
Model	meta-llama/llama-4-maverick

Temperature	0.0
Dataclass	Document

Normalized Score	62.57 %
Test time	unknown seconds

Prompt

Results

no valid result

Scoring

Fuzzy Score	F1 micro / macro		Micro precision/recall		Tue/False Positives
0.63	n/a	n/a	n/a	n/a	n/a	n/a	n/a	n/a
			Micro Precision	Micro Recall	Instances	TP	FP	FN

Costs / Pricing

Pricing Date: 8 months ago, 2025-10-17.

Tokens: 7.7K IT + 8.4K OT = 16.2K TT

Cost: 0.001$ + 0.005$ = 0.006$

Cite: Hindermann, Marti, Alberto, et al., (2025). RISE-UNIBAS/humanities_data_benchmark, 10.5281/zenodo.16941752

Result 82 of 119

Test T0140 at 2025-10-01

{'document-type': ['book-page'], 'writing': ['printed'], 'century': [20], 'language': ['en'], 'layout': ['list'], 'entry-type': ['bibliographic'], 'task': ['information-extraction']}

Configuration

Provider	openai
Model	gpt-4.1-mini

Temperature	0.0
Dataclass	Document

Normalized Score	64.83 %
Test time	unknown seconds

Prompt

Results

no valid result

Scoring

Fuzzy Score	F1 micro / macro		Micro precision/recall		Tue/False Positives
0.65	n/a	n/a	n/a	n/a	n/a	n/a	n/a	n/a
			Micro Precision	Micro Recall	Instances	TP	FP	FN

Costs / Pricing

Pricing Date: 9 months ago, 2025-10-01.

Tokens: 8.8K IT + 10.2K OT = 19.1K TT

Cost: 0.004$ + 0.016$ = 0.020$

Cite: Hindermann, Marti, Alberto, et al., (2025). RISE-UNIBAS/humanities_data_benchmark, 10.5281/zenodo.16941752

Result 83 of 119

Test T0225 at 2025-10-01

{'document-type': ['book-page'], 'writing': ['printed'], 'century': [20], 'language': ['en'], 'layout': ['list'], 'entry-type': ['bibliographic'], 'task': ['information-extraction']}

Configuration

Provider	anthropic
Model	claude-sonnet-4-5-20250929

Temperature	0.0
Dataclass	Document

Normalized Score	0.00 %
Test time	unknown seconds

Prompt

Results

no valid result

Scoring

Fuzzy Score	F1 micro / macro		Micro precision/recall		Tue/False Positives
0.00	n/a	n/a	n/a	n/a	n/a	n/a	n/a	n/a
			Micro Precision	Micro Recall	Instances	TP	FP	FN

Costs / Pricing

Pricing Date: 9 months ago, 2025-10-01.

Tokens: 12.8K IT + 8.3K OT = 21.1K TT

Cost: 0.038$ + 0.125$ = 0.164$

Cite: Hindermann, Marti, Alberto, et al., (2025). RISE-UNIBAS/humanities_data_benchmark, 10.5281/zenodo.16941752

Result 84 of 119

Test T0203 at 2025-10-01

{'document-type': ['book-page'], 'writing': ['printed'], 'century': [20], 'language': ['en'], 'layout': ['list'], 'entry-type': ['bibliographic'], 'task': ['information-extraction']}

Configuration

Provider	genai
Model	gemini-2.5-flash-lite

Temperature	0.0
Dataclass	Document

Normalized Score	54.45 %
Test time	unknown seconds

Prompt

Results

no valid result

Scoring

Fuzzy Score	F1 micro / macro		Micro precision/recall		Tue/False Positives
0.54	n/a	n/a	n/a	n/a	n/a	n/a	n/a	n/a
			Micro Precision	Micro Recall	Instances	TP	FP	FN

Costs / Pricing

Pricing Date: 9 months ago, 2025-10-01.

Tokens: 1.7K IT + 9.4K OT = 11.1K TT

Cost: 0.000$ + 0.004$ = 0.004$

Cite: Hindermann, Marti, Alberto, et al., (2025). RISE-UNIBAS/humanities_data_benchmark, 10.5281/zenodo.16941752

Result 85 of 119

Test T0139 at 2025-10-01

{'document-type': ['book-page'], 'writing': ['printed'], 'century': [20], 'language': ['en'], 'layout': ['list'], 'entry-type': ['bibliographic'], 'task': ['information-extraction']}

Configuration

Provider	openai
Model	gpt-4.1

Temperature	0.0
Dataclass	Document

Normalized Score	65.69 %
Test time	unknown seconds

Prompt

Results

no valid result

Scoring

Fuzzy Score	F1 micro / macro		Micro precision/recall		Tue/False Positives
0.66	n/a	n/a	n/a	n/a	n/a	n/a	n/a	n/a
			Micro Precision	Micro Recall	Instances	TP	FP	FN

Costs / Pricing

Pricing Date: 9 months ago, 2025-10-01.

Tokens: 7.2K IT + 10.1K OT = 17.3K TT

Cost: 0.014$ + 0.081$ = 0.095$

Cite: Hindermann, Marti, Alberto, et al., (2025). RISE-UNIBAS/humanities_data_benchmark, 10.5281/zenodo.16941752

Result 86 of 119

Test T0130 at 2025-10-01

{'document-type': ['book-page'], 'writing': ['printed'], 'century': [20], 'language': ['en'], 'layout': ['list'], 'entry-type': ['bibliographic'], 'task': ['information-extraction']}

Configuration

Provider	openai
Model	gpt-5-mini

Temperature	0.0
Dataclass	Document

Normalized Score	67.67 %
Test time	unknown seconds

Prompt

Results

no valid result

Scoring

Fuzzy Score	F1 micro / macro		Micro precision/recall		Tue/False Positives
0.68	n/a	n/a	n/a	n/a	n/a	n/a	n/a	n/a
			Micro Precision	Micro Recall	Instances	TP	FP	FN

Costs / Pricing

Pricing Date: 9 months ago, 2025-10-01.

Tokens: 7.4K IT + 28.2K OT = 35.6K TT

Cost: 0.002$ + 0.056$ = 0.058$

Cite: Hindermann, Marti, Alberto, et al., (2025). RISE-UNIBAS/humanities_data_benchmark, 10.5281/zenodo.16941752

Result 87 of 119

Test T0129 at 2025-10-01

{'document-type': ['book-page'], 'writing': ['printed'], 'century': [20], 'language': ['en'], 'layout': ['list'], 'entry-type': ['bibliographic'], 'task': ['information-extraction']}

Configuration

Provider	openai
Model	gpt-5

Temperature	0.0
Dataclass	Document

Normalized Score	68.53 %
Test time	unknown seconds

Prompt

Results

no valid result

Scoring

Fuzzy Score	F1 micro / macro		Micro precision/recall		Tue/False Positives
0.69	n/a	n/a	n/a	n/a	n/a	n/a	n/a	n/a
			Micro Precision	Micro Recall	Instances	TP	FP	FN

Costs / Pricing

Pricing Date: 9 months ago, 2025-10-01.

Tokens: 6.5K IT + 33.4K OT = 39.9K TT

Cost: 0.008$ + 0.334$ = 0.342$

Cite: Hindermann, Marti, Alberto, et al., (2025). RISE-UNIBAS/humanities_data_benchmark, 10.5281/zenodo.16941752

Result 88 of 119

Test T0133 at 2025-10-01

{'document-type': ['book-page'], 'writing': ['printed'], 'century': [20], 'language': ['en'], 'layout': ['list'], 'entry-type': ['bibliographic'], 'task': ['information-extraction']}

Configuration

Provider	openai
Model	o3

Temperature	0.0
Dataclass	Document

Normalized Score	66.68 %
Test time	unknown seconds

Prompt

Results

no valid result

Scoring

Fuzzy Score	F1 micro / macro		Micro precision/recall		Tue/False Positives
0.67	n/a	n/a	n/a	n/a	n/a	n/a	n/a	n/a
			Micro Precision	Micro Recall	Instances	TP	FP	FN

Costs / Pricing

Pricing Date: 9 months ago, 2025-10-01.

Tokens: 6.7K IT + 21.9K OT = 28.6K TT

Cost: 0.013$ + 0.175$ = 0.189$

Cite: Hindermann, Marti, Alberto, et al., (2025). RISE-UNIBAS/humanities_data_benchmark, 10.5281/zenodo.16941752

Result 89 of 119

Test T0170 at 2025-10-01

{'document-type': ['book-page'], 'writing': ['printed'], 'century': [20], 'language': ['en'], 'layout': ['list'], 'entry-type': ['bibliographic'], 'task': ['information-extraction']}

Configuration

Provider	mistral
Model	mistral-medium-2505

Temperature	0.0
Dataclass	Document

Normalized Score	66.14 %
Test time	unknown seconds

Prompt

Results

no valid result

Scoring

Fuzzy Score	F1 micro / macro		Micro precision/recall		Tue/False Positives
0.66	n/a	n/a	n/a	n/a	n/a	n/a	n/a	n/a
			Micro Precision	Micro Recall	Instances	TP	FP	FN

Costs / Pricing

Pricing Date: 9 months ago, 2025-10-01.

Tokens: 9.6K IT + 9.2K OT = 18.8K TT

Cost: 0.004$ + 0.018$ = 0.022$

Cite: Hindermann, Marti, Alberto, et al., (2025). RISE-UNIBAS/humanities_data_benchmark, 10.5281/zenodo.16941752

Result 90 of 119

Test T0141 at 2025-10-01

{'document-type': ['book-page'], 'writing': ['printed'], 'century': [20], 'language': ['en'], 'layout': ['list'], 'entry-type': ['bibliographic'], 'task': ['information-extraction']}

Configuration

Provider	openai
Model	gpt-4.1-nano

Temperature	0.0
Dataclass	Document

Normalized Score	32.02 %
Test time	unknown seconds

Prompt

Results

no valid result

Scoring

Fuzzy Score	F1 micro / macro		Micro precision/recall		Tue/False Positives
0.32	n/a	n/a	n/a	n/a	n/a	n/a	n/a	n/a
			Micro Precision	Micro Recall	Instances	TP	FP	FN

Costs / Pricing

Pricing Date: 9 months ago, 2025-10-01.

Tokens: 11.7K IT + 8.2K OT = 19.9K TT

Cost: 0.001$ + 0.003$ = 0.004$

Cite: Hindermann, Marti, Alberto, et al., (2025). RISE-UNIBAS/humanities_data_benchmark, 10.5281/zenodo.16941752

Search Test Runs

Search Results
Show compact results Refine Search New Search

Download JSON Download CSV

Test T0234 at 2025-10-17

Test T0140 at 2025-10-01

Test T0225 at 2025-10-01

Test T0203 at 2025-10-01

Test T0139 at 2025-10-01

Test T0130 at 2025-10-01

Test T0129 at 2025-10-01

Test T0133 at 2025-10-01

Test T0170 at 2025-10-01

Test T0141 at 2025-10-01

Search Test Runs

Search Results Show compact results Refine Search New Search Download Download JSON Download CSV

Test T0234 at 2025-10-17

Test T0140 at 2025-10-01

Test T0225 at 2025-10-01

Test T0203 at 2025-10-01

Test T0139 at 2025-10-01

Test T0130 at 2025-10-01

Test T0129 at 2025-10-01

Test T0133 at 2025-10-01

Test T0170 at 2025-10-01

Test T0141 at 2025-10-01

Search Results
Show compact results Refine Search New Search

Download JSON Download CSV