Search our LLM Benchmark Suite for Humanities Image Data and compare model performance.
Check out our leaderboard and find out which providers, models work best with which kind of data.
We currently have 12 datasets subjected to 1651 tests in 2293 runs.
Search for test-runs and compare results
This is an open and ongoing project. You're welcome to join us!