-
GAIA: a benchmark for General AI Assistants
Paper • 2311.12983 • Published • 251 -
Zephyr: Direct Distillation of LM Alignment
Paper • 2310.16944 • Published • 124 -
SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
Paper • 2502.02737 • Published • 261 -
Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
Paper • 2412.03304 • Published • 21
OpenEvals
AI & ML interests
LLM evaluation
Recent Activity
-
Find a leaderboard
🔍160Explore and discover all leaderboards from the HF community
-
YourBench
🚀45Generate custom evaluations from your data easily!
-
Example Leaderboard Template
🥇16Duplicate this leaderboard to initialize your own!
-
Run your LLM evaluations on the hub
🐢2Generate a command to run model evaluations
-
Open-LLM performances are plateauing, let’s make the leaderboard steep again
🏔127Explore and compare advanced language models on a new leaderboard
-
Open LLM Leaderboard
🏆14.1kTrack, rank and evaluate open LLMs and chatbots
-
open-llm-leaderboard/contents
Viewer • Updated • 4.58k • 16.9k • 25 -
open-llm-leaderboard/results
Preview • Updated • 14.2k • 20
-
GAIA: a benchmark for General AI Assistants
Paper • 2311.12983 • Published • 251 -
Zephyr: Direct Distillation of LM Alignment
Paper • 2310.16944 • Published • 124 -
SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
Paper • 2502.02737 • Published • 261 -
Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
Paper • 2412.03304 • Published • 21
-
Find a leaderboard
🔍160Explore and discover all leaderboards from the HF community
-
YourBench
🚀45Generate custom evaluations from your data easily!
-
Example Leaderboard Template
🥇16Duplicate this leaderboard to initialize your own!
-
Run your LLM evaluations on the hub
🐢2Generate a command to run model evaluations
-
Open-LLM performances are plateauing, let’s make the leaderboard steep again
🏔127Explore and compare advanced language models on a new leaderboard
-
Open LLM Leaderboard
🏆14.1kTrack, rank and evaluate open LLMs and chatbots
-
open-llm-leaderboard/contents
Viewer • Updated • 4.58k • 16.9k • 25 -
open-llm-leaderboard/results
Preview • Updated • 14.2k • 20