Chatbot Arena +

OproAI 2026Jun 23

This leaderboard is based on the following benchmarks.

  • Arena - a crowdsourced, randomized battle platform for large language models (LLMs). We use 6M+ user votes to compute Elo ratings.
  • AAII - Artificial Analysis Intelligence Index v3 aggregating 10 challenging evaluations.
  • ARC-AGI - Artificial General Intelligence benchmark v2 to measure fluid intelligence.

🔍

 

Model
Arena Elo
Coding
Vision
AAII
MMLU-Pro
ARC-AGI
Organization
License
🏆Claude Fable 51510156613108191.586AnthropicProprietary
🥇Claude Opus 4.8 Thinking1506156213117790.176AnthropicProprietary
🥇GPT-5.5-high1506156113127689.685OpenAIProprietary
🥇Claude Opus 4.7 Thinking150515601310769075.8AnthropicProprietary
🥇Gemini-3.1-Pro150515311309769177.1GoogleProprietary
🥇Claude Opus 4.8150415581306749062.2AnthropicProprietary
🥇Gemini-3.5-Flash150415351301749172.1GoogleProprietary
🥇Claude Opus 4.71503155413007389.962.1AnthropicProprietary
🥇Claude Opus 4.6 Thinking1503154513047389.769.2AnthropicProprietary
🥇Grok-4.201496151812797289.665.1xAIProprietary
🥇GPT-5.4-high1495153812907388.574OpenAIProprietary
🥇Gemini-3-Pro149215011308739033.6GoogleProprietary
🥇Claude Opus 4.61490153512987189.564.6AnthropicProprietary
🥇GLM-5.2 ✅148815257387.522.8Z.aiMIT
🥇Qwen3.7-Max1486150512897289.6AlibabaProprietary
🥇Muse Spark1485149312947187.3MetaProprietary
🥇Grok-4.1-Thinking14821483708926xAIProprietary
🥇ERNIE-5.1147514957187.1BaiduProprietary
🥇Claude Opus 4.5 (thinking-32k)146915107089.530.6AnthropicProprietary
🥇Claude Sonnet 4.6 Thinking146815111278718860.4AnthropicProprietary
🥈GLM-5.1146715067187.15.1Z.aiMIT
🥈DeepSeek-V4-Pro ✅146714917187.5DeepSeekMIT
🥈Seed2.0 Pro1466149512887087.8ByteDanceProprietary
🥈Kimi-K2.6-Thinking ✅1466149312867187.3MoonshotModified MIT
🥈Qwen3.5-Max146614937087.8AlibabaProprietary
🥈MiMo-V2.5-Pro146614857087XiaomiMIT
🥈GPT-5.2-high1465147012807287.552.9OpenAIProprietary
🥈Gemini-3-Flash146514691292718931.1GoogleProprietary
🥈GPT-5.41465146812757088.429.2OpenAIProprietary
🥈GPT-5.1-high1464146612507087.117.6OpenAIProprietary
🥈GPT-5.21464146512486887.426.7OpenAIProprietary
🥈Grok-4.1146314636888xAIProprietary
🥈Claude Opus 4.5146214966088.87.8AnthropicProprietary
🥈GPT-5.4-mini-high14611461688718.9OpenAIProprietary
🥈Claude Sonnet 4.61460150012776587.337.6AnthropicProprietary
🥈Gemini-2.5-Pro1460146512666386.24.9GoogleProprietary
🥈ERNIE-5.01458146112516586BaiduProprietary
🥈Qwen3.6-Plus145614827088.5AlibabaProprietary
🥈Minimax-M3145214857087.2MiniMaxNon-commercial
🥈GLM-51452146170875Z.aiMIT
🥈Kimi-K2.5-Thinking1451148012716987.111.8MoonshotModified MIT
🥈Qwen3.5-397B-A17B145014636987.85AlibabaApache 2.0
🥉Gemma-4-31B-it144914626385.2GoogleApache 2.0
🥉GLM-4.7144514606885.6Z.aiMIT
🥉DeepSeek-V4-Flash144514606486.2DeepSeekMIT
🥉GPT-5-high1444146012326687.19.9OpenAIProprietary
🥉Qwen3-Max144314686485.3AlibabaProprietary
🥉Gemma-4-26B-A4B-it144314605682.6GoogleApache 2.0
🥉Grok-41442145312216686.616xAIProprietary
🥉GLM-4.6144114585683.5Z.aiMIT
🥉GPT-5.114401450124366877.5OpenAIProprietary
🥉Kimi-K2-Thinking143814506784.8MoonshotModified MIT
🥉Claude Sonnet 4.5 (thinking-32k)143214856187.513.6AnthropicProprietary
🥉MiMo-V2-Pro143114416986.8XiaomiProprietary
🥉GLM-4.5143014485483.5Z.aiMIT
🥉Qwen3-VL-235B-A22B-Instruct1429145712464982.8AlibabaApache 2.0
🥉Mistral Large 3142814504081MistralApache 2.0
🥉ChatGPT-4o-latest (2025-03-26)1427143412423880.3OpenAIProprietary
🥉DeepSeek-R1-0528142614365784.91.3DeepSeekMIT
🥉Claude Opus 4.1 (thinking-16k)142514755987.8AnthropicProprietary
🥉o3-2025-04-161424144112186585.36.5OpenAIProprietary
🥉Grok-3-Preview-02-24142314394479.9xAIProprietary
🥉Hunyuan-Hy3142214486686.2TencentTencent
🥉DeepSeek-V3.2-Thinking142214386686.24DeepSeekMIT
🪙Gemini-3.1-Flash-Lite1421140212306086.2GoogleProprietary
🪙Claude Sonnet 4.5142014644986AnthropicProprietary
🪙LongCat-Flash-Chat142014614982.7MeituanMIT
🪙Claude Opus 4.1141914654787.3AnthropicProprietary
🪙Grok-4-Fast141914416285xAIProprietary
🪙Qwen3-235B-A22B-Instruct-2507141814574982.81.3AlibabaApache 2.0
🪙Nemotron-3-Ultra-550B-A55B141814527086.8NvidiaNvidia Open
🪙DeepSeek-V3.1-Thinking141814375885.1DeepSeekMIT
🪙DeepSeek-V3.2141814315283.7DeepSeekMIT
🪙Qwen3-Next-80B-A3B-Instruct141714565782.4AlibabaApache 2.0
🪙Amazon-Nova-Chat-11-10141714316283AmazonProprietary
🪙DeepSeek-V3.1141714304783.3DeepSeekMIT
🪙Minimax-M2.7141614486987.1MiniMaxNon-commercial
🪙Qwen3-235B-A22B-Thinking-2507141614426284.3AlibabaApache 2.0
🪙GPT-4.5-Preview14151419119040810.8OpenAIProprietary
🪙GPT-5-chat1413142112386586.67.5OpenAIProprietary
🪙Gemini-2.5-Flash1412142012355683.22.5GoogleProprietary
🪙Qwen3-VL-235B-A22B-Thinking1411143212156284.3AlibabaApache 2.0
🪙Hunyuan-Vision-1.5-Thinking14101419122086.2TencentProprietary
🪙Minimax-M2.51408143665874.9MiniMaxModified MIT
🪙MiMo-V2-Flash140214116184.8XiaomiMIT
🪙Mistral Medium 3.1140114123877.2MistralProprietary
🪙Hunyuan-T1-202507111400140986.2TencentProprietary
🪙Minimax-M2.1139914306487MiniMaxModified MIT
🪙MAI-1-Preview13981406MicrosoftProprietary
🪙Gemini-2.0-Pro-Exp-02-051398139611703680.5GoogleProprietary
🪙Gemini-2.0-Flash-Thinking-Exp-01-211397138312084079.8GoogleProprietary
🪙Step-3.5-Flash13871433StepFunApache 2.0
🪙GLM-4.5-Air138614104781.5Z.aiMIT
🪙Qwen3-30B-A3B-Instruct-2507138214254477.7AlibabaApache 2.0
🪙Kimi-K2-0905-Preview138214034982.4MoonshotModified MIT
🪙Qwen-VL-Max-2025-08-13138114401213AlibabaProprietary
🪙GPT-4.1-2025-04-141381139612074580.60.4OpenAIProprietary
🪙Kimi-K2-0711-Preview138014024782.4MoonshotModified MIT
🪙Claude Haiku 4.5137814364280AnthropicProprietary
🪙DeepSeek-V3-0324137713914281.9DeepSeekMIT
🪙Hunyuan-Turbos-202504161377139078TencentProprietary
🪙Claude Opus 4 (thinking-16k)1376143411885787.38.6AnthropicProprietary
🪙GPT-5-mini1375141912146282.84.4OpenAIProprietary
🪙DeepSeek-R1137313824884.41.3DeepSeekMIT
🪙Gemini-2.0-Flash-Exp1370137111913678.21.3GoogleProprietary
🪙Qwen3-235B-A22B136913944682.8AlibabaApache 2.0
🪙Mistral Medium 31369138711593776MistralProprietary
🪙gpt-oss-120b136813985980.8OpenAIApache 2.0
🪙Qwen2.5-Max136713733276.2AlibabaProprietary
🪙Claude Opus 413661405117345861.3AnthropicProprietary
🪙Grok-3-mini-high136613805682.8xAIProprietary
🪙o1-2024-12-171366137811665084.11.3OpenAIProprietary
🪙o4-mini-2025-04-161362138511926383.26.1OpenAIProprietary
🪙Step-3136014001190StepFunProprietary
🪙Qwen3-Coder-480B-A35B-Instruct135814064378.8AlibabaApache 2.0
🪙Nemotron-3-Super-120B-A12B135714016083.7NvidiaNvidia Open
🪙Gemma-3-27B-it1356135011632366.9GoogleGemma
🪙INTELLECT-313551376Prime IntellectMIT
🪙Claude Sonnet 4 (thinking-32k)1351141211875784.25.9AnthropicProprietary
🪙Minimax-M1135113695181.6MiniMaxApache 2.0
🪙Qwen3-32B134213764279.8AlibabaApache 2.0
🪙Llama-3.3-Nemotron-Super-49B-v1.5134013595081.4NvidiaNvidia Open
🪙Step-1o-Turbo-202506133913611182StepFunProprietary
🪙o3-mini-high133813805380.23OpenAIProprietary
🪙GPT-4.1-mini-2025-04-141338137011774078.1OpenAIProprietary
🪙Gemini-2.5-Flash-Lite1337136212054275.9GoogleProprietary
🪙Mistral-Small-3.2-25061337136111453068.1MistralApache 2.0
🪙Claude Sonnet 41335138411704483.71.3AnthropicProprietary
🪙Gemma-3-12B-it133513102259.5GoogleGemma
🪙DeepSeek-V3133413373375.2DeepSeekDeepSeek
🪙GPT-5-nano1333136311665277.22.6OpenAIProprietary
🪙QwQ-32B133213514676.4AlibabaApache 2.0
🪙GLM-4-Plus-01111332131078.6Z.aiProprietary
🪙Gemini-2.0-Flash-Lite1330133810972872.4GoogleProprietary
🪙Qwen-Plus-012513271339AlibabaProprietary
🪙Command A (03-2025)132713363071.2CohereCC-BY-NC-4.0
🪙Amazon-Nova-Chat-05-14132413373373.3AmazonProprietary
🪙Llama-3.1-Nemotron-Ultra-253B-v1132113454482.5NvidiaNvidia Open
🪙Step-2-16K-Exp13211313StepFunProprietary
🪙Qwen3-30B-A3B132013464077.7AlibabaApache 2.0
🪙Gemini-1.5-Pro-00213201311115832750.8GoogleProprietary
🪙o1-mini131813664174.20.8OpenAIProprietary
🪙o3-mini131813615179.12.1OpenAIProprietary
🪙Claude 3.7 Sonnet (thinking-32k)1316135511674583.70.9AnthropicProprietary
🪙gpt-oss-20b131513714773.6OpenAIApache 2.0
🪙Hunyuan-Turbo-011013141335TencentProprietary
🪙Llama-3.3-Nemotron-Super-49B-v1131013203878.5NvidiaNvidia Open
🪙OLMo-3-32b-think130613273675.9Ai2Apache-2.0
🪙Grok-2-08-13130512982670.9xAIGrok 2
🪙Gemma-3n-e4b-it130412971648.8GoogleGemma
🪙Yi-Lightning1303132101 AIProprietary
🪙GPT-4o-2024-05-131302130711342874.8OpenAIProprietary
🪙Claude 3.7 Sonnet1301134111453580.3AnthropicProprietary
🪙Claude 3.5 Sonnet (20241022)1299134011223177.2AnthropicProprietary
🪙Deepseek-v2.5-1210129613162267.2DeepSeekDeepSeek
🪙Athene-v2-Chat-72B12941320NexusFlowNexusFlow
🪙Gemma-3-4B-it129312651241.7GoogleGemma
🪙Llama-4-Maverick-17B-128E-Instruct1292131211354080.9MetaLlama 4
🪙GLM-4-Plus1292130170.2Z.aiProprietary
🪙Hunyuan-Large-2025-02-1012911311TencentProprietary
🪙Gemini-1.5-Flash-0021290127311372668GoogleProprietary
🪙GPT-4o-mini-2024-07-181289130010632264.8OpenAIProprietary
🪙GPT-4.1-nano-2025-04-141287131210602865.7OpenAIProprietary
🪙Llama-3.1-405B-Instruct-bf16128612992773.2MetaLlama 3.1
🪙Llama-3.1-Nemotron-70B-Instruct128512892469NvidiaLlama 3.1
🪙Qwen-Max-091912841296AlibabaProprietary
🪙Llama-3.1-405B-Instruct-fp8128412922773.2MetaLlama 3.1
🪙Yi-Lightning-lite1284128601 AIProprietary
🪙Claude 3.5 Sonnet (20240620)1283130911172775.1AnthropicProprietary
🪙Grok-2-mini-08-1312831279xAIProprietary
🪙Llama-4-Scout-17B-16E-Instruct1276129011263175.2MetaLlama 4
🪙Hunyuan-Standard-2025-02-1012761289TencentProprietary
🪙Llama-3.3-70B-Instruct127612792971.3MetaLlama 3.3
🪙Deepseek-v2.5127513062166.2DeepSeekDeepSeek
🪙GPT-4-Turbo-2024-04-091275128010872669.4OpenAIProprietary
🪙Qwen2.5-72B-Instruct127213022772AlibabaQwen
🪙Hunyuan-Large-Vision127012981183TencentProprietary
🪙Mistral-Small-3.1-24B-Instruct-25031269129511212265.9MistralApache 2.0
🪙Mistral-Large-2411126912842569.7MistralMRL
🪙Athene-70B12681274NexusFlowCC-BY-NC-4.0
🪙GPT-4-1106-preview126712692363.7OpenAIProprietary
🪙GPT-4-0125-preview12661261OpenAIProprietary
🪙Claude 3 Opus1265126910202269.6AnthropicProprietary
🪙Llama-3.1-70B-Instruct126512682267.6MetaLlama 3.1
🪙Amazon Nova Pro 1.0126212829792769.1AmazonProprietary
🪙Llama-3.1-Tulu-3-70B12601251Ai2Llama 3.1
🪙Claude 3.5 Haiku (20241022)1256128710952163.4AnthropicProprietary
🪙magistral-medium-2506125313073675.3MistralProprietary
🪙Reka-Core-202409041252123820Reka AIProprietary
🪙Reka-Core-2024072212501226Reka AIProprietary
🪙Qwen-Plus-082812421263AlibabaProprietary
🪙Jamba-1.5-Large124212441657.2AI21 LabsJamba Open
🪙Deepseek-v2-API-062812401260DeepSeekDeepSeek
🪙Mistral-Small-3-24B-Instruct-2501123812512265.2MistralApache 2.0
🪙Deepseek-Coder-v2-0724123712861558.5DeepSeekDeepSeek
🪙Yi-Large123612381458.601 AIProprietary
🪙Gemma-2-27B-it123612261857.5GoogleGemma
🪙Qwen2.5-Coder-32B-Instruct123512792363.5AlibabaApache 2.0
🪙Amazon Nova Lite 1.0123312539892359AmazonProprietary
🪙Gemma-2-9B-it-SimPO12331211PrincetonMIT
🪙Command R+ (08-2024)12331200743.2CohereCC-BY-NC-4.0
🪙Gemini-1.5-Flash-8B-0011231122810421756.9GoogleProprietary
🪙Llama-3.1-Nemotron-51B-Instruct12311227NvidiaLlama 3.1
🪙GLM-4-052012301237Z.aiProprietary
🪙Nemotron-4-340B-Instruct12291220NvidiaNvidia Open
🪙Aya-Expanse-32B12291211637.7CohereCC-BY-NC-4.0
🪙Reka-Flash-202409041225120820Reka AIProprietary
🪙Llama-3-70B-Instruct122412161457.4MetaLlama 3
🪙Claude 3 Sonnet122312329831457.9AnthropicProprietary
🪙OLMo-2-0325-32B-Instruct122312151451.1Ai2Apache-2.0
🪙Phi-4122212422671.4MicrosoftMIT
🪙Reka-Flash-2024072212181201Reka AIProprietary
🪙Amazon Nova Micro 1.0121512281853.1AmazonProprietary
🪙Gemma-2-9B-it12131194849.5GoogleGemma
🪙Hunyuan-Standard-256K12091244TencentProprietary
🪙Command R+ (04-2024)12091184642.7CohereCC-BY-NC-4.0
🪙Qwen2-72B-Instruct120812061962.2AlibabaQianwen
🪙Claude 3 Haiku120012089501050AnthropicProprietary
🪙Llama-3.1-Tulu-3-8B12001197Ai2Llama 3.1
🪙Qwen-Max-042811991208AlibabaProprietary
🪙Ministral-8B-241011981219838.9MistralMRL
🪙GLM-4-011611981209Z.aiProprietary
🪙DeepSeek-Coder-V2-Instruct11961259DeepSeekDeepSeek
🪙Command R (08-2024)11951180133.8CohereCC-BY-NC-4.0
🪙Llama-3.1-8B-Instruct119312031047.6MetaLlama 3.1
🪙Jamba-1.5-Mini11931197AI21 LabsJamba Open
🪙Aya-Expanse-8B11931184231.2CohereCC-BY-NC-4.0
🪙Qwen1.5-110B-Chat1180119211AlibabaQianwen
🪙Yi-1.5-34B-Chat1178118101 AIApache-2.0
🪙Claude-111781161AnthropicProprietary
🪙Qwen1.5-72B-Chat11721175AlibabaQianwen
🪙Mistral Medium11711172949.1MistralProprietary
🪙Llama-3-8B-Instruct11711164740.5MetaLlama 3
🪙Command R (04-2024)11691141133.7CohereCC-BY-NC-4.0
🪙InternLM2.5-20B-chat11681179InternLMOther
🪙Mixtral-8x22b-Instruct-v0.1116811751253.7MistralApache 2.0
🪙Gemma-2-2b-it11631130GoogleGemma
🪙Granite-3.1-8B-Instruct11581191IBMApache 2.0
🪙Claude-2.011581160948.6AnthropicProprietary
🪙Gemini-1.0-Pro-00111551125GoogleProprietary
🪙Zephyr-ORPO-141b-A35b-v0.111501144HuggingFaceApache 2.0
🪙Claude-2.1114611581049.5AnthropicProprietary
🪙GPT-3.5-Turbo-061311451164946.2OpenAIProprietary
🪙Qwen1.5-32B-Chat11441163AlibabaQianwen
🪙Phi-3-Medium-4k-Instruct114411461154.3MicrosoftMIT
🪙Starling-LM-7B-beta11391151NexusflowApache-2.0
🪙Mixtral-8x7B-Instruct-v0.111381136338.7MistralApache 2.0
🪙GPT-3.5-Turbo-031411381136OpenAIProprietary
🪙Granite-3.1-2B-Instruct11361166IBMApache 2.0
🪙Qwen1.5-14B-Chat11351144AlibabaQianwen
🪙Claude-Instant-111351136143.4AnthropicProprietary
🪙Yi-34B-Chat1134112901 AIYi
🪙Tulu-2-DPO-70B11271120Ai2Ai2 ImpACT
🪙DBRX-Instruct-Preview11261141DatabricksDBRX
🪙WizardLM-70B-v1.011261093MicrosoftLlama 2
🪙Llama-2-70B-chat11221099640.7MetaLlama 2
🪙Nous-Hermes-2-Mixtral-8x7B-DPO11191103NousResearchApache-2.0
🪙Llama-3.2-3B-Instruct11181097534.7MetaLlama 3.2
🪙Phi-3-Small-8k-Instruct11171123MicrosoftMIT
🪙OpenChat-3.5-010611141119OpenChatApache-2.0
🪙Starling-LM-7B-alpha11141104UC BerkeleyCC-BY-NC-4.0
🪙Vicuna-33B11131091LMSYSNon-commercial
🪙DeepSeek-LLM-67B-Chat111111066DeepSeekDeepSeek
🪙Snowflake Arctic Instruct11091101SnowflakeApache 2.0
🪙Granite-3.0-8B-Instruct11081115IBMApache 2.0
🪙NV-Llama2-70B-SteerLM-Chat11061047NvidiaLlama 2
🪙OpenChat-3.511031077OpenChatApache-2.0
🪙Gemma-1.1-7B-it11021105GoogleGemma
🪙OpenHermes-2.5-Mistral-7B11001083NousResearchApache-2.0
🪙pplx-70B-online10991055Perplexity AIProprietary
🪙Mistral-7B-Instruct-v0.210971094124.5MistralApache-2.0
🪙Llama-2-13b-chat10931077640.6MetaLlama 2
🪙Granite-3.0-2B-Instruct10911104IBMApache 2.0
🪙SOLAR-10.7B-Instruct-v1.010911073Upstage AICC-BY-NC-4.0
🪙Qwen1.5-7B-Chat10901110AlibabaQianwen
🪙Phi-3-Mini-4K-Instruct-June-2410881098MicrosoftMIT
🪙Dolphin-2.2.1-Mistral-7B10881049CognitiveApache-2.0
🪙WizardLM-13b-v1.210841048MicrosoftLlama 2
🪙Phi-3-Mini-4k-Instruct10821102MicrosoftMIT
🪙MPT-30B-chat10761055MosaicMLCC-BY-NC-SA-4.0
🪙Zephyr-7B-beta10761053HuggingFaceMIT
🪙CodeLlama-34B-instruct10731065MetaLlama 2
🪙Llama-3.2-1B-Instruct10671063120MetaLlama 3.2
🪙Qwen2.5-VL-32B-Instruct1149AlibabaApache 2.0
🪙Step-1o-Vision-32k (highres)1119StepFunProprietary
🪙Qwen2.5-VL-72B-Instruct1104AlibabaQwen
🪙Pixtral-Large-241110882470.1MistralMRL
🪙Qwen-VL-Max-11191056AlibabaProprietary
🪙Qwen2-VL-72b-Instruct1044AlibabaQwen
🪙Step-1V-32K1043StepFunProprietary
🪙Molmo-72B-09241010Ai2Apache 2.0
🪙Pixtral-12B-24091006947.3MistralApache 2.0
🪙InternVL2-26B1003OpenGVLabMIT
🪙Llama-3.2-90B-Vision-Instruct9972067.1MetaLlama 3.2
🪙Hunyuan-Standard-Vision-2024-12-31996TencentProprietary
🪙Aya-Vision-32B992CohereCC-BY-NC-4.0
🪙Qwen2-VL-7B-Instruct988AlibabaApache 2.0
🪙Yi-Vision97801 AIProprietary
🪙Llama-3.2-11B-Vision-Instruct9641146.4MetaLlama 3.2
🪙Molmo-7B-D-0924957Ai2Apache 2.0

OproAI

SWE-bench +

💡 AAII v3 incorporates 10 evaluations: MMLU-Pro, Humanity’s Last Exam, AA-LCR, GPQA Diamond, AIME, IFBench, SciCode, LiveCodeBench, Terminal-Bench Hard, and 𝜏²-Bench Telecom.

💻 Arena Elo ratings are computed by this notebook. Higher values are better for all benchmarks. Empty cells mean not available.

Transition from online Elo rating system to Bradley-Terry model

We adopted the Elo rating system for ranking models since the launch of the Arena. It has been useful to transform pairwise human preference to Elo ratings that serve as a predictor of winrate between models. Specifically, if player A has a rating of RA and player B a rating of RB, the probability of player A winning is

{\displaystyle E_{\mathsf {A}}={\frac {1}{1+10^{(R_{\mathsf {B}}-R_{\mathsf {A}})/400}}}~.}

ELO rating has been used to rank chess players by the international community for over 60 years. Standard Elo rating systems assume a player’s performance changes overtime. So an online algorithm is needed to capture such dynamics, meaning recent games should weigh more than older games. Specifically, after each game, a player’s rating is updated according to the difference between predicted outcome and actual outcome.

This algorithm has two distinct features:

  1. It can be computed asynchronously by players around the world.
  2. It allows for players performance to change dynamically – it does not assume a fixed unknown value for the players rating.

This ability to adapt is determined by the parameter K which controls the magnitude of rating changes that can affect the overall result. A larger K essentially put more weight on the recent games, which may make sense for new players whose performance improves quickly. However as players become more senior and their performance “converges” then a smaller value of K is more appropriate. As a result, USCF adopted K based on the number of games and tournaments completed by the player (reference). That is, the Elo rating of a senior player changes slower than a new player.

When we launched the Arena, we noticed considerable variability in the ratings using the classic online algorithm. We tried to tune the K to be sufficiently stable while also allowing new models to move up quickly in the leaderboard. We ultimately decided to adopt a bootstrap-like technique to shuffle the data and sample Elo scores from 1000 permutations of the online plays. This provided consistent stable scores and allowed us to incorporate new models quickly.

In the context of LLM ranking, there are two important differences from the classic Elo chess ranking system. First, we have access to the entire history of all games for all models and so we don’t need a decentralized algorithm. Second, most models are static (we have access to the weights) and so we don’t expect their performance to change. However, it is worth noting that the hosted proprietary models may not be static and their behavior can change without notice. We try our best to pin specific model API versions if possible.

To improve the quality of our rankings and their confidence estimates, we are adopting another widely used rating system called the Bradley–Terry (BT) model. This model actually is the maximum likelihood (MLE) estimate of the underlying Elo model assuming a fixed but unknown pairwise win-rate. Similar to Elo rating, BT model is also based on pairwise comparison to derive ratings of players to estimate win rate between each other. The core difference between BT model vs the online Elo system is the assumption that player’s performance does not change (i.e., game order does not matter) and the computation takes place in a centralized fashion.