1 min readfrom Machine Learning

GoBench: Evaluating LLMs on the game of Go [R]

GoBench: Evaluating LLMs on the game of Go [R]
GoBench: Evaluating LLMs on the game of Go [R]

GoBench evaluates LLMs on 9x9 Go games against a ladder of KataGo opponents, from random to superhuman.

It measures general reasoning ability, strongly correlates with ARC-AGI 2 (r=0.83 correlation), and remains highly unsaturated.

GPT-6 Astra max achieves 2500 Elo, much lower than the best KataGo, which achieves 4400 Elo.

With coding tools and two hours of preparation before evaluation, Codex with Astra achieves 3560 Elo.

I will keep the leaderboard updated as long as it is not saturated.

leaderboard: https://rolandgao.com/blog/gobench/

code: https://github.com/RolandGao/gobench

paper: https://github.com/RolandGao/gobench/blob/main/paper/gobench2.pdf

x: https://x.com/Roland65821498/status/2100253388562723298?s=20

submitted by /u/Roland31415
[link] [comments]

Want to read more?

Check out the full article on the original site

View original article

Tagged with

#GoBench
#LLMs
#Go
#KataGo
#Elo
#Reasoning
#GPT-6 Astra
#Codex
#ARC-AGI 2
#Unsaturated
#9x9 Go
#Superhuman
#Machine Learning
#Coding Tools
#Leaderboard
#General Ability
#Correlation
#Evaluation
#Preparation