1 min readfrom Machine Learning

Evaluated 6 frontier LLMs (GPT-5.4, Claude Sonnet 4.6, Claude Opus 4.7, Gemini Pro/Flash, Grok 4.3) on political, gender, and racial bias across 8 benchmarks (~20,600 examples) [R]

I ran a solo evaluation project benchmarking six current frontier models: GPT-5.4, Claude Sonnet 4.6, Claude Opus 4.7, Gemini Pro, Gemini Flash, and Grok 4.3. I tested tham across 8 established bias/fairness datasets (WinoBias, BBQ Race/Ethnicity, SeeGULL, OpinionsQA, cajcodes Political Bias, Hyperpartisan News, Political Compass).

On PoliticalCompass, I found that all LLMs were left leaning except Grok, but across other Political Bias benchmarks, all six LLMs leaned left, including Grok. So Grok self-reports as right-leaning but behaves left-leaning when actually classifying content or answering policy questions.

Another interesting result I found is that the Refusal behavior on BBQ race data was interesting. On questions that involved race, and the correct answer must be answered with race, GPT-5.4 refused 20.3% of the time, Claude Opus 4.7 13.8%, Grok 9.5%, Claude Sonnet 4.6 and Gemini Pro ~5%.

Limitations: solo, non-peer-reviewed project. No multi-run averaging on every dataset, single prompt template per task.

Full data, per-model breakdowns, and methodology: https://www.civicsparklearning.org/ai-nonprofit-dashboard

submitted by /u/marggggggggg
[link] [comments]

Want to read more?

Check out the full article on the original site

View original article

Tagged with

#LLMs
#Bias
#Fairness
#GPT-5.4
#Political Bias
#Claude Opus
#Racial Bias
#Benchmarking
#Claude Sonnet
#Gender Bias
#Gemini Pro
#Grok
#BBQ Race/Ethnicity
#Political Compass
#Gemini Flash
#WinoBias
#Refusal
#SeeGULL
#OpinionsQA
#Hyperpartisan News