top of page
Search

AI Ranking Series Takeaways: What We Learned Testing ChatGPT, Claude, and Gemini

  • GridironIQ
  • Aug 9
  • 4 min read


Over the past 2 months at GridironIQ, I ran an experiment across our social channels to test how artificial intelligence understands the NFL. I took three of the most advanced AI models on the market—ChatGPT, Claude, and Gemini—and tasked them with ranking the top 10 players at every position in the NFL right now.


After compiling the lists, running the numbers, and cross-referencing their outputs with actual expert consensus, the results were eye-opening.


Here are the biggest takeaways from our AI ranking series, how each model approaches football, and what this tells us about the current state of AI in sports analysis.


The Big Takeaways

1. Claude Has Zero Ball Knowledge

Let’s not beat around the bush: Claude struggled mightily.

Throughout the series, Anthropic’s model consistently put players on the wrong teams, leaned on heavily outdated rosters, and left off clear top-tier talent in favor of players who haven't been elite in years. If you were drafting a fantasy team or building a roster based on Claude's board, you’d be starting guys who were traded two seasons ago.


2. No AI Compares to a Real NFL Analyst

If there’s one overriding takeaway from this entire experiment, it’s that AI is not replacing real human tape-watchers anytime soon. Every single model made blunders at some point in the series. Whether it was hallucinating team rosters, misinterpreting player positions, or serving up objectively terrible rankings (like ranking backup-level talent over All-Pros), none of these models possess true "football instinct." They don't watch game film; they process text.


3. Gemini Came Out on Top (85% Accuracy)

Out of the three models tested, Gemini was comfortably the best. When cross-referenced against consensus rankings from top NFL analysts and film evaluators, Gemini proved to be the most accurate, posting an 85% accuracy rating across all positional top-10 lists. Its ability to pull real-time information kept its rosters clean and its rankings grounded in current performance.


How Each AI Builds Its NFL Rankings

Why did the three models produce such wildly different lists? It comes down to how each AI is built, trained, and updated. Here is a breakdown of the three distinct "personas" that emerged during the series.


1. Gemini: The "Recency & Momentum" Tracker

  • Primary Bias: Live context, recent season volume, and current headlines.


  • How It Ranks Players: Gemini heavily weighs the most recent 1–2 seasons. If a player had a dominant breakout campaign last year, Gemini immediately rewards them with a top spot over an established veteran who had a down year or suffered an injury.


  • Why This Happens: Gemini is directly integrated with live Google Search indexing. When prompted for a "top 10 right now," it pulls heavily from fresh articles, recent PFF grades, and live injury reports rather than static historical datasets.


  • NFL Ranking Impact:


    • Quick on Breakouts: Rapidly elevated young stars into top-tier spots faster than the other models.


    • Penalizes Missed Time: Swiftly demoted elite veterans who missed significant time due to injuries in the prior season.


2. ChatGPT: The "Consensus & Media Hype" Standard Bearer

  • Primary Bias: Mainstream consensus, name recognition, and career accolades.


  • How It Ranks Players: ChatGPT produces lists that closely mirror mainstream sports networks (ESPN, NFL Network, or Madden ratings). It defaults to safety, placing heavy emphasis on Pro Bowl appearances, All-Pro selections, market size, and playoff reputation.


  • Why This Happens: ChatGPT’s training favors high-probability text combinations representing public consensus. It reflects what the majority of the internet writes about a player rather than raw underlying efficiency.


  • NFL Ranking Impact:


    • Big-Market Favoritism: Heavily favors stars on high-visibility teams (Eagles, Cowboys, Chiefs, 49ers) where online discussion volume is highest.


    • Sticky Legacy Standings: Slower to drop declining veterans because years of positive career text still dominate its training data.


    • Plays It Safe: Unanimously safe at positions #1 through #3, rarely taking bold risks at the top.


3. Claude: The "Systemic & Efficiency" Analyst

  • Primary Bias: Methodological criteria, advanced efficiency, and positional nuance.


  • How It Ranks Players: Claude attempts to evaluate players based on role-specific metrics (like EPA/play, pass-block win rate, or pressure rate) rather than simple box-score totals.


  • Why This Happens: Anthropic’s fine-tuning emphasizes structured reasoning and multi-variable logic over quick search lookups. It tries to weigh context—like how a poor offensive line affects a quarterback's raw stats.


  • NFL Ranking Impact:


    • Better on the Line: More logical at non-stat-heavy positions like Offensive Line (OT/OG/C) where raw box scores don't exist.


    • Efficiency Over Volume: More willing to rank a hyper-efficient player on a bad team over a high-volume player on a good team.


    • Roster Disconnect: Because it lacks live web grounding in standard prompts, it consistently suffered from hallucinated rosters and outdated team assignments.


At a Glance: Model Comparison Matrix

Dimension

ChatGPT (OpenAI)

Gemini (Google)

Claude (Anthropic)

Model Archetype

The Consensus Historian

The Real-Time Newsroom

The Analytical Scout

Primary Driver

Public sentiment, high-frequency web mentions, lifetime accolades

Live news context, current-season stats, recent performance trends

Scheme fit, advanced efficiency metrics, contextual logic

Young Breakout Players

Lagging: Holds them in the #7–10 range until multi-year consensus builds

Aggressive: Rapidly shoots mid-season breakout stars into top spots

Measured: Elevates them based on per-snap efficiency rather than raw volume

Injured / Aging Veterans

Sticky: Retains star veterans top-side due to years of positive training text

Penalizing: Swiftly drops or omits players missing extended playing time

Adjusted: Keeps them high if underlying per-snap metrics remained elite

Non-Stat Positions (OL / DB)

Name-Driven: Relies heavily on Pro Bowl and All-Pro name recognition

Team-Driven: Evaluates players based on their overall team unit rankings

Trait-Driven: Weighs individual metrics like Pass Block Win Rate

Main System Bias

Legacy & Hype: Overweights big-market teams and media narratives

Recency Bias: Overreacts to small sample sizes or short hot streaks

Metric Rigidness: Can miss real-world context and hallucinate team rosters

Analyst Equivalent

The TV Panelist (ESPN / Madden Ratings)

The Fantasy Manager (Waiver-wire focus & recent form)

The Front Office Evaluator (PFF / Advanced Analytics)

Final Verdict

While AI can be a useful tool for aggregating stats, generating base comparisons, or spark-plugging debate, it cannot replace human football evaluation.


If you're looking for real-time accuracy, Gemini takes the crown among AI models due to its search integration. But at the end of the day, no algorithm can substitute for watching the film, understanding scheme context, and knowing the game from a human perspective.

 
 
 

Comments


Stay connected with GridironIQ for the latest football analytics and trends.

  • Instagram
  • TikTok

224-334-8387

GridironIQ

© 2035 by GridironIQ. Powered and secured by Wix

bottom of page