AI Model ELO Tracker Reveals How Flagship Models Degrade Over Time
Noticed your favorite AI model feels dumber after a few weeks? This ELO tracker proves it's not just you — it's a measurable trend.

Every developer knows that feeling: a shiny new AI model launches and it's a beast, but weeks later it seems... off. One builder decided to quantify this phenomenon and created a live dashboard tracking historical ELO ratings for flagship models.
Instead of a spaghetti chart of every variant, the tracker intelligently plots a single continuous curve per major AI lab. It clearly shows both the generational leaps and the slow performance decay — the infamous "nerfing." Yes, folks, your gut feeling is data-driven: models do get dumber over time.
The dashboard is mobile-friendly and includes dark mode. But the creator admits a blind spot: Arena AI mostly tests API endpoints, while consumer chat UIs often add heavy system prompts, safety wrappers, or silently switch to quantized models under load. API benchmarks miss this "nerfing" that everyday users experience.
The author is looking for historical ELO datasets that scrape outputs from consumer web UIs rather than raw APIs. Got leads? The project is open source — contributions welcome.
METABYTE studio comment: We've noticed models forgetting their skills over time too — just like devs after a vacation. If you need your AI to stay sharp (and not act like a student during finals), hit us up for monitoring setups.
NEXT STEP
Liked the approach?
We apply the same principles to client projects: AI, automation, products that don't die after launch.