ZAYA1-8B: 8B MoE Model with 760M Active Params Matches DeepSeek-R1 on Math
Firethering's new open-source ZAYA1-8B model delivers impressive math and coding performance, rivaling much larger models like DeepSeek-R1.

The open-source AI landscape just got a spicy new addition: ZAYA1-8B by Firethering. It's an 8-billion-parameter MoE (Mixture of Experts) model, but here's the kicker — it only activates 760 million parameters per query. Sounds like a cheat code? Kinda works like one.
On AIME 2024 and MATH-500 benchmarks, ZAYA1-8B scores comparable to DeepSeek-R1 (671B params) and outperforms Qwen2.5-Math-7B-Instruct. Coding benchmarks? LiveCodeBench and Codeforces — solid results.
Why developers should care:
- Efficiency: 760M active params vs 8B means major savings in memory and inference speed.
- Open-source: fine-tune it, customize it, embed it in your product.
- Specialization: optimized for math and code — perfect for EdTech, ML pipelines, or dev tools.
The model is already on Hugging Face, with chat (Instruct) and base versions. Future plans include long-context support and multimodality.
METABYTE studio's take: ZAYA1-8B is a textbook example of how smart architecture (MoE) can deliver big performance without breaking the bank. For startups looking to add a math/coding AI assistant, this is a real alternative to expensive APIs. We'd definitely consider integrating it into our own products.
NEXT STEP
Liked the approach?
We apply the same principles to client projects: AI, automation, products that don't die after launch.