Unsloth + NVIDIA: LLM Training Gets 2-4x Faster
Unsloth teams up with NVIDIA to speed up large language model training, making it more accessible for developers.

Training large language models (LLMs) can take weeks and require a whole farm of GPUs. But things are about to change. Unsloth, known for its efficient fine-tuning library, has announced a collaboration with NVIDIA.
What's New?
The partnership has optimized key operations like fast RMS normalizations and RoPE kernels, boosting training speed by 2–4x compared to previous versions. Now Unsloth supports not only NVIDIA H100 but also older cards (RTX 3090, 4090, A100) and AMD MI250.
Key Improvements:
- 2-4x speedup on H100.
- Compatibility with various GPU architectures.
- Free access to new kernels.
For developers, this means faster experimentation and lower cloud costs. For startups, it's a chance to run their own LLM without buying expensive hardware.
METABYTE studio comment: Speeding up model training is exactly what anyone working with AI needs. If you're looking to integrate an LLM into your product but worry about costs, these optimizations make the technology more accessible. We keep an eye on trends and are ready to help with integration.
NEXT STEP
Liked the approach?
We apply the same principles to client projects: AI, automation, products that don't die after launch.