I Tried to Make AI Writing Sound Human by Banning AI Words Through logit_bias
I tried to make AI writing sound more human with logit_bias, an API setting that changes how likely a model is to…
Google's latest AI model Gemini 2.5 Pro has emerged as arguably the best-performing AI model to date. This experimental model leapfrogged the previous leader (Anthropic's Claude 3.7 "Sonnet") on major benchmarks.
But who really trusts benchmarks anyway? What's more compelling is that I can attest from my own extensive usage that Gemini 2.5 Pro consistently outperforms Claude 3.7 Sonnet in real-world applications.
In short, Google now claims the AI crown with Gemini 2.5 Pro's superior reasoning, coding, and complex task performance.
What's more significant is how Google achieved this milestone. Unlike most cutting-edge models, Gemini 2.5 Pro doesn't run on NVIDIA's GPUs at all – it runs on Google's own Tensor Processing Units (TPUs). Google has invested years developing these proprietary AI chips, which are fundamental to Gemini's training and inference processes. At the Google Cloud Next '24 event, Google confirmed that the Gemini model was trained entirely on TPUs (specifically the latest "Trillium" TPU v6) and is also served on TPUs for inference. In other words, the world's top AI model completely avoids NVIDIA hardware in favor of Google's in-house silicon.
This is a bold departure from the norm. NVIDIA's GPU accelerators (like the A100 and H100) have been the de facto platform for training advanced AI models across the industry. Yet Google's TPUs have quietly reached parity or better in performance. TPU v5p, Google's current production chip, is roughly 2.8× faster than NVIDIA's flagship H100 for training large models. Google's newest sixth-gen Trillium TPU pushes the envelope even further. In practice, Google can train and fine-tune massive models on its TPU pods without relying on any NVIDIA GPUs. Google has developed the complete AI technology stack internally – from custom chips to software frameworks – making it independent from Nvidia for its AI infrastructure.
If a state-of-the-art model no longer needs NVIDIA chips, it marks a significant shift in the AI industry. For years, NVIDIA has dominated AI data centers with an estimated 70–95% market share in AI accelerators. This dominance has been due not just to powerful GPUs but also to the CUDA software ecosystem and widespread adoption. However, Google's Gemini 2.5 Pro is living proof that an alternative hardware platform can reach the very top. It demonstrates that cutting-edge AI can be achieved without NVIDIA's technology.
The implications for NVIDIA are profound. Google's success with TPU-powered Gemini suggests that Big Tech is willing and able to chart its own path, reducing dependence on NVIDIA. Every TPU that Google deploys is one less GPU NVIDIA sells – and Google deploys a lot of TPUs. In fact, Google has been using TPUs for its AI workloads since 2015 and does "not depend on Nvidia GPUs for the majority of their projects." This erodes the notion that only NVIDIA can deliver the performance needed for the highest-end AI. It also puts pressure on NVIDIA's dominance: if more companies follow Google's lead, NVIDIA could see its de facto monopoly on AI compute begin to slip.
We're already seeing signs of that broader shift. Amazon has developed its own Trainium and Inferentia chips for AWS to reduce reliance on NVIDIA, and Microsoft is likewise investing in custom AI silicon (like the new "Athena"/MAIA accelerator) to avoid over-dependence on GPUs. But Google's TPU strategy is arguably the most advanced example – Google has completely sidestepped NVIDIA for its flagship model, something that would have seemed unthinkable a few years ago. In essence, NVIDIA's once unchallenged grip on AI hardware is now facing credible challengers. The takeaway: the future of AI hardware may be more diversified, with specialized chips (TPUs, ASICs, etc.) sharing the stage with GPUs rather than GPUs singularly ruling.
Perhaps most telling is that even Apple – a company known for end-to-end control of its tech – has quietly leaned on Google's AI hardware. Recent reports reveal that Apple skipped NVIDIA GPUs entirely when training its new internal foundation models. Instead, Apple leveraged thousands of Google's TPU chips to build its advanced AI, as disclosed in an official research paper. Specifically, Apple used TPUv4 and TPUv5 pods (over 8,000 TPU chips) to train its "Apple Foundation Model" for the Apple Intelligence features announced at WWDC 2024. This is a stunning development: Apple essentially went to Google's cloud for the cutting-edge hardware needed, rather than buying NVIDIA GPUs or using another provider.
Apple's partnership (or at least patronage) of Google's TPU infrastructure underscores a broader trend. It signals that Google's AI stack is so compelling that even rival tech giants are willing to use it when it meets their needs. It also highlights a possible industry realignment – for certain AI tasks, Google's TPU-powered cloud is an attractive alternative to NVIDIA-based offerings. If Apple, of all companies, is willing to rely on Google's chips for critical AI workloads, one can imagine other enterprises might follow suit for the performance or efficiency benefits. For NVIDIA, this represents a concerning development: traditional GPU customers might choose alternative technologies if they provide competitive advantages.
The rise of Gemini 2.5 Pro on TPUs is a landmark moment. It proves that the best AI model on the planet can run on non-NVIDIA silicon, shattering the assumption that NVIDIA GPUs are the only game in town for top-tier AI. Google has achieved a full-stack victory – owning the model and the hardware – and in doing so has diminished NVIDIA's perceived necessity. None of this means NVIDIA is suddenly irrelevant (far from it, they remain deeply entrenched and still advancing GPU tech), but it does mean the playing field is evolving.
For the AI industry, this could spark a healthier hardware competition and drive innovation across different platforms. Companies building frontier models now have a precedent for going off the beaten path of CUDA and GPUs. And for NVIDIA, Google's TPU-powered triumph is a wake-up call that the race for AI dominance won't be won on brand legacy alone – it will be won by whoever delivers the most powerful, efficient compute for the next generation of AI. Today, that title belongs to Google's TPU and its Gemini 2.5 Pro model – the best AI model that, tellingly, doesn't run on NVIDIA chips.
Give Vroni a GitHub issue, bug report, spec, or rough idea. It reads the repo, plans the change, writes code, runs checks, and works toward a review-ready pull request.
Take a look at vroni.com