For months, the AI race has looked like a battle between companies racing to build the smartest models. Behind the scenes, another contest has been unfolding who gets access to the computing power, models, and infrastructure needed to compete. Google’s decision to limit Meta’s access to Gemini is a reminder that the AI industry isn’t just about building better technology. Increasingly, it’s about deciding who gets to use it, on what terms, and in what quantities.


The story that broke in the Financial Times on June 28 reads, on the surface, like a routine supplier capacity issue. Google couldn’t meet Meta’s demand, so access got capped. Employees were told to use tokens more efficiently. Inconvenient, but hardly dramatic. The more interesting story is what this episode reveals about how frontier AI actually works as an economic and competitive system and why the companies that control the compute are accumulating a different kind of power than the companies building the models that run on it.

Google’s Relationship With Meta Has Changed

The relationship between Google and Meta has never been straightforward. They’ve competed for digital advertising revenue for years, operating in overlapping markets for audience attention and marketing budgets. What’s newer is that they’ve become simultaneously competitors and customers of each other. That dynamic has no clean precedent in the technology industry’s history.


Meta had been relying on Gemini, which proved more capable than its own open source Llama models for certain tasks, to automate safety processes such as removing harmful content. (Semantic Scholar) That’s a significant admission embedded in a dry infrastructure story. Meta a company that has invested billions in its own Llama family of open source models and positioned itself as a serious frontier AI player was quietly depending on a direct competitor’s model for core operational functions. Content moderation at Meta’s scale is not a peripheral experiment. It’s fundamental infrastructure. The fact that Gemini was embedded there says something about the performance gap between what Meta has built internally and what it needed for production critical applications.

When Rivals Become Customers


Meta was also in discussions with Google Cloud about leveraging Gemini models for its advertising business. (arxiv) Advertising is where Meta generates the overwhelming majority of its revenue. If Gemini was being evaluated for ad targeting applications, the dependency goes deeper than safety tooling it touches the commercial engine of the company. Against that backdrop, Around March, Google told Meta directly: it could not meet the full Gemini capacity Meta had sought to purchase. The shortfall disrupted and delayed multiple internal AI projects as a result. (NeurIPS) Several other Google clients were affected too, though none as severely as Meta, given the scale of its demand.


Meta responded by directing employees to optimize their usage of AI tokens the units that measure compute consumption for AI projects. (ACM Digital Library) For a company that has publicly committed to providing “personal superintelligence to billions of people,” rationing AI tokens internally is an unusual position to be in.

Why Gemini Access Matters

The obvious question is: why was Meta using Gemini at all, given that it has its own models? The answer gets at something important about how frontier AI capability is actually distributed.


Llama is genuinely capable. Meta’s open-source models have attracted an enormous developer community and perform competitively on standard benchmarks. But “competitive on benchmarks” and “ready for production safety systems at two billion users” are different standards. Meta has used Gemini for coding, customer service, advertising tools, and content moderation placing the limitation near day to day operations across multiple core business lines. (ACM Digital Library) These are not experiments. They are workflows that run constantly, at massive scale, where model reliability and performance directly affect business outcomes.


The compute dimension matters as much as the model capability. Gemini API requests more than doubled between March and August 2025. (arxiv) That growth rate reflects an adoption curve steeper than most forecasts anticipated, and it’s creating resource allocation problems that no amount of spending can solve instantly. Data centers take years to build. Chip supply chains have long lead times. Demand for AI compute is growing faster than anyone’s ability to supply it. That gap is exactly what Meta ran into.

The Infrastructure Wall


On Alphabet’s April 29 earnings call, Google Cloud revenue rose 63% year over year to more than $20 billion in the first quarter of 2026, while the backlog of unmet demand nearly doubled to more than $460 billion. Sundar Pichai explicitly said Google Cloud was “compute constrained” in the near term and that revenue would have been higher if Google could meet demand. (Bloomberg) That last sentence is worth sitting with. Google is leaving money on the table not because it lacks customers, not because the pricing is wrong, but because there isn’t enough physical infrastructure to serve everyone who wants access. The company’s own growth is now colliding with the limits of what can be built fast enough.

The Real Reason Companies Limit AI Access

The instinct is to read access restrictions as strategic decisions Google choosing not to serve Meta because Meta is a competitor. The reality, in this case, is more mundane and more structurally interesting: Google genuinely cannot build compute infrastructure fast enough to meet demand. Even paying customers willing to spend enormous sums are hitting the ceiling.


Google plans to spend $180 to $190 billion on infrastructure in 2026 and is leasing additional capacity from SpaceX and xAI a deal reportedly worth around $920 million per month to ease the crunch. (Semantic Scholar) When the world’s largest search company is renting compute from a rocket company and an AI startup to keep up with its own customers, the scale of the infrastructure problem becomes concrete. Anthropic has signed a similar arrangement. The demand for AI inference the work of actually running models in response to user queries has grown so fast that it has outpaced even the capital commitments of the best resourced companies on Earth.

How Google Is Managing the Crunch


On May 17, 2026, Google imposed compute-based usage limits across all Gemini Apps. The limits vary by prompt complexity, model tier, and conversation length, and they refresh every five hours up to a weekly cap. Google’s own Gemini Apps help page ties the limits to capacity directly if capacity changes, Google may cut limits for free users before paying subscribers, and during periods of high demand, may withhold compute heavy features from free users entirely. (CNBC)


This is what a genuine resource constraint looks like when it hits a mass market product. It’s not a pricing decision or a strategic access limitation it’s triage. When you can’t serve everyone at once, you build a hierarchy. Enterprise customers with contracts get priority. Free users get reduced access. And even enterprise customers including Meta can find themselves capped when their demand exceeds what was contracted.

A Pattern Technology Has Seen Before


History offers a useful parallel. In cloud computing’s early years, whoever controlled the underlying infrastructure held leverage over everyone who needed it regardless of budget or capability. Server access was once a competitive differentiator. Semiconductor allocation shaped entire industries before chips became abundant. AI compute is following the same pattern. The companies controlling it hold a form of power distinct from and in some ways more durable than the companies building the models that run on it.

What This Says About Competition in AI

The Google Meta situation illustrates a competitive dynamic that most AI coverage misses: the most consequential battles in AI right now are not between models. They are between infrastructure positions.


The situation is particularly revealing given that Meta is simultaneously a competitor to Google in the AI space. Meta builds its own Llama family of open source models, yet still leaned heavily on Google’s Gemini infrastructure for internal projects. That’s not hypocrisy or strategic confusion on Meta’s part it’s a rational response to the reality that no single company has built sufficient AI capability across every dimension it needs. Even the most ambitious AI labs have gaps, and filling those gaps with a competitor’s infrastructure is sometimes the fastest path.


In February 2026, Meta agreed to rent Google’s Tensor Processing Units Google’s own AI accelerator chips in a separate arrangement. (Semantic Scholar) So Meta is simultaneously building its own frontier models, renting Google’s chips to train them, and using Google’s Gemini models for production applications where Llama falls short. The companies competing most fiercely at the AI frontier are also deeply interdependent at the infrastructure layer. That’s a strange equilibrium, and it’s one that access restrictions can destabilize quickly.

The Unintended Consequence


For Google, the strategic position here is complicated. Restricting a competitor’s access to Gemini, even involuntarily through capacity constraints, accelerates that competitor’s incentive to build alternatives. Those restrictions appear to have accelerated Meta’s shift toward in-house infrastructure the company cut 8,000 jobs while redirecting 7,000 employees to new AI teams. Each wall Meta hits with external providers strengthens the internal business case for independence. Ironically, the short term pain may be the exact pressure Meta needed to stop depending on a competitor’s infrastructure altogether.

What Developers and Businesses Should Expect

The implications of this story extend well beyond the relationship between two technology giants. Almost every company deploying AI today builds on external infrastructure. The Google Meta situation previews dynamics that will hit most of them eventually.


Inference the work of running models after training now represents the highest cost in AI. Every prompt a user sends consumes compute, and as usage grows, that bill compounds fast. Embedding AI deeper into business processes doesn’t scale linearly. Each new user triggers more queries, more inference cycles, and more pressure on a supply chain already stretched to its limits.


The practical consequence for businesses is that “we have API access” is no longer a permanent condition. It’s a current condition, subject to the capacity decisions of whoever controls the underlying infrastructure. When a project hits a Google cap, the system loses access and the task fails. Buying more seats does not change fixed system limits. Meta kept paid access but practical usage stayed constrained. Internal tools competing for the same model family made it worse. (ACM Digital Library) This is not a theoretical risk. It is what happened to one of Google’s largest enterprise customers, and the restrictions have been in place for months.

How to Build for This Reality


Most organizations should do what Meta is now being forced to do: reduce single-vendor dependency. That doesn’t mean abandoning major providers. It means building architectures flexible enough to route workloads elsewhere when one provider hits a ceiling. Token usage needs the same active monitoring as cloud spend with alerts, budgets, and contingency routing. Unlimited-access assumptions will get expensive fast.

Final Thoughts

One version of this story focuses on the drama two tech giants, a supplier dispute, competitive tension beneath a commercial surface. That framing isn’t wrong. It’s just the less important part.


The more durable insight is about the shape of power in AI. The companies building the most capable models get enormous attention. The companies controlling the compute have a quieter form of leverage. It doesn’t show up in benchmark results or release announcements. It shows up when a company as large as Meta has to tell its own employees to ration AI usage.

The Shift Nobody Is Talking About

For users, more AI capability is moving behind paid tiers. The deeper question is who controls the compute, who can afford it, and on what terms. Model releases get the headlines. Infrastructure decisions get the power.


For years, the AI industry was defined by which model was best. That question is shifting. Reliable access to enough compute to run those models at scale now matters just as much as the models themselves. Frontier AI access is becoming what oil, semiconductors, and cloud infrastructure were in earlier technology cycles a constraint that shapes strategy, determines winners, and rewards whoever secured it first. Companies that understand this are already building toward infrastructure independence. Those that don’t are discovering it the hard way, measured in delayed projects and efficiency mandates from management.

FAQS

Why did Google limit Meta’s access to Gemini?

Google did not deliberately cut off a competitor. Google’s own infrastructure is running at capacity the company’s cloud backlog of unmet demand exceeded $460 billion as of early 2026, and even with $180 to $190 billion in planned infrastructure spending this year, supply cannot immediately match demand. Meta had sought more Gemini computing capacity than Google was able to provide at the time, and the shortfall resulted in access limits that disrupted some of Meta’s internal AI projects. Several other Google enterprise customers faced similar, though less severe, constraints.

Does Meta still use Google’s AI technology?

Yes, in multiple ways. Gemini capacity limits haven’t ended the commercial relationship Meta keeps paid access to Gemini, though at constrained levels for certain workloads. In February 2026, Meta also agreed to rent Google’s Tensor Processing Units for its own AI development. Both companies compete fiercely in AI. Both remain significant customers of each other’s infrastructure. No clean analogue exists for that kind of relationship in earlier technology competition.

Will this affect Gemini users?

It already has, in a limited way. Since May 17, 2026, Google introduced compute-based usage limits on Gemini Apps for all users, based on prompt complexity, model choice, and conversation length. Free users face weekly caps, and Google withholds compute-intensive features like Deep Research during high-demand periods. Paid subscribers get higher limits but face the same underlying constraints. Google’s own documentation warns that limits can change without notice. Infrastructure conditions now drive that experience as much as pricing tier does.

Why are AI companies becoming more protective of their models?

Two dynamics are converging here. First, compute is genuinely scarce. When infrastructure runs at capacity, providers make allocation decisions that look like restrictions even when scarcity drives them, not strategy. Second, frontier AI model access has become a strategic asset. Companies controlling it hold leverage over competitors who depend on it, and that leverage grows as AI embeds deeper into business operations. The U.S. government’s decisions to restrict Anthropic’s Fable 5 and Mythos reflect the same logic at a national scale treating access to capable AI as something to manage strategically rather than distribute freely.

What does this mean for the future of AI competition?

The most important AI battles are shifting from model capability to infrastructure control. Companies running capable models reliably at scale, without depending on a competitor’s infrastructure build advantages that compound over time. This is already driving massive infrastructure investment across the industry Meta is spending up to $135 billion on AI infrastructure, Google is spending $180 to $190 billion, Microsoft and Amazon are on similar trajectories. The competition to control AI compute is, in some ways, more consequential than the competition to build the best model. Models improve on a cycle measured in months. Infrastructure advantages take years to build and are much harder to replicate quickly.

Further Reading on BEXORN

Curious how this fits into the bigger picture? These articles go deeper:

Why Google’s Best AI Scientists Are Leaving for Anthropic and OpenAI

Behind Every AI Delay Is a Much Bigger Story: Why Gemini 3.5 Pro Slipped to July

The AI Gap Between China and the US Is Getting Smaller

Inside the OpenAI IPO: What Public Markets Mean for the Future of AI


Leave a Reply

Your email address will not be published. Required fields are marked *

Never Miss the Stories Shaping AI

Receive reporting on artificial intelligence, Big Tech, startups, and emerging technology from around the world.