The demand for inference has gone vertical. Hyperscaler capex on AI infrastructure is expected to exceed $1T+ next year, driven by hyperscalers, frontier labs and neoclouds. AI data centers are ultimately token factories monetized by inference. However, the reality today is that the infrastructure supply chain is extremely tight, and existing GPU fleets often run suboptimally at 10-50% utilization. In a world constrained by compute and power, maximizing efficiency, throughput, and performance is mission-critical.
Gimlet Labs was purpose-built for these constraints, delivering unprecedented performance and enabling massive token volumes at very low latency with its multi-silicon inference cloud. The company is now scaling to hundreds of megawatts of managed heterogeneous infrastructure with billions of dollars of contracted revenue. For these reasons and more, we’re thrilled to partner with Zain Asgar and the Gimlet Labs team as a major investor in their $300M Series B financing.
The Multi-Trillion, Multi-Silicon Opportunity
Per Gartner, the AI Infrastructure market is expected to grow to $2.8 trillion in 2030, up from $1.5 trillion in 2026. The vast majority of this infrastructure will power inference, but not all of this spend will go to a single architecture. Given the binding constraints of power and capacity, we believe it will be a heterogeneous market with multiple winners where software like Gimlet’s optimizes inference across silicon to deliver the best throughput per megawatt.
GPUs are extremely powerful, but each stage of inference faces a different bottleneck, and standalone GPUs are not optimal for all of them. Prefill is compute-bound; decode is memory bandwidth-bound. GPUs excel at raw compute, while AI accelerators with SRAM-based architectures offer far higher memory bandwidth. An accelerator optimized for one phase is structurally suboptimal for the other, and these gaps compound with agentic workloads that chain models and tool calls together over many steps.
Heterogeneous silicon for inference is already here. NVIDIA licensed Groq’s IP for $20B and now ships Groq 3 LPX for disaggregated inference in Vera Rubin. OpenAI signed a 750MW deal with Cerebras and unveiled its own custom inference chip, Jalapeño. And last but not least, Anthropic has created multiple partnerships across NVIDIA, AMD, Google, Amazon, and more in addition to building its own inference silicon in-house. Multi-sourcing silicon solves for capacity and diversification, but the next step – maximizing performance, throughput and efficiency across the system – is exactly what Gimlet Labs is built for.
Gimlet’s Multi-Silicon & Agent-Native Inference Cloud
Gimlet capitalizes on each system’s comparative advantage to achieve an entirely new Pareto frontier of performance, delivering up to 10x gains in inference throughput and interactivity within the same power envelope. Its software intelligently slices and orchestrates AI inference workloads, matching them to the most appropriate hardware. Compute-bound workloads land on GPUs and memory bandwidth-bound workloads land on inference chips built for extremely fast data access.
This optimization via disaggregation of inference workloads yields a powerful advantage: more tokens faster and cheaper. It is no surprise that the world’s leading AI semiconductor companies, including NVIDIA, AMD, Intel, Arm, Cerebras and d-Matrix, have partnered with Gimlet to support their chips in its multi-silicon architecture.
The Team Behind the Backbone of AI Inference
Problems this deep get solved by teams who have lived them, and Gimlet Labs’ founders have spent their careers at the intersection of low-level infrastructure software, hardware and AI. Having earned their insights the hard way, they were early in recognizing that multi-silicon inference was inevitable.
Zain Asgar, Michelle Nguyen, Omid Azizi, Natalie Serrino and James Bartlett are the co-founders of Gimlet and were the team behind Pixie, the eBPF-based Kubernetes observability company acquired by New Relic, where Zain went on to serve as a General Manager. In addition to being an entrepreneur, Zain (CEO) serves on the faculty at Stanford, where his research has spanned AI systems and custom chip design. Omid (Head of Hardware Platforms) earned his PhD at Stanford after a career in chip design and computer architecture. And Michelle (Head of Engineering), Natalie (leading KForge) and James (System Architect) are accomplished engineers who have shipped infrastructure at scale, used in production around the world. Fluency in silicon and fluency in software rarely live inside one founding team, and Gimlet has both.
At Sapphire Ventures, we pride ourselves on partnering with Enterprise AI founders building Companies of Consequence. Gimlet sits squarely at the intersection of the two most powerful forces in technology today: the rise of AI agents and the historic buildout of the infrastructure that serves them.
We’ve backed companies building foundational infrastructure and toolchains like Baseten, LangChain, Temporal, and Weights & Biases. Gimlet sits at the deepest layer of that stack. We believe whoever enables heterogeneous compute at scale will define a generation of infrastructure, and we believe Gimlet is the team to do it.
We’re honored to partner with Gimlet on what comes next.
To learn more, visit gimletlabs.ai. And if building the backbone of AI inference sounds like your kind of problem, the team is hiring.
Key Takeaways
- Gimlet Labs raised a $300M Series B, with Sapphire Ventures partnering as a major investor.
- AI infrastructure spend is projected to reach $2.8T by 2030, up from $1.5T in 2026, per Gartner, yet existing GPU fleets often run at only 10% to 50% utilization. Gimlet’s multi-silicon inference cloud solves this by orchestrating inference workloads across GPUs and AI accelerators, matching each stage to the hardware built for it.
- Gimlet is scaling to hundreds of megawatts of managed heterogeneous infrastructure with billions of dollars in contracted revenue, and has partnered with NVIDIA, AMD, Intel, Arm, Cerebras and d-Matrix to support their chips within its architecture.
- Sapphire’s conviction is rooted in the founding team. Zain Asgar, Michelle Nguyen, Omid Azizi, Natalie Serrino and James Bartlett previously built Pixie, an observability company acquired by New Relic, and bring deep fluency across AI systems, chip design and infrastructure software.
Legal disclaimer
This article is for informational purposes only. Nothing presented within this article is intended to constitute investment advice, and under no circumstances should any information provided herein be used or considered as an offer to sell or a solicitation of an offer to buy an interest in any investment fund managed by Sapphire. Information provided reflects Sapphire’s views as of a time, whereby such views are subject to change at any point and Sapphire shall not be obligated to provide notice of any change. Companies mentioned in this article are a representative sample of portfolio companies in which Sapphire has invested in which the author believes such companies fit the objective criteria stated in commentary, which do not reflect all investments made by Sapphire. A complete alphabetical list of investments made by Sapphire’s Growth strategy is available here. No assumptions should be made that investments listed above were or will be profitable. Due to various risks and uncertainties, actual events, results or the actual experience may differ materially from those reflected or contemplated in these statements. Nothing contained in this article may be relied upon as a guarantee or assurance as to the future success of any particular company. Past performance is not indicative of future results.