Modular
Modular builds a unified compute layer for AI, eliminating the need for developers to rewrite code when moving between different hardware platforms.
Profile
Modular provides a software platform that optimizes AI model inference across different hardware without requiring code rewrites.
Modular builds a unified compute layer for AI, eliminating the need for developers to rewrite code when moving between different hardware platforms. Founded in 2022 by Chris Lattner (creator of LLVM and Swift) and Tim Davis (former Google Brain leader), the company offers MAX, a full-stack AI inference framework, and Mojo, a systems programming language optimized for GPU acceleration. In September 2025, Modular raised $250 million in Series B funding at a $1.6 billion valuation, bringing total capital raised to $380 million.
The company now employs approximately 337 people across offices in San Francisco, Los Altos, Boston, and Edinburgh. Modular's platform enables inference optimization across NVIDIA, AMD, Intel, and ARM hardware without code changes. Named customers include Inworld AI, which achieved 70% faster audio synthesis latency and 60% lower costs by adopting Modular's stack, and Hippocratic AI, which saw 30% faster P99 latency for real-time patient conversations.
The company has also partnered with AWS, Oracle, and major hardware vendors. In 2025, Modular released over 450,000 lines of production-grade code under the Apache 2.0 license and announced plans to open-source the Mojo compiler in 2026. The company competes directly with vLLM, TensorRT-LLM, and SGLang in the inference optimization space, but differentiates through hardware abstraction and language-level optimization.
Modular's positioning targets a market frustrated by NVIDIA lock-in, though leadership has stated the goal is enabling competition rather than defeating NVIDIA. The company faces execution risk around adoption by larger enterprises and must prove long-term performance advantages over entrenched open-source alternatives.
Who buys this
- AI application companies building real-time voice and text inference (speech synthesis, conversational AI)
- Healthcare and biotech firms deploying patient-facing AI agents at scale
- Cloud providers and infrastructure vendors (AWS, Oracle, Lambda Labs) offering multi-hardware support
- ML teams at enterprises running inference on heterogeneous GPU fleets (NVIDIA, AMD, custom silicon)
Publicly disclosed clients
- Inworld AI
- Hippocratic AI
- Amazon Web Services (AWS)
- Oracle
- SF Compute
Strengths and what to watch
Strengths
- Founding team credibility: Chris Lattner (LLVM, Swift, Cloud TPUs) and Tim Davis (Google Brain) bring deep systems and ML infrastructure expertise
- Significant capital base ($380M total, $1.6B valuation) and strong investor backing (GV, Greylock, General Catalyst, US Innovative Technology Fund) de-risks R&D and hiring
- Demonstrated customer wins with measurable performance gains: Inworld achieved 70% latency improvement and 60% cost reduction; Hippocratic AI saw 30% P99 latency gains on production workloads
Watch for
- Adoption stagnation: despite strong technical benchmarks, enterprise adoption may lag if vLLM and TensorRT-LLM (NVIDIA) remain entrenched; MAX must prove ROI on migration effort
- Mojo language adoption risk: the language ecosystem remains immature; open-sourcing the compiler in 2026 is a bet that community-driven development will reach parity with proprietary alternatives
- Hardware dependency: market success hinges on AMD, Intel, and ARM executing on GPU roadmaps; a dominant player (NVIDIA) controlling software and hardware creates asymmetric competition
Recent moves
- 7mo ago Modular 2025 year in review: Inworld TTS tops Artificial Analysis Speech Leaderboard with 70% faster latency
- 8mo ago Modular announces Mojo 1.0 Beta and roadmap to production release in summer 2026
- 11mo ago Modular raises $250M Series B at $1.6B valuation to scale unified AI compute layer
- 11mo ago Hippocratic AI partners with Modular to achieve 30% faster P99 latency for real-time patient conversations
Key Information
- Industry
- Enterprise ML Platforms
- Founded
- 2022
Frequently Asked Questions
What is Modular AI software?
Modular provides a unified compute layer for AI inference that works across NVIDIA, AMD, Intel, and ARM hardware without code changes. Founded in 2022, it offers MAX, a full-stack inference framework, and Mojo, a systems programming language for GPU acceleration.
Who founded Modular?
Modular was founded in 2022 by Chris Lattner, creator of LLVM and Swift programming languages, and Tim Davis, former Google Brain leader. Together they brought deep expertise in systems design and ML infrastructure, attracting major investors including GV and Greylock.
What customer results has Modular achieved?
Inworld AI achieved 70% faster audio synthesis latency and 60% lower costs using Modular's stack. Hippocratic AI saw 30% faster P99 latency for real-time patient conversations. Both customers demonstrated significant performance improvements on production workloads, validating Modular's multi-hardware optimization approach.
Does Modular break NVIDIA lock-in?
Modular enables AI inference optimization across NVIDIA, AMD, Intel, and ARM without vendor lock-in. The platform competes with vLLM and TensorRT-LLM by offering hardware abstraction at the language level. Modular's goal is enabling competition, not defeating NVIDIA, though it targets enterprises frustrated by single-vendor dependencies.
How much did Modular raise in Series B?
Modular raised $250 million in Series B funding in September 2025 at a $1.6 billion valuation. This brings total capital raised to $380 million, de-risking the company's R&D and hiring. Investors include GV, Greylock, General Catalyst, and the US Innovative Technology Fund.
When is Mojo 1.0 coming?
Modular announced Mojo 1.0 Beta with a production release target of summer 2026. The company plans to open-source the Mojo compiler in 2026 to support community-driven development. In 2025, Modular released 450,000 lines of production-grade code under the Apache 2.0 license.
How Modular compares
Direct head-to-head against 3 competitors. Picked by 7wData.
Modular
- Positioning
- Modular provides a software platform that optimizes AI model inference across different hardware without requiring code rewrites.
- Customer segments
- AI application companies building real-time voice and text inference (speech synthesis, conversational AI)
- Strengths
- Founding team credibility: Chris Lattner (LLVM, Swift, Cloud TPUs) and Tim Davis (Google Brain) bring deep systems and ML infrastructure expertise
- Watch for
- Adoption stagnation: despite strong technical benchmarks, enterprise adoption may lag if vLLM and TensorRT-LLM (NVIDIA) remain entrenched; MAX must prove ROI on migration effort
- Recent moves
- Modular raises $250M Series B at $1.6B valuation to scale unified AI compute layer
Inferact (vLLM)
- Positioning
- Commercial entity built by vLLM's original maintainers, most widely deployed open-source LLM serving framework, running at Amazon, LinkedIn, and Roblox.
- Customer segments
- Hyperscalers, consumer app platforms, and enterprise ML teams running self-hosted GPU inference at scale, primarily infrastructure engineers and MLOps leads.
- Strengths
- vLLM adopted by Amazon, LinkedIn, and Roblox in production. Speculative decoding cut Roblox latency 50% serving 4 billion tokens weekly.
- Watch for
- Most customers run vLLM free. Inferact's commercial model is unproven, with enterprise contracts unreported as of mid-2026.
- Recent moves
- January 2026: Inferact launched with $150M seed at $800M valuation, co-led by Andreessen Horowitz and Lightspeed.
RadixArk (SGLang)
- Positioning
- Berkeley research spinout commercializing SGLang, deployed on hundreds of thousands of GPUs at Google, Microsoft, xAI, and Oracle.
- Customer segments
- Frontier AI labs and hyperscalers running reasoning model workloads at scale, ML infrastructure leads at large model providers.
- Strengths
- SGLang runs production inference for Google, xAI, and Microsoft. Radix Attention KV cache reuse reduces compute costs on shared prefixes.
- Watch for
- Seed-stage company as of May 2026 with enterprise support infrastructure still unbuilt. Customers scaling beyond community support may hit SLA gaps.
- Recent moves
- May 2026: RadixArk launched with $100M seed at $400M valuation, led by Accel with NVIDIA and AMD as investors.
NVIDIA TensorRT-LLM
- Positioning
- NVIDIA's native inference optimization toolkit. Deepest GPU kernel access and highest raw throughput for NVIDIA hardware deployments.
- Customer segments
- Cloud providers and enterprises standardized on NVIDIA GPUs. Data center and MLOps teams optimizing H100 and Blackwell deployments.
- Strengths
- Native FP8 and NVFP4 quantization achieving 10,000+ output tokens per second on H100. No overhead from hardware abstraction layers.
- Watch for
- NVIDIA hardware lock-in: migrating to AMD or ARM silicon requires entirely separate toolchains, raising switching costs on heterogeneous GPU fleets.
- Recent moves
- March 2026: NVIDIA Dynamo 1.0 launched as open-source inference OS, integrating TensorRT-LLM, vLLM, and SGLang for multi-node deployments.
Sources
- www.modular.com — Series B funding round ($250M), valuation ($1.6B), total capital raised ($380M), investor names, founding year
- www.modular.com — Founders (Chris Lattner, Tim Davis), leadership team, office locations, company mission
- www.modular.com — Inworld customer results (70% latency improvement, 60% cost reduction), technical details on MAX Framework adoption
- www.modular.com — Open source release of 450K+ lines of code under Apache 2.0, MAX kernels, Mojo standard library
- www.modular.com — Mojo 1.0 Beta announcement, summer 2026 release timeline, plan to open-source compiler
- www.modular.com — Hippocratic AI partnership, performance metrics (30% P99 latency improvement, 22% mean latency improvement)
- www.modular.com — 2025 product milestones, Inworld TTS 1 Max performance, Mammoth announcement, open source progress
- getlatka.com — Employee count estimates (278-299 in 2025, 337 as of Feb 2026), revenue data