Modal
Modal Labs, commonly known as Modal, is an AI infrastructure company that provides a serverless cloud platform optimized for inference, training, batch processing, and sandboxed code execution.
Profile
Modal provides a serverless cloud platform where developers can run AI inference, training, and batch jobs using Python code, with automatic scaling and per-second billing.
Modal Labs, commonly known as Modal, is an AI infrastructure company that provides a serverless cloud platform optimized for inference, training, batch processing, and sandboxed code execution. Founded by Erik Bernhardsson, the company has raised significant venture capital, including an $87 million Series B in late 2025 that valued it at $1.1 billion. By February 2026, TechCrunch reported that Modal was in early talks to raise another round at a roughly $2.5 billion valuation, with General Catalyst as a potential lead investor.
The company’s annualized revenue run rate was estimated at approximately $50 million at that time. Modal’s platform is built around a Python SDK that lets developers define compute, hardware, and dependencies in code, with sub-second container cold starts and automatic scaling from zero to thousands of GPUs. It supports a range of workloads including LLM and multi-modal inference, fine-tuning, reinforcement learning rollouts, and secure sandboxes for agents.
Modal routes workloads across multiple clouds and regions in real time, and claims SOC 2 and HIPAA compliance. Customers include Decagon, Runway, Physical Intelligence, Suno, Chai Discovery, Lovable, Quora, and Reducto. The company competes with other GPU-as-a-service providers like Replicate, Together AI, and Fal.ai, as well as the major cloud providers. Modal’s rapid valuation growth and customer traction reflect strong demand for flexible, developer-friendly AI infrastructure, but also raise questions about sustainability given the capital intensity of GPU provisioning and the competitive landscape.
Who buys this
- AI startups and scale-ups needing flexible GPU infrastructure for inference and training
- Enterprise teams deploying AI agents or sandboxed code execution environments
- Research labs and biotech firms running computationally intensive simulations or molecular modeling
Publicly disclosed clients
- Decagon
- Runway
- Physical Intelligence
- Suno
- Chai Discovery
- Lovable
- Quora
- Reducto
Strengths and what to watch
Strengths
- Sub-second container cold starts and instant autoscaling from zero to thousands of GPUs, which reduces idle cost and developer friction
- Python-native SDK that allows developers to specify compute, hardware, and dependencies in a single code file, simplifying deployment
- Multi-cloud routing and real-time GPU scheduling across regions, providing flexibility and resilience without capacity planning
Watch for
- Capital intensity: Modal must continuously invest in GPU capacity across multiple clouds; any downturn in AI infrastructure spending could pressure margins
- Competitive pressure from well-funded rivals (Replicate, Together AI, Fal.ai) and hyperscalers (AWS, GCP, Azure) that offer similar serverless GPU services
- Valuation risk: The reported $2.5 billion valuation target in early 2026, if closed, would represent a rapid multiple on ~$50 million ARR, raising expectations for growth that may be hard to sustain
Recent moves
Key Information
- Industry
- GPU Cloud / Infra
- Founded
- 2021
- Headquarters
- New York City
Frequently Asked Questions
What is Modal and what does it do for AI developers?
Modal is a serverless cloud platform where developers run AI inference, training, and batch jobs using Python code. It handles automatic scaling from zero to thousands of GPUs and charges per second, reducing idle costs and simplifying deployment.
How does Modal's Python SDK simplify AI deployment?
Modal's Python SDK lets developers define compute, hardware, and dependencies in a single code file. This approach eliminates complex infrastructure setup, enabling sub-second container cold starts and instant autoscaling, so teams focus on code rather than capacity planning.
What types of AI workloads does Modal support?
Modal supports LLM and multi-modal inference, fine-tuning, reinforcement learning rollouts, batch processing, and sandboxed code execution. It also handles computationally intensive simulations for research labs and biotech firms, with real-time multi-cloud GPU scheduling across regions.
Is Modal compliant with SOC 2 and HIPAA standards?
Yes, Modal claims SOC 2 and HIPAA compliance, making it suitable for enterprise teams and regulated industries. This allows organizations to deploy AI agents or sandboxed code execution environments while meeting security and privacy requirements.
Who are some notable customers using Modal?
Notable Modal customers include Decagon, Runway, Physical Intelligence, Suno, Chai Discovery, Lovable, Quora, and Reducto. These companies use Modal for AI inference, training, and batch processing, reflecting its traction among AI startups and scale-ups.
How does Modal compare to competitors like Replicate or Together AI?
Modal competes with GPU-as-a-service providers like Replicate, Together AI, and Fal.ai, as well as major cloud providers. Its differentiators include sub-second cold starts, a Python-native SDK, multi-cloud routing, and per-second billing, but it faces capital intensity and competitive pressure.
Sources
- modal.com — Company product description, customer logos, and workload categories
- techcrunch.com — Details on $2.5B valuation talks, $50M ARR estimate, and General Catalyst as potential lead
- modal.com — Series B announcement of $87M at $1.1B valuation (referenced in TechCrunch article)
- modal.com — Customer case studies and logos for Decagon, Runway, Physical Intelligence, Suno, Chai Discovery, Lovable, Quora, Reducto