Contents
Baseten
Key facts
Baseten is an artificial intelligence infrastructure company headquartered in San Francisco, founded in 2019 [5][8][19]. It operates a platform for deploying, serving, and training machine learning models in production, describing itself as "a training and inference platform" [2] and marketing its homepage as "The platform for high-performance inference" [7]. Contrary Research characterizes Baseten as a cloud provider with an "inference-first infrastructure platform" specializing in deployment and scaling of production-ready AI models, distinct from training-focused or general-purpose providers [5]. In June 2026 the company announced a $1.5 billion Series F at a $13 billion valuation, its fourth fundraise in 18 months [14].
History
Co-founders Tuhin Srivastava (CEO) and Phil Howes (Chief Scientist) grew up together in Australia and met Amir Haghighat (CTO) in 2012 as early employees at the creator marketplace Gumroad [4][5]. At Gumroad, Srivastava and Howes were data scientists who became full-stack engineers while applying machine learning to fraud and content moderation [4]; Haghighat served as head of engineering [5] and later led machine learning teams at Clover Health [4]. Srivastava and Howes subsequently co-founded Shape, an HR data analytics provider acquired by Reflektive in 2018 [4][5]. The fourth co-founder, Pankaj Gupta, was a software engineer at Uber [5]. Greylock's April 2022 announcement names three co-founders, while later company materials name four, including Gupta [4][13].
In late 2019, months after the Shape acquisition, Srivastava pitched Greylock on a company to reduce machine learning pipeline costs and iteration cycles, arguing that otherwise "most teams are never going to ship" [4]. Baseten raised a seed round in 2019 co-led by Greylock and South Park Commons [4][12]. The company was named after the base-ten blocks used to teach arithmetic, by analogy to simplifying machine learning [5]. The founding thesis cited a McKinsey estimate that industries could unlock more than $5 trillion from machine learning, while even simple models could take six or more months to ship [4][5].
After more than 18 months of building, Baseten opened its platform on April 26, 2022, announcing $20 million in combined seed and Series A financing led by Greylock [4][12]. Srivastava later wrote that the product "was raw (and clunky)" [12]. The original platform was a software toolkit for technical data science teams, organized around four pillars: Serve (model serving), Integrate (serverless APIs), Design (a drag-and-drop interface builder), and Ship (moving from notebooks to users) [12]. Early adopters included Pipe, Patreon, and Primer, applying it to content moderation, underwriting, data labeling, and audio transcription [4].
The company's focus shifted after OpenAI launched ChatGPT on November 30, 2022, and open-source models such as Stable Diffusion and Whisper demonstrated demand [9]. By its March 2024 Series B, Baseten described itself as "the most performant, scalable, and reliable way to run your machine learning workloads – in our cloud or yours," with Truss as "our open-source standard for serving models in production" [13]. All four co-founders remained with the company as of February 2026 [5].
Products and technology
Baseten's platform takes a model, whether an open-source LLM from Hugging Face, a fine-tuned checkpoint, or a custom model, and turns it into a production API endpoint with autoscaling, observability, and optimized serving, handling containerization, GPU scheduling across multiple clouds, and engine-level optimizations such as TensorRT-LLM compilation [2].
The product line includes hosted Model APIs with OpenAI-compatible endpoints for open models such as DeepSeek, Qwen, and GLM [2]; dedicated inference for high-scale and custom workloads [7][5]; and Truss, an open-source framework that packages models into deployable containers, with config-only deployments for popular architectures [2][15]. Chains orchestrates multi-step pipelines such as retrieval-augmented generation, with each step running on its own hardware; Baseten claims Chains powers "6x better GPU usage and cutting latency in half" [2][7]. Training infrastructure covers LoRA fine-tuning and reinforcement learning through the Loops SDK, plus bring-your-own-script Training Jobs on H100 or H200 GPUs [2]. Named inference engines include Engine-Builder-LLM (TensorRT-LLM compilation), BIS-LLM for large mixture-of-experts models, and BEI for embedding, reranking, and classification [2].
The underlying Baseten Inference Stack combines an inference runtime (custom kernels, speculative decoding, quantization, KV cache optimizations) with inference-optimized infrastructure (request routing, autoscaling, multi-cloud capacity management) [17]. Models scale to zero when idle and scale up within seconds [2][5]. Baseten claims failover across 10 clouds [9]; Contrary Research reports support for more than 10 providers [5]. The company operates no proprietary data centers, running workloads on public clouds such as AWS and Google Cloud [5][19]. Deployment options are the fully managed Baseten Cloud, self-hosted deployment inside a customer's VPC, and a hybrid model in which overflow traffic spills to Baseten Cloud [5][7]. Baseten states it never stores inference inputs or outputs and is SOC 2 Type II certified and HIPAA compliant [5].
Funding
After the 2019 seed and the April 2022 Series A announcement [4][12], Baseten raised a $40 million Series B in March 2024 led by IVP and Spark [13]; a $75 million Series C announced February 19, 2025 at an $825 million valuation, co-led by IVP and Spark [10][19]; a $150 million Series D in 2025 led by BOND, with Jay Simons joining the board [9]; and a $300 million Series E announced around January 2026 at a $5 billion valuation, led by IVP and CapitalG [11]. The $1.5 billion Series F in June 2026, at a $13 billion valuation, was led by Altimeter Capital, Conviction Partners, and Spark Capital and co-led by Sands Capital and Wellington Management [14]. Total disclosed funding stood at $585 million as of February 2026 [5], reaching approximately $2.085 billion with the Series F [14].
Business model and customers
Baseten runs AI models for clients at the inference stage, when trained models generate outputs in response to queries, on infrastructure rented from cloud providers including Amazon and Google [19]. An enterprise tier lets customers supply their own infrastructure [19]. By drawing on multiple providers, Baseten offers access to more GPUs than a single cloud's supply and shields clients from provider-side disruptions [19]. Head of marketing Mike Bilodeau said clients often see inference costs fall 40 percent or more relative to homegrown architectures [19].
The customer base grew from early adopters such as Patreon and Pipe [4] to more than 100 enterprise customers plus hundreds of smaller companies by early 2025 [19]. The Series F announcement named Cursor, Notion, Lovable, Harvey, HubSpot, OpenEvidence, Abridge, Decagon, and Parallel as customers [14].
Revenue grew sixfold in the fiscal year ending January 2025, according to Srivastava [19]. Baseten reported that inference volume grew 100x in the year preceding the Series E [11], and that revenue grew 20x and inference volume 40x in the year preceding the Series F [14]. Headcount rose from about 60 employees in February 2025 [19] to 200 by February 2026 [5].
Competition and risks
An East Wind analysis groups Baseten with Modal as offering an "in-between" experience, with more infrastructure control than API-only rivals such as Replicate, Fireworks AI, and Deepinfra [3]. CNBC names Salesforce-backed Together AI as a competitor and notes Baseten competes with AI model companies and hedge funds for talent [19]. Baseten claims that "in 95% of head-to-head bakeoffs, we beat out competitors by 40-50% in performance" [9].
The Wall Street Journal described the Series F as a "split-priced round" in which some investors came in at a $13 billion valuation and others at $11 billion, "a tactic startups use to boost headline valuation and make lead investors look good on paper" [6]. The East Wind analysis warns that price and performance among top competitors are similar while margins are not, and that "most startups will not be able to absorb their COGS (cost of goods sold) in the long run" [3]. Baseten's dependence on cloud providers surfaced operationally in a June 26, 2026 incident attributed to a cloud provider outage [1]. Marketing materials claim 99.99% uptime [7][17], while the contractual SLA for Dedicated Inference commits to 99.9% monthly availability, with service credits capped at 40% of the monthly bill and exclusions for third-party hosting failures [18]. Baseten's status page recorded 12 incidents in June 2026 and 15 in July 2026 [1].
Watch
References
- 1.Baseten Status - Incident History↩
- 2.Baseten overview - Baseten↩
- 3.A Deep Dive on AI Inference Startups - by Kevin Zhang↩
- 4.Self-Serve Apps for ML Teams | Greylock↩
- 5.Report: Fivetran Business Breakdown & Founding Story | Contrary Research↩
- 6.AI inference startup Baseten reportedly raising $1.5B months after its last mega-round | TechCrunch↩
- 7.Inference Platform: Deploy AI models in production | Baseten↩
- 8.Meet the engineers behind Baseten↩
- 9.Announcing Baseten’s $150M Series D ↩
- 10.Announcing Baseten’s $75M Series C↩
- 11.Announcing Baseten's $300M Series E↩
- 12.Announcing our Series A↩
- 13.Announcing our Series B↩
- 14.Announcing our Series F↩
- 15.Why we built and open-sourced a model serving solution↩
- 16.Cloud Pricing↩
- 17.The Baseten Inference Stack | Guides↩
- 18.Service Level Agreement↩
- 19.AI startup Baseten raises $75 million following DeepSeek's emergence↩