Contents
Fireworks AI
Key facts
Fireworks AI is an American artificial intelligence infrastructure company headquartered in Redwood City, California [17][18]. Founded in 2022, it operates a managed cloud platform for training, fine-tuning, and serving open and enterprise-owned AI models in production [9][17].
History
Fireworks AI was founded in 2022 by Lin Qiao and six co-founders who left the PyTorch team at Meta [9][16][18]; Contrary Research gives a specific founding date of October 2022 [17]. Six of the seven co-founders have Meta backgrounds; the seventh, Chenyu Zhao, was a Google Vertex AI lead [13]. Qiao, the CEO, studied computer science at Fudan University, earned a Ph.D. at the University of California, Santa Barbara, and worked at IBM and LinkedIn before Meta, where she led the PyTorch effort and grew an AI infrastructure team from five engineers to more than 300 [17]. The founding thesis held that the future of AI "wouldn't be controlled by a handful of powerful foundation model labs, but distributed across thousands of enterprises that want to own and customize their own AI products" [9]. The company's initial mission was to make production-scale inference accessible and affordable, with focus areas of reducing computing costs, improving response times, and making inference easier to operate at scale [17]. It later expanded into model training [15].
Products and technology
Fireworks operates a managed platform that consolidates model hosting, GPU provisioning, autoscaling, performance optimization, monitoring, traffic routing, and failover into one system [17]. Its core is a proprietary disaggregated inference engine with custom kernels, speculative decoding, disaggregated KV caching, and multi-node expert parallelism for large mixture-of-experts models [11]. Named proprietary technologies include FireAttention, a custom CUDA kernel-based serving stack; FireOptimizer; and Multi-LoRA, which consolidates many fine-tuned variants onto one base-model deployment [18][11]. FireAttention V4, announced in May 2025, added FP4 precision support on NVIDIA B200 GPUs, and independent benchmarking showed throughput above 250 tokens per second on DeepSeek V3 [5]. The company keeps the engine proprietary while contributing fixes back to the open-source serving stacks vLLM and SGLang [2].
The platform offers three deployment tiers: serverless, billed per token with OpenAI- and Anthropic-compatible APIs; on-demand dedicated deployments, billed per GPU second; and reserved capacity with bring-your-own-cloud support [11][12]. Training capabilities include supervised fine-tuning, DPO, KTO, and reinforcement learning through a self-serve SDK, and trained models serve on the same stack as inference [2][11]. A Fireworks Training Preview was announced in April 2026 [6][7].
The model catalog covers hundreds of open-source models across text, image, audio, and multimodal formats, with day-zero support for new releases [9][11]; families offered include Meta's Llama series, Mistral, Qwen, and DeepSeek [15]. In July 2026 the company released Fireworks Nexus, a routing and cost-control platform for engineering organizations that uses a difficulty-scoring model to direct requests between open and closed models, with a quoted 50 to 75 percent reduction in AI coding spend [20].
Fireworks does not own its GPU fleet; it sources capacity from more than 20 third-party suppliers and operates across regions in the US, Europe, and APAC [16][18]. Enterprise features include zero data retention by default, SSO, audit logs, HIPAA and SOC2 compliance, and airgapped deployments [18].
Business model and customers
Fireworks is a B2B managed infrastructure platform with usage-based pricing: serverless inference per token, fine-tuning per training token, reinforcement fine-tuning per GPU-hour, and dedicated deployments per GPU-second [18][12]. Go-to-market combines self-serve API access with enterprise contracts that add reserved capacity, higher rate limits, and private deployment [18].
The company reported more than 10,000 customer companies as of October 2025, a tenfold increase from its Series B, along with "hundreds of thousands of developers" [9]. Named customers include Cursor, Harvey, Uber, Shopify, DoorDash, Notion, Samsung, and Upwork [9][6][16]. Cursor, described as Fireworks' most prominent customer case [3], builds its Composer models on Fireworks' training stack with a focus on reinforcement learning, while Fireworks manages rollout and inference [2].
Funding and financials
Fireworks raised a $25 million Series A in early 2024 [18]; a $52 million Series B in July 2024 at a $552 million valuation, led by Sequoia Capital [6]; a $250 million Series C in October 2025 at a $4 billion valuation, co-led by Lightspeed Venture Partners, Index Ventures, and Evantic [9]; and a $1.505 billion Series D in July 2026 at a $17.5 billion valuation, led by Atreides Management, Index Ventures, and TCV [10][16]. Total capital raised exceeds $1.8 billion [17][18]. Forge Global's secondary marketplace lists a $15.54 billion Series D valuation, which conflicts with the announced $17.5 billion post-money figure [14][10].
Annualized revenue surpassed $280 million at the Series C [9] and $1 billion at the Series D, a figure Quartz reports grew fivefold year over year [10][16]. Daily token volume grew from 140 billion at the Series B [6] to more than 10 trillion at the Series C [9] and more than 40 trillion by mid-2026, with over 95 percent of tokens coming from models specialized on customers' proprietary data [10][16]. Gross margins are in the 30 to 40 percent range according to Qiao, or roughly 50 percent per Sacra's estimate, reflecting GPU infrastructure costs in cost of goods sold; the company states it is prioritizing growth speed over margin optimization [3][18]. Headcount stood at roughly 200 as of mid-2026, with plans to triple the workforce before year-end [16].
Partnerships, competition, and risks
In March 2026 Fireworks announced a partnership bringing its open-model inference into Microsoft Foundry on Azure as first-party models, and separately acquired Hathora to deepen real-time global compute orchestration [7][1]. At NVIDIA GTC 2026, Jensen Huang described Fireworks as "the TSMC of AI Factories" [4].
The company classifies itself as a "direct provider" that controls both the API endpoint and the underlying GPU capacity, listing Together AI, Baseten, Groq, Cerebras, and Replicate as peers in that category [8]. Other named competitors include closed-model providers OpenAI, Anthropic, and Google, and hyperscaler platforms such as AWS Bedrock and Google Vertex AI [15][18]. Sacra identifies three principal risks to the business: commoditization of inference as open-source frameworks such as vLLM and SGLang improve, capture by hyperscaler partners whose platforms absorb inference workloads, and hardware concentration given the company's reliance on NVIDIA GPUs sourced from third parties [18].
Watch
References
- 1.Introducing Fireworks AI on Microsoft Foundry: Bringing high performance, low latency open model inference to Azure | Microsoft Azure Blog↩
- 2.Most AI Startups Are Scaling Into Bankruptcy | Lin Qiao, CEO of Fireworks|Gradient Dissent — BigGo Finance↩
- 3.Are OpenAI & Anthropic Overvalued? The Open-Source AI Reality with Fireworks AI, Lin Qiao|20VC — BigGo Finance↩
- 4.Own Your Specialized Intelligence | Fireworks ↩
- 5.FireAttention V4: Industry-Leading Latency and Cost Efficiency with FP4↩
- 6.Fireworks Raises $52M Series B to Lead Industry Shift to Compound AI Systems↩
- 7.Fireworks AI↩
- 8.Inference Providers vs. API Routers: Where Do Your Tokens Actually Come From? | Fireworks↩
- 9.Fireworks Raises $250M Series C to Power the Future of Enterprise AI↩
- 10.Fireworks Secures $1.5 Billion in Series D Funding↩
- 11.Inference | Fireworks↩
- 12.Fireworks - Pricing↩
- 13.Fireworks - Team↩
- 14.Fireworks AI IPO Timeline and Financing Details - Forge↩
- 15.Lin Qiao’s Fireworks Bets on Specialized Models Over General A.I. Hype | Observer↩
- 16.Fireworks AI raises $1.5 billion Series D at $17.5 billion valuation↩
- 17.Report: Fireworks AI Business Breakdown & Founding Story | Contrary Research↩
- 18.Fireworks AI revenue, valuation & funding | Sacra↩
- 19.AI Startups Struggle for Fast Access to DeepSeek Models - Business Insider↩
- 20.Fireworks AI Releases Fireworks Nexus: A Drop-In Routing and Cost-Control Layer That Moves Routine Coding Work to Open-Weight Models - MarkTechPost↩