Cartesia is a Daly City-based voice AI infrastructure company, founded by researchers from the Stanford AI Lab, that builds the speech synthesis and transcription models at the core of most production voice agent deployments. Its flagship products are the Sonic series (TTS) and Ink series (STT), both built on proprietary State Space Model (SSM) architecture — a fundamentally different approach from transformer-based models that enables native streaming with dramatically lower latency. Sonic 3.5 (the current flagship) is the only product in existence with model latency of less than 100 ms, outperforming its next best alternative by a factor of four, while Ink 2 provides streaming STT with native turn detection co-optimized with Sonic to minimize the full pipeline round-trip. Cartesia raised $64M in a Series A led by Kleiner Perkins and serves customers including Quora (Poe), running 100+ voices across millions of users, and high-volume operators running 20M+ outbound calls per month.
Cartesia
Ultra-low latency TTS and STT built from the ground up for real-time voice agents
Compliance
SOC 2
Key Features
- Sonic 3.5 (real-time TTS): The world’s fastest text-to-speech model, streaming the first byte of audio in just 90ms with refined prosodic rhythm, natural intonation, superior pacing, and a wider emotional range — designed for live conversational voice agents where every 100ms of additional latency impacts user perception of naturalness.
- Ink 2 (streaming STT with native turn detection): A streaming speech-to-text model optimized for real-time voice agents with native turn detection — using context to intelligently decide when the user is done speaking versus mid-sentence pausing, reducing false interruptions in conversations.
- Instant voice cloning: Clone any voice with 10 seconds of audio, with high speaker similarity that preserves accent, emotion, and speaking style — enabling brands to deploy a consistent proprietary voice across their agent interactions.
- Full pipeline optimization: Sonic (TTS) and Ink (STT) are co-designed and co-optimized as a unified pipeline, unlike vendor-assembled pipelines using separate best-of-breed components — enabling sub-90ms TTS and 100ms transcript latency with native turn detection in a single integration.
- Multilingual support (40+ languages): Native-speaker quality voices and localization across 40+ languages and accents, with emotion, tone, and speaker identity preserved across language transfer — enabling global voice agent deployments from a single voice configuration.
- Emotion and prosody control: Fine-grained control over speech attributes including emotion tags (including [laughter]), volume, speed, and tone — allowing voice agent builders to create contextually appropriate emotional responses beyond monotone synthesis.
Use Cases
- For voice AI platforms needing the lowest possible latency: An AI voice agent company replaces its incumbent TTS provider with Cartesia Sonic 3.5, achieving sub-100ms first audio and improving the naturalness of agent responses in live calls.
- For high-volume outbound voice agent deployments: A company running millions of outbound calls per month uses Cartesia for peak concurrency at scale, with consistent latency and voice quality under load.
- For brands requiring a proprietary voice: An enterprise deploys a cloned version of its brand voice across all AI-powered customer interactions, maintaining brand consistency across millions of conversations.
Pricing MODELS
Usage-based
Pricing Summary
Usage-based, credit-based pricing. Free tier available for development. Pro, Startup, and Scale tiers offer increasing credit allocations at monthly subscription rates, with TTS billed at 1 credit per character and STT billed per second. Enterprise is custom-quoted with volume discounts, SLA guarantees, and dedicated support.
Company Size Fit
Enterprise Mid-market SMB
Technical Snapshot
API Available
Yes
LLM Provider
Proprietary
Open Source
No
Deployment Options
Cloud Saas, On-premise