Skip to main content
Cartesia’s models are portable enough to run on widely available GPU hardware. In the table below we show the recommended concurrency for our TTS and STT model workers. See Metrics for more details on performance metrics.

Compatibility Matrix

Kubernetes and tooling

GPU

MIG (Multi-Instance GPU)

When choosing hardware you need to consider the tradeoffs between latency (TTFA), and throughput. See the table below for the metrics on the different set of GPUs we test on:
The benchmarks below are for Sonic 3.5 and require release tag sonic-20260503 or later. Updated April 2026.
With these you’ll setup your per worker configurations. For handling your application’s scaling requirements, you’ll need to configure autoscaling behavior. See autoscaling for more details.