Deepgram now exposes GPU-level inference metrics for its speech-to-text and text-to-speech models on Amazon SageMaker. This integration reveals specific feature usage and billing data previously locked inside vendor containers. Practitioners can now align capacity planning with actual hardware utilization. It solves a persistent observability trade-off for self-hosted audio AI deployments.