Who's hiring in speech AI this week

Speech-AI hiring right now is concentrated in the model layer. Deepgram, ElevenLabs, Rime, Otter and Sanas are all adding research or applied-science roles — and where they post US compensation, it's high: $213–328K at Deepgram, $230–265K at Otter, up to $250K at DeepScribe. Healthcare documentation and in-car voice are both still hiring too. Twelve roles this week, most of them remote-friendly.

ASR & transcription

  • Senior Applied Scientist, Speech — Otter.ai · Mountain View, CA · $230K–$265K
    Architects the large-scale ASR and TTS behind Otter's meeting transcription and summarization; owns models end to end and sets ML-infra direction. 5+ yrs, PhD preferred. View posting
  • Sr. AI Engineer — Uniphore · Palo Alto, CA · $169K–$233K
    Develops and fine-tunes NLP/ASR and multimodal foundation models for enterprise CX and contact-center analytics. 5+ yrs AI/NLP/ASR, PyTorch + Hugging Face. View posting
  • AI Engineer — DeepScribe · Remote (US) · $150K–$250K
    Ships LLM-powered clinical-documentation apps and optimizes the real-time speech-recognition pipelines behind ambient notes. 3+ yrs production AI; ASR / Whisper a plus. View posting
  • ML/AI Engineer — Verbit · Kyiv, UA
    Builds and serves ASR models in production — inference optimization, GPU utilization, streaming — and tracks WER/DER/EER. Prior speech experience not required. View posting

Text-to-speech & voice

  • Director of Research, Text to Speech — Deepgram · SF / Ann Arbor / Remote · $213K–$328K
    Sets TTS research direction and picks the technical bets — hands-on leadership, not pure management. View posting
  • Machine Learning Scientist — Rime · Remote (US)
    Designs and trains speech-synthesis models (autoregressive, non-AR, full-duplex S2S) and iterates on neural-codec representations with in-house linguists. PhD or equivalent. View posting

Voice agents

  • Forward Deployed Engineer (India) — Cartesia · Bangalore · ₹70L–₹90L + equity
    Embeds with enterprise customers to deploy Cartesia's voice AI across cloud, VPC and on-prem. Backend-heavy; 4+ yrs production systems, real-time / telephony a plus. View posting
  • Speech LLM Engineer, Voice-First Agentic AI — NVIDIA · Ho Chi Minh City / Hanoi
    Builds speech, audio and multimodal LLMs for voice-first agentic AI, and drives the benchmarking framework for voice agents — quality, latency, reliability, safety. 2+ yrs speech/LLM, PyTorch; NeMo / Riva a plus. View posting

Embedded & audio infra

  • Partner Engineer, Audio & ECNR (Automotive) — SoundHound · Beijing
    Integrates SoundHound's in-vehicle audio pipeline and echo cancellation with automakers. Customer-facing; 5+ yrs SWE, C++/Python/Linux, signal-processing background. View posting

Research

  • Research Engineer — ElevenLabs · Remote (global)
    Trains and post-trains audio foundation models and designs the benchmarks that decide whether a new iteration is genuinely better. 3+ yrs ML; no degree required. View posting
  • Research Scientist, Model Evaluation — Sanas · Palo Alto, CA
    Defines how Sanas measures model quality across accent translation, noise cancellation and speech enhancement — evaluation infra, perceptual metrics, human listening studies. 4+ yrs speech/audio research. View posting
  • Principal AI Researcher, LLM — Cerence · Remote (US/Canada) or Aachen
    Builds LLM-based dialog, destination-entry and search for in-car assistants, tuned for resource-constrained automotive platforms. PhD preferred. Applications close Sept 30. View posting

Hiring for a speech-tech role? Tell us about it and we'll include it here and in the digest. Browse everything by specialty on the homepage.