Who's hiring in speech AI this week

Nine roles this week, spread across five different kinds of speech-tech work: research, voice-agent engineering, speech analytics, platform infrastructure, and a few broader product and design roles at speech-first companies. Dolby is hiring a spatial-audio researcher for XR content; Wispr's ML engineer role tops out at $400K building the sub-500ms inference behind its voice interface; Deepgram needs someone to port its speech models onto hardware it doesn't control (non-NVIDIA edge devices); and Retell AI is paying $180–300K for a forward-deployed engineer fluent in Mandarin to support its China accounts. A genuinely mixed week.

Research

  • Research Scientist, Spatial Audio & AI — Dolby · Atlanta, GA · $140K–$170K + bonus + equity
    Designs AI models and multimodal foundation models for spatial and XR audio content creation in Dolby's Advanced Technology Group, combining deep learning with perceptually grounded audio signal processing. PhD preferred; strong audio signal analysis background, Python, PyTorch or TensorFlow. View posting

Voice agents & real-time speech

  • ML Engineer — Wispr · San Francisco (onsite) · $250K–$400K + equity
    Builds the voice interface behind Wispr's Flow and Notetaker products: sub-500ms LLM inference at scale, plus fine-tuning and RL to personalize the underlying speech models. Previous founding or startup experience preferred; fluency in Python and LLM development required. View posting
  • Senior Forward Deployed Engineer (Mandarin fluency required) — Retell AI · Redwood City, CA (onsite, relocation provided) · $180K–$300K + equity
    Embeds with enterprise customers — a large share based in China — to take Retell's voice-AI call-center agents from prototype to production: LLM prompting, full-stack integrations with CRMs and scheduling tools. 2+ years in a customer-facing technical role; fluent Mandarin and English required. View posting

Speech analytics

  • Backend Engineer, Communications Capture — Gong · Dublin, Ireland (hybrid, 3 office days/week) · compensation not listed
    Builds and scales the pipelines that capture customer calls, emails and video-conferencing audio into Gong's Revenue Intelligence platform — the raw data underneath its speech-based analytics. Java/Spring Boot, Kafka/SQS, AWS/GCP/Azure. 5+ years backend engineering; no speech-specific background required. View posting

Platform & edge infrastructure

  • Applied ML Engineer, Edge Devices — Deepgram · Remote (US) · $155K–$245K + equity + bonus
    Ports Deepgram's speech models to non-NVIDIA and edge hardware — quantization, operator swaps, graph rewrites — and validates accuracy and latency on real devices. Required: hands-on production experience deploying models to edge or non-NVIDIA hardware; strong Python and PyTorch. Speech or streaming-model experience is a plus, not required. View posting
  • ML Data & Platform Engineer — Speechmatics · Cambridge, UK (hybrid, 2–3 office days/week) · compensation not listed
    Owns the pipelines and platform behind Speechmatics' speech-AI models — sourcing and validating training data, then training, evaluating and serving the models in production, including GPU/distributed-training optimization. Strong Python/SQL, Docker/Kubernetes, MLOps experience. View posting

Other roles at speech-tech companies

  • Software Engineer, Product — Descript · San Francisco or Remote (US) · $220K–$265K + equity
    Full-stack product engineering across Descript's audio/video editor and its AI co-editor, Underlord (TypeScript, React, Node). General product engineering, not audio-ML specific. 5+ years shipping customer-facing product features end to end. View posting
  • Lead UX Designer, Product Design — Rev · Austin, TX (hybrid) or Remote (US) · compensation not listed
    Senior/staff individual-contributor design role shaping Rev's ASR and conversational-intelligence products end to end, from discovery through shipped UI in Figma. A design role, not engineering or research. View posting
  • Software Engineer — Picovoice · Vancouver, Canada (onsite; no visa sponsorship) · $75K–$150K CAD + equity
    General software engineering (C, Python, TypeScript) at an on-device voice, language and vision AI company. A background in speech recognition, synthesis or speaker recognition is called out as a plus, not required. 2+ years experience. View posting

Hiring for a speech-tech role? Tell us about it and we'll include it here and in the digest. Browse everything by specialty on the homepage.