Nine roles this week, spread across five different kinds of speech-tech work: research, voice-agent engineering, speech analytics, platform infrastructure, and a few broader product and design roles at speech-first companies. Dolby is hiring a spatial-audio researcher for XR content; Wispr's ML engineer role tops out at $400K building the sub-500ms inference behind its voice interface; Deepgram needs someone to port its speech models onto hardware it doesn't control (non-NVIDIA edge devices); and Retell AI is paying $180–300K for a forward-deployed engineer fluent in Mandarin to support its China accounts. A genuinely mixed week.
Research
- Research Scientist, Spatial Audio & AI — Dolby
· Atlanta, GA · $140K–$170K + bonus + equity
Designs AI models and multimodal foundation models for spatial and XR audio content creation in Dolby's Advanced Technology Group, combining deep learning with perceptually grounded audio signal processing. PhD preferred; strong audio signal analysis background, Python, PyTorch or TensorFlow. View posting
Voice agents & real-time speech
- ML Engineer — Wispr
· San Francisco (onsite) · $250K–$400K + equity
Builds the voice interface behind Wispr's Flow and Notetaker products: sub-500ms LLM inference at scale, plus fine-tuning and RL to personalize the underlying speech models. Previous founding or startup experience preferred; fluency in Python and LLM development required. View posting - Senior Forward Deployed Engineer (Mandarin fluency required) — Retell AI
· Redwood City, CA (onsite, relocation provided) · $180K–$300K + equity
Embeds with enterprise customers — a large share based in China — to take Retell's voice-AI call-center agents from prototype to production: LLM prompting, full-stack integrations with CRMs and scheduling tools. 2+ years in a customer-facing technical role; fluent Mandarin and English required. View posting
Speech analytics
- Backend Engineer, Communications Capture — Gong
· Dublin, Ireland (hybrid, 3 office days/week) · compensation not listed
Builds and scales the pipelines that capture customer calls, emails and video-conferencing audio into Gong's Revenue Intelligence platform — the raw data underneath its speech-based analytics. Java/Spring Boot, Kafka/SQS, AWS/GCP/Azure. 5+ years backend engineering; no speech-specific background required. View posting
Platform & edge infrastructure
- Applied ML Engineer, Edge Devices — Deepgram
· Remote (US) · $155K–$245K + equity + bonus
Ports Deepgram's speech models to non-NVIDIA and edge hardware — quantization, operator swaps, graph rewrites — and validates accuracy and latency on real devices. Required: hands-on production experience deploying models to edge or non-NVIDIA hardware; strong Python and PyTorch. Speech or streaming-model experience is a plus, not required. View posting - ML Data & Platform Engineer — Speechmatics
· Cambridge, UK (hybrid, 2–3 office days/week) · compensation not listed
Owns the pipelines and platform behind Speechmatics' speech-AI models — sourcing and validating training data, then training, evaluating and serving the models in production, including GPU/distributed-training optimization. Strong Python/SQL, Docker/Kubernetes, MLOps experience. View posting
Other roles at speech-tech companies
- Software Engineer, Product — Descript
· San Francisco or Remote (US) · $220K–$265K + equity
Full-stack product engineering across Descript's audio/video editor and its AI co-editor, Underlord (TypeScript, React, Node). General product engineering, not audio-ML specific. 5+ years shipping customer-facing product features end to end. View posting - Lead UX Designer, Product Design — Rev
· Austin, TX (hybrid) or Remote (US) · compensation not listed
Senior/staff individual-contributor design role shaping Rev's ASR and conversational-intelligence products end to end, from discovery through shipped UI in Figma. A design role, not engineering or research. View posting - Software Engineer — Picovoice
· Vancouver, Canada (onsite; no visa sponsorship) · $75K–$150K CAD + equity
General software engineering (C, Python, TypeScript) at an on-device voice, language and vision AI company. A background in speech recognition, synthesis or speaker recognition is called out as a plus, not required. 2+ years experience. View posting
Hiring for a speech-tech role? Tell us about it and we'll include it here and in the digest. Browse everything by specialty on the homepage.