A research-heavy week. Deepgram is hiring on two fronts — a Director to own its text-to-speech research program end to end, and research staff to build the data factory feeding its next generation of models. Google DeepMind wants an audio-to-audio research scientist for open-mic, low-latency dialog. David AI, fresh off a $50M Series B, is paying up to $360K for an audio ML researcher. On the engineering side, Together AI is making a foundational hire to own voice-model serving on H100s and B200s, and Ford is looking for someone who can sit in an acoustics lab with a head-and-torso simulator and write the C++ that automates it.
Research & modeling
- Director, Text-to-Speech Synthesis Research — Deepgram
· Remote (US), Ann Arbor, MI or San Francisco · $262.6K–$328.3K (SF/NYC/Seattle), $213K–$266.3K elsewhere + equity + bonus
Owns Deepgram's TTS program end to end — research strategy, technical bets, and the models that ship — across neural audio modeling, prosody, controllability, multilingual speech and voice consistency. A hands-on leadership role: you lead ICs and tech lead managers while still designing experiments and diagnosing model failures. Requires a track record personally training large-scale speech-generation models and setting research direction under real uncertainty. View posting - Audio to Audio Research Scientist — Google DeepMind
· Mountain View, CA or New York, NY · $207K–$300K + 20% bonus target + equity
Builds audio-first models that can plan and orchestrate complex dialog, including tool use, and works with infra teams on open-mic conversation that stays fluid at low latency. Also owns the datasets and benchmarks that measure A2A conversationality. PhD in CS/EE/CE or equivalent; generative audio, multimodal modeling or on-device AI preferred. View posting - Research Scientist — David AI
· San Francisco (onsite) · $210K–$360K + equity
Speech and audio models plus the production inference systems around them at an audio-data research company that raised a $50M Series B from Meritech and NVIDIA. Expect both classical DSP and modern audio ML, and end-to-end ownership from proof-of-concept to deployment. Requires 5+ years of professional audio ML experience, strong Python and PyTorch. View posting - Applied Scientist, Speech AI — Amazon (Prime Video)
· Seattle, WA · $142.8K–$193.2K + sign-on + RSUs
Leads speech and audio generation research for Prime Video content localization — end-to-end speech-to-speech architecture, plus source separation, enhancement and mixing for production. You define the research roadmap for the area and publish. 3+ years building models for business applications; PhD, or Master's plus 4+ years. Speech synthesis and foundation-model experience called out as a big plus. View posting - Speech & Audio Research Engineer — Qualcomm
· San Diego, CA · $155.4K–$233.2K + bonus + RSU eligibility
Fundamental DSP and machine-learning research in Qualcomm's Multimedia R&D Group: conversational speech codecs for wireless networks, real-time communication for XR, and embedded neural models including quantization and optimization. One of the few roles on this list where classical signal processing — linear prediction, adaptive filtering, spectrum estimation — is a hard requirement alongside deep learning. Master's plus 3+ years, or PhD plus 2+. View posting
Training data & evaluation
- Research Staff, Data Science — Deepgram
· Remote (US), Ann Arbor, MI or San Francisco · $150K–$220K + equity + 10% bonus
Builds the "data factory" behind Deepgram's next generation of voice models: acquisition, preparation and synthesis pipelines, characterization of messy conversational audio, and benchmarking methodology for conversational voice systems. Wants someone who has owned a full data stack from a blank page. Speech and audio experience is listed as a nice-to-have, not a requirement. View posting
Voice agents, serving & real-time speech
- Senior Machine Learning Engineer, Voice AI — Together AI
· San Francisco · $200K–$260K + equity
Owns the model serving layer for voice workloads — TRT-LLM and SGLang internals, batching for streaming audio, GPU profiling on H100s/H200s/B200s — for models like Whisper, Parakeet, Orpheus and Kokoro. A foundational hire on a small team. Requires 5+ years in ML engineering with hands-on LLM serving-engine work; speech and audio experience is a strong plus but explicitly not required. View posting - Media Software Engineer, Speech (Senior–Staff) — Cantina
· Sunnyvale or San Francisco (hybrid) · $180K–$270K + equity
Low-level media pipelines and WebRTC infrastructure behind live conversations with AI characters, across iOS, Android and web — cutting latency across audio streaming and speech pipelines, and integrating ASR and synthesis models. Heavy C/C++ systems role: memory management, concurrency, async I/O. 3+ years as a software engineer; speech ML familiarity preferred, not required. View posting - Machine Learning Engineer — Guava
· Downtown Los Angeles (onsite) · $140K–$180K + equity
Full model lifecycle across ASR, TTS, intent recognition, dialogue and summarization for regulated production phone calls, including LLM components and model governance. Small enough that you root-cause real production failures — ASR errors, hallucinations, turn-taking — and turn them into model fixes. Applications go to hiring@goguava.ai rather than an ATS. View posting
Automotive & in-cabin audio
- Product Development Engineer, Acoustics Test & Automation — Ford
· Dearborn, MI (hybrid, 4+ days onsite) · $99.6K–$146.6K (grade 7) or $115K–$162.9K (grade 8)
Validates in-cabin voice and audio performance and then automates the testing: hands-on work with head-and-torso simulators, audio analyzers and oscilloscopes, paired with a C++/Python framework that catches signal and timing drift. Expects familiarity with ITU-T P.1100/P.1110/P.1120 for hands-free telephony and voice recognition. No visa sponsorship for this one. View posting
Hiring for a speech-tech role? Tell us about it and we'll include it here and in the digest. Browse everything by specialty on the homepage.