Who's hiring in speech AI this week

Fourteen roles, and the money is loud this week. xAI has the widest band we've ever printed — $150K to $450K for a voice-model engineer on Grok. Decagon is hiring twice at $200–400K, once for speech research and once for the staff engineer who owns the voice runtime. Cohere has opened an audio-specific inference role, which is a notable move for a company known for text models. And Innodata is hiring something you almost never see posted: a Principal Speech Data Linguist at $160–185K, IPA fluency required.

Research & modeling

  • Software Engineer, Voice Model — xAI · Palo Alto, CA · $150K–$450K + equity
    Owns the Grok voice model pipeline end to end: speech data curation and synthetic generation, pre-training and post-training of speech-language models with SFT and RL, evaluation harnesses for accuracy, latency and expressiveness, then production integration. Wants a Python expert comfortable with JAX/PyTorch, Spark and Ray, and distributed training on Kubernetes. The range is unusually wide even by frontier-lab standards — assume level is negotiable. View posting
  • Senior Research Engineer, Audio and Speech — Decagon · San Francisco or New York City (in-office) · $200K–$400K + equity
    Builds the models and agent harnesses behind Decagon's real-time voice agents — full-duplex systems handling turn-taking, interruptions and overlapping speech, plus work on recognition, VAD, endpointing and generation across speakers and languages. Requires 4+ years in speech or multimodal ML and hands-on experience with autoregressive, diffusion, flow-matching or codec-based speech models. Their team publishes — worth reading their write-ups on flow-matching TTS with RL before you apply. View posting
  • Senior Applied Scientist, Real-Time Conversational AI — Amazon (AGI) · Sunnyvale, CA (onsite) · $192.2K–$260K + sign-on + RSUs
    Large-scale multimodal foundation models for real-time speech and audio generation, including reward models that capture naturalness and conversational quality, and RL for natural timing. You own a research area across the full lifecycle, from pre-training and architecture through post-training alignment and real-time deployment. 5+ years in ML, or a PhD plus 6+ years; requires hands-on foundation-model training experience. View posting
  • Multimodal AI Researcher, Audio — Dolby · Atlanta, GA · $140.7K–$170K + bonus (equity for some roles)
    Generative modeling for audio in Dolby's Advanced Technology Group: text-to-audio, video-to-audio and image-to-audio architectures, multimodal audio-video-text representations, source separation, speech enhancement and text-to-music, on the Machine Reasoning and Perception team. PhD required, along with a publication record at NeurIPS/ICLR/ICML or ICASSP/Interspeech. Research-track role — papers, IP and tech transfer to product groups. View posting

Speech data & evaluation

  • Research Scientist, Speech & Audio — Innodata · Remote (US) · $160K–$185K
    Designs the data specifications and evaluation methodology that frontier labs buy — benchmarks for ASR, TTS, speech-to-speech, diarization and audio-language models, with a focus on robustness across accents, noise and code-switching, plus ablation studies proving which data decisions actually moved the model. Roughly 5+ years in speech or audio ML; hands-on with ESPnet, NeMo, SpeechBrain or Kaldi, and forced alignment. View posting
  • Principal Speech Data Linguist — Innodata · Remote (US) · $160K–$185K
    A rare one: owns transcription and segmentation standards end to end — verbatim and phonetic (IPA) conventions, timestamping, speaker labeling, code-switching — plus error taxonomies, inter-annotator agreement and the human-in-the-loop design question of where human expertise still beats ASR. Typically 8+ years in speech-data quality, a linguistics or phonetics degree, genuine IPA fluency, and working knowledge of Whisper, forced alignment and ELAN. View posting

Inference, serving & real-time platforms

  • Audio Inference Engineer, Model Efficiency — Cohere · Remote or New York, San Francisco, Toronto, Montreal · $205K–$380K (CA/NY/WA), $175K–$325K other US states, CA$295K–CA$535K Toronto/Montreal + equity
    Pushes latency, throughput and quality on Cohere's audio model serving — profiling the system, finding bottlenecks, and building for streaming and duplex real-time workloads. C++ and Python, with GPU programming, multi-GPU parallelization and vLLM/SGLang/TensorRT-LLM experience as strong pluses. No minimum in-office requirement, but the team clusters in EST and PST. View posting
  • Staff Software Engineer, Voice Agent — Decagon · San Francisco · $200K–$400K + equity
    The architecture side of the same voice platform: owns Decagon's real-time voice runtime and its multi-quarter roadmap, sets reliability, testing and observability standards for live calls, and builds the frameworks that make voice systems debuggable. 8+ years with real technical leadership; speech recognition, VAD and streaming-protocol experience is listed as "even better if," not required — this is a real-time systems role first. View posting
  • Senior Software Engineer, Real Time Audio — Qualitate · New York City (in-office) · compensation not listed
    Owns the live layer behind an AI moderator that runs thousands of voice-based expert interviews a month: capture, streaming, speech recognition, model orchestration and synthesis, plus the WebSocket infrastructure holding sessions stable and the path from thousands to tens of thousands of concurrent conversations. 5+ years, with shipped real-time production experience on WebSockets, WebRTC or streaming media. View posting

Voice products & applied ASR

  • Senior AI Engineer, Voice Platform — ClickUp · Remote (US) · $200K–$250K
    Owns the AI systems behind ClickUp's voice platform: streaming ASR, VAD and audio processing, accuracy gains through context injection (user names, teams, custom vocabulary, language detection), LLM post-processing for filler removal and formatting, and voice-to-action parsing into structured workspace commands. Also benchmarks and integrates third-party ASR (Whisper, AssemblyAI, Fireworks) on cost, latency and accuracy. View posting
  • Machine Learning Engineer, Voice — Speak · San Francisco · $200K–$300K + equity
    Trains and deploys the ASR models behind an AI language tutor, owns the pronunciation model that gives learners precise feedback, builds metrics for ASR performance across tasks and languages, and expands the system into new languages and markets. Wants deep experience training large models on GPUs and owning ML pipelines POC-to-production — speech or audio background is listed only as a bonus. View posting
  • Research Engineer — Assort Health · San Francisco · $190K–$240K + equity
    Builds the models and pipelines behind healthcare voice agents that handle patient calls — instruction following, tool calling, retrieval and memory, plus speech and audio research questions and the evaluation protocols that measure them. 5+ years total with 3+ in AI/ML, and prior LLM post-training in production. The company reports $222M raised and a $120M Series C at a $1.2B valuation. View posting
  • ASR Engineer — Clera · San Francisco Bay Area (hybrid, 3 days) · $150K–$200K · no visa sponsorship
    Foundational hire at an early-stage AI consumer-hardware startup, owning the transcription pipeline end to end: a cloud ASR system with a narrowly scoped on-device component, judged on latency, small-word accuracy and voice-print reliability. 3+ years building production ASR pipelines, plus Core ML or TensorFlow Lite experience. Expect coordination with R&D and hardware teams in China, and very little written spec. View posting

Automotive & in-cabin audio

  • Audio DSP Engineer — ALTEN Technology USA · Palo Alto, CA (onsite 4 days, remote Fridays) · $160K–$175K
    Classical audio DSP for an EV infotainment stack: acoustic echo cancellation and residual echo suppression, adaptive beamforming (MVDR), noise and wind-noise reduction, VAD, AGC, direction-of-arrival estimation and speech enhancement, plus the mixing and routing layer. Prior voice communication or voice recognition commercialization and automotive IVI experience are the differentiators here; DSP Concepts Audio Weaver is a plus. View posting

Hiring for a speech-tech role? Tell us about it and we'll include it here and in the digest. Browse everything by specialty on the homepage.