Who's hiring in speech AI this week

This was a research week. Spotify has stood up an Artist-First AI Music Lab and is hiring generative-audio scientists at every level; Google, Sesame and Rad AI all posted senior speech and audio research roles in the same few days. Where the top of the range is listed it's high — Decagon has a speech research-engineer role at $200–400K, Sesame up to $320K, Amazon up to $260K for real-time speech inference. The other thread is the real-time voice stack: streaming ASR, endpointing, barge-in and low-latency serving turn up across five postings. Eleven roles this week, most of them remote-friendly.

Research

  • Research Scientist, Generative Audio — Spotify · New York or Seoul · $133K–$190K + equity
    Spotify's new Artist-First AI Music Lab is hiring across all seniority levels for diffusion and flow-matching research in vocal and speech synthesis, post-training alignment (DPO, RLHF, KTO) and text-guided audio editing. PhD, with publications at venues like ICASSP, Interspeech or ISMIR. View posting
  • ML Research Scientist, Audio Algorithms — Google · Mountain View or Irvine, CA · $174K–$252K + equity
    Develops and trains ML for audio — neural noise suppression, acoustic modeling, speech enhancement — in JAX or TensorFlow. PhD plus 2+ years leading a research agenda. View posting
  • Research Scientist — Sesame · SF, Bellevue or New York (onsite) · $190K–$320K
    Deep-learning research across NLP, speech and computer vision for Sesame's lifelike voice agents, working the full stack from model architecture to training and inference infra. Published large-scale deep-learning work; Master's or PhD preferred. View posting
  • Staff ML Research Scientist — Rad AI · SF (4 days onsite) or US remote · $190K–$260K + equity
    Applied research across LLMs, retrieval, speech and multimodal modeling for radiology-report AI, owned end to end from problem framing to production. MS or PhD plus 7+ years (or PhD plus 5). View posting

Voice agents & real-time speech

  • Research Engineer, Audio and Speech — Decagon · SF or New York (onsite) · $200K–$400K + equity
    Builds the models and agent harness behind Decagon's real-time voice agents — streaming ASR, voice-activity detection, endpointing, turn-taking and speech generation, plus full-duplex and speech-to-speech work. 2+ years in speech, audio ML or multimodal ML. View posting
  • Senior Inference Engineer, AGI — Amazon · Boston, Sunnyvale or Seattle · $167K–$260K by location + RSUs
    Owns low-latency inference for real-time multimodal conversational AI: co-designs speech and audio model architecture for servability, builds the streaming runtime, and the RL and evaluation infra behind it. PhD, or Master's plus 6+ years; 2+ years optimizing model inference. View posting
  • Sr Software Engineer, Audio Intelligence — NRG · Remote to start, then hybrid Seattle · $150K–$185K
    Conversational and audio AI for the smart home — STT, TTS, LLMs and audio understanding across mobile, panel and camera devices, with wake-word and streaming-audio work on the preferred list. 5+ years with a Bachelor's, or 2+ with a Master's. View posting
  • Senior Agent Design Engineer — PolyAI · Remote (US) · $110K–$125K + equity
    Senior IC role at the intersection of conversational UX and implementation: owns user journeys, dialogue and tone, plus the agentic systems underneath — multi-agent loops, evaluation pipelines, tool access. 5+ years in a technical or design-adjacent role; Python you can build in; voice craft (latency, barge-in, ASR-error recovery) is a plus. View posting
  • Staff Machine Learning Engineer, Gen AI (Voice & Speech) — Weave · Remote (US) · compensation not listed
    Staff-level generative-AI role focused on voice and speech for Weave's small-business communication platform. View posting

Platform & audio infrastructure

  • Audio Systems Engineer — Meta · Sunnyvale, CA or Redmond, WA · $144K–$204K + bonus + equity
    Audio system design, acoustic simulation and DSP for next-generation wearables: tuning and debugging real-time audio capture, with beamforming, echo cancellation and noise suppression on the preferred list. 6+ years; MATLAB and Python. This is a hardware-audio role, not ASR. View posting
  • Software Engineer, Platform — Speechify · Remote (US) · compensation not listed
    Backend behind Speechify's TTS product — the public TTS API, payments, subscriptions, auth and consumption metering across 50M+ users. TS/Node and GCP required. This is a backend platform role, not speech modeling. View posting

Hiring for a speech-tech role? Tell us about it and we'll include it here and in the digest. Browse everything by specialty on the homepage.