Who's hiring in speech AI this week

Twelve roles, with the widest pay spread we've printed so far. The top is Inflection AI, which wants a principal to own its whole real-time voice stack at a $400K–$550K base, the highest floor we've listed. Near the bottom is a new-grad research role on ByteDance's Seed Speech team that still pays a $218K–$388K base. Three themes run through this week's postings. Evaluation shows up as a job in its own right at Sesame, Meta and Inflection. Two of the twelve roles are for people who lead speech teams. And three companies (ASAPP, LILT and Cloaked) are hiring engineers to make streaming ASR → LLM → TTS pipelines hold up in production. If you're looking for remote work, Deepgram's new technical enablement role is the only fully remote listing this week.

Leading speech teams

  • Principal Research & Engineering, Realtime Voice AI — Inflection AI · Palo Alto, CA · $400K–$550K base + equity
    A hands-on lead who sets the technical roadmap for Inflection's real-time voice stack, the voice side of Pi. That covers streaming ASR, TTS, speech-to-speech, speech LLMs, turn-taking, barge-in and latency, plus build-vs-buy-vs-train decisions, with a 1,000-GPU cluster for experiments. The posting asks for evaluation that goes past WER, scoring interruption handling, emotional fit and task success. You'd also mentor a team while staying close to the code. View posting
  • Speech Science Technology Manager — Motorola Solutions (Theatro) · Richardson, TX · $130K–$178K + incentive bonus
    Leads the speech team behind Theatro, Motorola's voice-driven communication platform for retail store staff. The role architects the real-time pipeline (voice-command detection, VAD, ASR, TTS, noise suppression, echo cancellation, keyword spotting) and tunes ASR engines such as Whisper, Deepgram, Cerence and Sensory for a fixed command set. It also works with hardware teams on microphone arrays. Requires 7+ years in speech with at least 2 in management. The posting has been up 30+ days, so apply soon. View posting

Research & modeling

  • Research Scientist II, Speech AI Lab — Adobe Research · San Francisco, CA · $187.1K–$270.95K in California ($142.7K–$270.95K US range) + bonus
    A senior audio research scientist for Adobe's speech generative AI and multimodal work, covering speech and audio generation, audio representations and large-scale training. The lab expects independent research leadership: first-author papers, patents and prototypes that product teams can ship. PhD preferred (a research-focused master's is accepted), and 3+ years of industry research is strongly preferred. View posting
  • Research Engineer, Language — Wearables Polyglot AI — Meta · Redmond, WA + 3 other locations · $154K–$217K + bonus + equity
    Voice LLMs for speech recognition, translation and synthesis on Reality Labs wearables, deployed both on servers and on the device itself. Datasets and evaluation frameworks are part of the job, not an afterthought. Requires 5+ years building speech or language models and shipping them to production. Multilingual and edge-deployment experience are listed as pluses. View posting
  • Research Engineer — Sesame · San Francisco, Bellevue or New York (onsite) · $190K–$320K + stock options
    An evaluation-first role on the team building Sesame's voice companion. You'd own the offline and live eval pipelines for its speech and multimodal models, build the dataset-curation tooling and monitoring, and scale training and inference for LLM-sized workloads. Requires expert-level PyTorch and evaluation metrics that "actually predict user happiness." This is a different opening from the Research Scientist role we listed on September 8. View posting
  • Audio ML Engineer (Research) — HARMAN · Northridge, CA (hybrid) · $134.25K–$196.9K
    Perception models for HARMAN's Intelligent Audio research group: quality prediction, artifact detection, acoustic scene classification and listener-preference modeling. The models have to fit embedded and cloud budgets, using quantization, pruning and distillation where needed. Asks for 5+ years of applied ML, at least 2 of them on audio, speech or acoustics. The posting says shipped product impact counts for more than credentials. View posting

Real-time voice systems

  • Senior Speech Software Engineer — ASAPP · New York or Mountain View (hybrid) · $215K–$235K + performance bonus
    Half model tuning, half infrastructure. On the model side you'd adapt ASR and TTS for noisy call-center audio and improve TTS prosody and number pronunciation. On the systems side you'd build the multi-threaded server frameworks that run thousands of concurrent streaming ASR → LLM → TTS sessions. Requires 5+ years of distributed systems in Go or Python and hands-on ASR/TTS work. You'll need to know your codecs too: Opus, G.711, jitter and packet loss. View posting
  • Machine Learning Engineer, Real-Time Speech Translation — LILT · Washington, D.C. or Boston ($129K–$161K), Indianapolis ($120K–$150K) · hybrid · US citizenship required
    Owns the backend for LILT's new live-translation product. The ASR and MT models already exist; the job is wiring them into a low-latency streaming system on Ray Serve and GPU Kubernetes, including confidence scoring that sends uncertain segments to human linguists. Requires 3+ years of production Python with asyncio and hands-on WebSocket/gRPC streaming. The first languages are Japanese, Korean and English. US citizenship is a contract requirement. View posting
  • Senior Software Engineer, Voice AI — Cloaked · New York, NY (hybrid) · $200K–$230K + equity + bonus
    Works on Call Guard, an AI agent that answers unknown callers, holds a live conversation and blocks scams. It has screened 50M+ calls so far. You'd cut end-to-end STT → LLM → TTS latency and build defenses against voice-cloned attackers. Wants 5+ years and production experience with Whisper or Deepgram, ElevenLabs or Cartesia, and LiveKit or Pipecat. Telephony and SIP experience is a plus. View posting
  • Product Engineer, Systems Architect — Wispr · San Francisco (onsite) · $220K–$300K (L5) or $270K–$350K (L6) + equity · visa sponsorship
    Owns the detection, inference, state and delivery systems behind Wispr Flow, its system-wide dictation app, and Notetaker. A lot of the work is proving those systems are correct when no ground truth exists. The posting says speech experience is not required, and it names signal processing, ranking, fraud and ML evaluation as backgrounds that translate. That makes it one of the few ways into a voice company without a speech background. View posting

Technical go-to-market

  • Technical GTM Enablement Manager — Deepgram · Remote (US) · compensation not listed
    Deepgram's first hire in this role. You'd build the onboarding, playbooks and certifications that bring its sales engineers, solutions architects and customer engineers up to speed. That includes standards for technical discovery, proof-of-concept evaluations, architecture guidance and handoffs, plus readiness work before each product launch. Requires 5+ years in customer-facing technical roles and hands-on fluency with APIs and SDKs. You don't need to write production code, but "hand-waving does not pass." Experience with STT, TTS or voice agents is a plus. A good fit if you came up through solutions engineering and want a non-engineering role at a speech company. View posting

New grad

  • Research Scientist Graduate, Seed Model — Speech (2027 start) — ByteDance · San Jose, CA · $218.4K–$387.6K base + bonus + RSUs
    A new-grad role on the Seed Speech team, building speech foundation models for both understanding and generation: data construction, instruction tuning, alignment, and gains in recognition, synthesis and robustness. A BS is the minimum and an MS is preferred, with internship experience in speech or audio a plus. You can apply to at most two ByteDance roles worldwide and they're reviewed on a rolling basis, so apply early and put your graduation date on your resume. View posting

Hiring for a speech-tech role? Tell us about it and we'll include it here and in the digest. Browse everything by specialty on the homepage.