Staff AI Engineer (Hybrid)
Summary
Staff-level applied ML engineer building GenAI and voice agents (ASR, TTS, SLMs, speech-to-speech) for Stryker's medical devices, deployed both on-device and in the cloud. Owns architecture, evaluation frameworks, and safety guardrails to make stochastic models behave deterministically enough for FDA-regulated clinical use, primarily in Python (with C++ for embedded).
We're hiring a Staff AI Engineer to build GenAI and voice agents for medical devices, deployed both on-device and in the cloud. You'll own the technical direction for these systems — connecting clinical use cases to the models behind them (ASR, TTS, SLMs, speech-to-speech) while working within tight on-device limits latency, memory, and reliability. This is a hands-on applied ML role, and the core challenge is making stochastic models behave predictably enough for clinical use: bounding them with deterministic architecture, building evaluation sets and frameworks, and designing safety guardrails that hold up in a regulated environment. At the staff level, you'll set the architecture and evaluation standards the rest of the team builds against, take on the hardest technical bets first, and align device software, data, validation, clinical, and regulatory teams around the safety, effectiveness, and quality of AI-enabled features across the product lifecycle — consistent with FDA guidance and good machine learning practice.
What You Will Do
Own the technical direction of GenAI and multimodal agent (voice, text, vision) capabilities: translate product needs into robust, testable AI system designs, drive the architecture across components, and carry the highest-risk pieces from prototype through validation-ready implementation.
Architect stateful agentic systems (intent handling, tool/function calling, dialog and device-state management, interruption handling and recovery) that behave deterministically where safety requires it, with well-defined interface contracts between AI components and device software.
Manage and mitigate stochastic model behavior: design layered guardrails (deterministic validation, plausibility bounds, model-based checks), define what the model is and is not permitted to decide, and make those boundaries testable.
Set the standard for evaluation of multimodal GenAI systems: comprehensive test sets and automated harnesses covering task and intent accuracy, robustness under realistic clinical audio conditions, conversational quality, and responsiveness.
Develop and evaluate real-time speech and language components (speech recognition, synthesis, and dialog/turn handling), balancing model quality against the latency, memory, and reliability constraints of medical hardware.
Evaluate, select, integrate, and fine-tune off-the-shelf and small-footprint models (SLMs, domain-adapted ASR) for domain-specific terminology; own the buy/adapt/build decisions and their justification.
Establish safety, bias, and performance metrics tailored to voice and generative systems, and produce documentation supporting QMS and regulatory submissions.
Instrument systems for traceability: structured logging, audit trails of agent actions, and reproducible evaluation runs suitable for a regulated development process.
Mentor junior engineers, ensure engineering quality through design and code review, and communicate AI constraints and trade-offs clearly to product, clinical, and regulatory stakeholders.
Stay abreast of the rapidly evolving GenAI model, speech, and agent-architecture landscape; identify which advances matter for the roadmap and pragmatically incorporate them.
What You Need (Minimum Required Qualifications)
Bachelor's Degree in Computer Science, Machine Learning, Electrical Engineering, Biomedical Engineering, Mathematics, or related field.
4+ years of AI/ML engineering experience, OR Master's Degree in the above fields and 2+ years.
Preferred Qualifications (Strongly Desired)
Strong proficiency in Python; optionally, working proficiency in C++ or other relevant languages for performance-critical and embedded integration work
Hands-on experience designing custom evaluation metrics, evaluation harnesses, and test-set generation (both synthetic and human) for stochastic AI systems.
Experience bringing statistical rigor to AI evaluation and validation (acceptance criteria, confidence intervals, subgroup and robustness analysis), ideally in collaboration with validation or quality functions.
Experience building production systems around LLM function calling / tool use, including schema design, context management for small-context models, and prompt/configuration versioning.
Demonstrated recent experience in at least one of: NLP/LLM systems, speech processing (ASR/TTS), or the development and evaluation of generative AI applications.
Experience fine-tuning ASR or small language models for domain-specific vocabulary.
Experience with real-time voice interfaces: streaming audio, turn-taking and interruption handling, and robustness to noisy environments.
Track record of technical leadership: setting direction for a significant system or product area and delivering it through/with other engineers.
Experience setting technical direction across multiple teams or leading the architecture of a multi-component product area.
Experience deploying and optimizing models on edge compute devices under latency/memory constraints.
Strong problem-solving, detail orientation, and critical-thinking skills; excellent communication and interpersonal skills, with the ability to effectively communicate complex technical concepts to non-technical stakeholders; general knowledge of the healthcare market.
$133,400 - $222,300 USD Annual