GPU-based microservices handle the core speech pipeline—STT consumption, word-level alignment, diarization, and speaker identification. The primary challenge isn't just serving models, but establishing rigorous data-driven evaluation to separate real accuracy gains from benchmark noise on difficult sports audio. The Senior ML Engineer will own the speech services, lead fair model bake-offs, and hold full authority over which models reach production.
The platform delivers real-time AI video processing and automated content generation for professional sports leagues globally. The underlying pipeline operates in high-noise live broadcast environments, demanding tight latency budgets, high throughput, and robust handling of overlapping speech and crowd noise.
Technology Stack: The platform delivers real-time AI video processing and automated content generation for professional sports leagues globally. The underlying pipeline operates in high-noise live broadcast environments, demanding tight latency budgets, high throughput, and robust handling of overlapping speech and crowd noise.
Own and scale the core speech pipeline covering alignment, diarization, speaker identification, and enrollment signature-matching workflows.
Optimize GPU inference performance for throughput, memory footprint, and cost through batching, precision tuning, and compilation.
Build and maintain labeled sports test benchmarks stratified by speaker count, audio quality, language, and background noise.
Define and track production evaluation metrics including DER, WER, word-level speaker attribution, speaker-count error, latency, and GPU cost.
Conduct rigorous model bake-offs and write clear decision memos detailing trade-offs, licensing constraints, and statistical uncertainty.
Maintain a shadow or A/B deployment framework to gate every production release with empirical evidence.
3–5+ years of experience shipping ML systems into production, with 2+ years dedicated to speech or audio pipelines.
Deep Python and PyTorch expertise, including hands-on GPU profiling, memory optimization, and inference acceleration.
Hands-on experience with at least two key speech domains: diarization (pyannote, NeMo/Sortformer, VBx), speaker embeddings (ECAPA-TDNN), or ASR and forced alignment (Whisper, Parakeet, wav2vec-family).
Strict evaluation rigor: demonstrated experience building test sets, computing DER/WER/EER correctly (handling collars and reference pitfalls), and running statistical hypothesis tests.
Demonstrated ability to critically analyze academic papers or model cards and reproduce published claims on internal datasets.
Hands-on experience deploying containerized GPU services in production environments using Docker and queue/API architectures.
Hands-on experience with Azure ecosystem tools (Service Bus, Blob Storage) and event-driven microservices.
Experience with domain adaptation or fine-tuning speech models on noisy, broadcast, or far-field sports audio.
Experience handling multilingual STT pipelines, particularly with Hebrew, Arabic, or Spanish.
Experience designing shadow deployment paths, experiment tracking workflows, and annotation team processes.
Direct technical ownership of production GPU microservices where benchmark evaluations directly determine what reaches live users.
Applied research opportunity on challenging sports audio conditions (crowd noise, overlapping commentators, PA systems) rather than clean laboratory datasets.
Clear, evidence-based engineering culture where architectural and model choices are driven by data-backed decision memos rather than hype.