Job Summary
We are seeking a Speech Deep Learning Scientist to develop and enhance advanced Speech AI solutions and improve conversational AI experiences for millions of users. This role will focus on speech synthesis model development, speech data processing, model evaluation, and driving innovation in speech technologies
Responsibilities
- Train Speech Synthesis Mel Spectrogram and Vocoder models
- Measure and benchmark model performance
- Maintain TTS model evaluation systems
- Analyze model accuracy and bias and recommend improvements and next steps
- Improve processes for speech data processing, augmentation, filtering, and TTS training set preparation
- Gather knowledge on TTS datasets for training and evaluation
- Characterize performance and quality metrics across platforms for various Speech AI components
- Collaborate with multiple teams on new product features and enhancements to existing products
- Participate in code development and reviews, design document reviews, use case reviews, and test plan reviews
- Help innovate, identify problems, recommend solutions, and perform triage in a collaborative team environment
Requirements
- 5+ years of relevant experience
- Excellent programming skills in Python
- Strong fundamentals in programming, optimization, and software design
- Strong knowledge of machine learning and deep learning techniques, algorithms, and tools, including exposure to Autoregressive Speech Language Models, Audio Diffusion, and Flow Matching models
- Knowledge of deep learning applications for speech synthesis, Large Language Models, and speech to speech translation
- Hands on experience with speech technologies such as speech synthesis and voice cloning
- Experience training speech models
- Experience with the PyTorch deep learning framework
- Exposure to speech digital signal processing and feature extraction techniques including FFT, MFCC, and Mel Spectrograms
- General background with version control and code review tools such as Git, Gerrit, and GitLab
- Strong collaborative and interpersonal skills with a proven ability to guide and influence within a dynamic matrix environment Preferred Qualifications
- Native or near native fluency in a non English language including Spanish, Mandarin, German, Japanese, Russian, French, UK English, Arabic, Hindi, Korean, Italian, or Portuguese
- Experience developing multilingual code switched TTS, voice cloning, and cross lingual voice cloning solutions
- Experience developing WFST and neural network based Text Normalization and Inverse Text Normalization solutions
- Experience working with G2P systems across multiple languages
- Strong personal interest in learning, researching, and creating technologies related to foreign languages, linguistics, phonetics, phonology, and language technology
- Comfortable working in a fast paced, highly collaborative, and dynamic environment
- Strong C++ programming skills
- Familiarity with GPU technologies including CUDA, CuDNN, and TensorRT
- Experience deploying machine learning models on data center, cloud, and embedded systems
Tools and Technologies:
- Python
- PyTorch
- Machine Learning
- Deep Learning
- Speech Synthesis
- Voice Cloning
- Speech to Speech Translation
- Autoregressive Speech Language Models
- Audio Diffusion
- Flow Matching
- FFT
- MFCC
- Mel Spectrograms
- Git
- Gerrit
- GitLab
- C++
- CUDA
- CuDNN
- TensorRT
Pay: $50.00 - $55.00 per hour
Experience:
- Deep learning: 5 years (Required)
- autoregressive speech language models (LMs): 2 years (Required)
Work Location: Remote