Back to jobs
BLAND
North America

Machine Learning Research Intern, Audio

San Francisco, CA, USA
2026-08-27

Role Description

**The Role: Machine Learning Research Intern, Audio** ----------------------------------------------------- As a Research Intern at Bland, you will own a focused research project across our voice stack: speech-to-text, large language models, neural audio codecs, or text-to-speech. You will work alongside our research team on the same problems they are working on, not on a side track built to keep interns busy. We scope internships around a single meaningful question that can be answered in the time you have. The goal is a result worth shipping, publishing, or both. Interns here regularly see their work reach production systems handling millions of calls. **What You Will Do** -------------------- **Own a research question end to end** * Take one well-scoped problem from literature review through implementation, experimentation, and results. * Design ablations that isolate what actually caused an improvement. * Present your findings to the research team and defend the methodology. **Work on real systems** * Train and evaluate models on large-scale, real-world telephony audio, including the accents, noise, and artifacts that make production speech hard. * Use our distributed GPU infrastructure rather than toy-scale setups. * Where the result warrants it, work with engineers to move it toward production. **Choose your depth** Depending on your background and interests, your project may focus on: * Expressive and controllable text-to-speech, including prosody and emotion modeling * Neural audio codecs and discrete or continuous speech representations * ASR robustness for telephony, accents, and code switching * Real-time and streaming inference under latency constraints * Full-duplex conversation and turn-taking dynamics **What Makes You a Great Fit** ------------------------------ **Research foundations** * Currently pursuing a MS or PhD in ML, CS, EE, or a related field, or equivalent research experience. * Comfortable reading a paper and reimplementing it without hand-holding. * Experience with self-supervised, generative, or multimodal modeling. **Audio or speech grounding** * Hands-on work with speech or audio models, whether TTS, ASR, codecs, or audio representation learning. * Strong intuition for audio quality and what makes synthetic speech sound wrong. * Prior publications or open source contributions in speech or language AI are a strong signal, though not required. **Engineering ability** * Fluent in PyTorch and comfortable in a real codebase. * Able to run your own experiments on GPU clusters without waiting to be unblocked. **How You Show Up** ------------------- * You identify the single experiment that validates an idea in days, not months. * You measure everything and let data drive decisions. * You are honest about negative results, because they are how we narrow the search. * You are obsessed with making voice agents sound truly human. * You use AI tools aggressively to amplify your own impact. **Benefits** ------------ * Competitive intern compensation * Mentorship from researchers working on frontier voice AI * Every tool you need to succeed * Beautiful office in Levi's Plaza, SF with rooftop views * A real shot at a return offer

Machine Learning Research Intern, Audio

BLAND

Sign Up →