The board
- Where
- San Francisco
- Posted
- Aug 28
The Role: Machine Learning Research Intern, Audio
As a Research Intern at Bland, you will own a focused research project across our voice stack: speech-to-text, large language models, neural audio codecs, or text-to-speech. You will work alongside our research team on the same problems they are working on, not on a side track built to keep interns busy.
We scope internships around a single meaningful question that can be answered in the time you have. The goal is a result worth shipping, publishing, or both. Interns here regularly see their work reach production systems handling millions of calls.
What You Will Do
What you'd do
- Take one well-scoped problem from literature review through implementation, experimentation, and results.
- Design ablations that isolate what actually caused an improvement.
- Present your findings to the research team and defend the methodology.
- Train and evaluate models on large-scale, real-world telephony audio, including the accents, noise, and artifacts that make production speech hard.
- Use our distributed GPU infrastructure rather than toy-scale setups.
- Where the result warrants it, work with engineers to move it toward production.
- Expressive and controllable text-to-speech, including prosody and emotion modeling
- Neural audio codecs and discrete or continuous speech representations
- ASR robustness for telephony, accents, and code switching
- Real-time and streaming inference under latency constraints
What you get
- Competitive intern compensation
- Mentorship from researchers working on frontier voice AI
- Every tool you need to succeed
- Beautiful office in Levi's Plaza, SF with rooftop views
- A real shot at a return offer