Back to search
Listing unavailableData & AnalyticsOn-site

Machine Learning Intern

Bland · On-site

This listing is no longer verified as available.

MeritLog keeps this source-backed description for reference. Availability is not verified, and there is no application link here.

Last seen by MeritLog September 8, 2026Source: AshbySource version: ashby-public-job-posting-v1

Source: the employer's Ashby job board. Open the original listing for current details. Availability is not verified for this retained page.

Job details

Work model
On-site
Salary
Not listed by source
Location
San Francisco

What the role asks for

What you'd do

  • Take one well-scoped problem from literature review through implementation, experimentation, and results.
  • Design ablations that isolate what actually caused an improvement.
  • Present your findings to the research team and defend the methodology.
  • Train and evaluate models on large-scale, real-world telephony audio, including the accents, noise, and artifacts that make production speech hard.
  • Use our distributed GPU infrastructure rather than toy-scale setups.
  • Where the result warrants it, work with engineers to move it toward production.
  • Expressive and controllable text-to-speech, including prosody and emotion modeling
  • Neural audio codecs and discrete or continuous speech representations
  • ASR robustness for telephony, accents, and code switching
  • Real-time and streaming inference under latency constraints
  • Full-duplex conversation and turn-taking dynamics
  • Currently pursuing a MS or PhD in ML, CS, EE, or a related field, or equivalent research experience.
  • Comfortable reading a paper and reimplementing it without hand-holding.
  • Experience with self-supervised, generative, or multimodal modeling.
  • Hands-on work with speech or audio models, whether TTS, ASR, codecs, or audio representation learning.
  • Strong intuition for audio quality and what makes synthetic speech sound wrong.
  • Prior publications or open source contributions in speech or language AI are a strong signal, though not required.
  • Fluent in PyTorch and comfortable in a real codebase.
  • Able to run your own experiments on GPU clusters without waiting to be unblocked.
  • You identify the single experiment that validates an idea in days, not months.
  • You measure everything and let data drive decisions.
  • You are honest about negative results, because they are how we narrow the search.
  • You are obsessed with making voice agents sound truly human.
  • You use AI tools aggressively to amplify your own impact.

Parsed by MeritLog from the employer’s own posting. The full description follows below.

Job description

THE ROLE: MACHINE LEARNING RESEARCH INTERN, AUDIO As a Research Intern at Bland, you will own a focused research project across our voice stack: speech-to-text, large language models, neural audio codecs, or text-to-speech. You will work alongside our research team on the same problems they are working on, not on a side track built to keep interns busy. We scope internships around a single meaningful question that can be answered in the time you have. The goal is a result worth shipping, publishing, or both. Interns here regularly see their work reach production systems handling millions of calls. WHAT YOU WILL DO Own a research question end to end - Take one well-scoped problem from literature review through implementation, experimentation, and results. - Design ablations that isolate what actually caused an improvement. - Present your findings to the research team and defend the methodology. Work on real systems - Train and evaluate models on large-scale, real-world telephony audio, including the accents, noise, and artifacts that make production speech hard. - Use our distributed GPU infrastructure rather than toy-scale setups. - Where the result warrants it, work with engineers to move it toward production. Choose your depth Depending on your background and interests, your project may focus on: - Expressive and controllable text-to-speech, including prosody and emotion modeling - Neural audio codecs and discrete or continuous speech representations - ASR robustness for telephony, accents, and code switching - Real-time and streaming inference under latency constraints - Full-duplex conversation and turn-taking dynamics WHAT MAKES YOU A GREAT FIT Research foundations - Currently pursuing a MS or PhD in ML, CS, EE, or a related field, or equivalent research experience. - Comfortable reading a paper and reimplementing it without hand-holding. - Experience with self-supervised, generative, or multimodal modeling. Audio or speech grounding - Hands-on work with speech or audio models, whether TTS, ASR, codecs, or audio representation learning. - Strong intuition for audio quality and what makes synthetic speech sound wrong. - Prior publications or open source contributions in speech or language AI are a strong signal, though not required. Engineering ability - Fluent in PyTorch and comfortable in a real codebase. - Able to run your own experiments on GPU clusters without waiting to be unblocked. HOW YOU SHOW UP - You identify the single experiment that validates an idea in days, not months. - You measure everything and let data drive decisions. - You are honest about negative results, because they are how we narrow the search. - You are obsessed with making voice agents sound truly human. - You use AI tools aggressively to amplify your own impact. BENEFITS - Competitive intern compensation - Mentorship from researchers working on frontier voice AI - Every tool you need to succeed - Beautiful office in Levi's Plaza, SF with rooftop views - A real shot at a return offer

Keep exploring

Available Data & Analytics roles

These current listings are available to explore now.

Search all jobs

Privacy choices

Analytics and advertising stay off unless you allow them. Private data stays out.

Read the privacy notice