Member of Technical Staff (Data Scientist, Evals)
Perplexity · Hybrid
This listing is no longer verified as available.
MeritLog keeps this source-backed description for reference. Availability is not verified, and there is no application link here.
Source: the employer's Ashby job board. Open the original listing for current details. Availability is not verified for this retained page.
Job details
- Work model
- Hybrid
- Salary
- Not listed by source
- Location
- San Francisco
What the role asks for
What you'd do
- Architect and maintain automated evaluation pipelines to assess answer quality across Perplexity's products, ensuring high standards for accuracy and helpfulness
- Design evaluation sets and methods specifically to measure the impact of tool calls (particularly web search retrieval) on the final answer's quality
- Develop VLM-based solutions to programmatically evaluate how final answers render visually across different platforms and devices
- Continuously review public benchmarks and academic evaluations for their applicability to the Perplexity product, adapting and incorporating them into our regular performance measurements
- Operate within a small, high-impact team where your evaluation metrics directly shape product changes, collaborating closely with technical leadership to measure and improve Answer Quality
What they're asking for
- PhD or MS in a technical field or equivalent experienceEducation
- 4+ years of experience in data science or machine learningExperience
- Strong proficiency in Python and SQL (expected to write production-grade code)Skill
- Experience building within a modern cloud data stack, specifically AWS and DatabricksSkill
- Comfortable with agentic coding workflows and using AI-assisted development tools to iterate fasterSkill
- 1+ years of experience working with LLMs at scale, specifically with LLM-as-a-judge setupsExperiencePreferred
- Prior experience working on customer-facing web products or consumer apps, with real user traffic at scaleSkillPreferred
- A strong research background, with experience applying research methods to real-world ML problemsSkillPreferred
- Experience defining evaluation metrics (e.g., factual consistency, hallucination rate, retrieval precision) and building ground truth datasetsSkillPreferred
Parsed by MeritLog from the employer’s own posting. The full description follows below.
Job description
Perplexity serves tens of millions of users daily with reliable, high-quality answers grounded in an LLM-first search engine and our specialized data sources. We aim to use the latest models as they are released, but the intelligence frontier is a jagged one, and popular benchmarks do not effectively cover our use cases. In this role, you will build specialized evals to improve answer quality across Perplexity, covering search-based LLM answers and other scenarios popular with our users. RESPONSIBILITIES - Architect and maintain automated evaluation pipelines to assess answer quality across Perplexity's products, ensuring high standards for accuracy and helpfulness - Design evaluation sets and methods specifically to measure the impact of tool calls (particularly web search retrieval) on the final answer's quality - Develop VLM-based solutions to programmatically evaluate how final answers render visually across different platforms and devices - Continuously review public benchmarks and academic evaluations for their applicability to the Perplexity product, adapting and incorporating them into our regular performance measurements - Operate within a small, high-impact team where your evaluation metrics directly shape product changes, collaborating closely with technical leadership to measure and improve Answer Quality QUALIFICATIONS - PhD or MS in a technical field or equivalent experience - 4+ years of experience in data science or machine learning - Strong proficiency in Python and SQL (expected to write production-grade code) - Experience building within a modern cloud data stack, specifically AWS and Databricks - Comfortable with agentic coding workflows and using AI-assisted development tools to iterate faster PREFERRED QUALIFICATIONS - 1+ years of experience working with LLMs at scale, specifically with LLM-as-a-judge setups - Prior experience working on customer-facing web products or consumer apps, with real user traffic at scale - A strong research background, with experience applying research methods to real-world ML problems - Experience defining evaluation metrics (e.g., factual consistency, hallucination rate, retrieval precision) and building ground truth datasets
Keep exploring
Available Data & Analytics roles
These current listings are available to explore now.