Back to search
Data & AnalyticsHybrid

Member of Technical Staff (Answer Quality & Evals)

Perplexity · Hybrid

Apply
Last seen by MeritLog September 12, 2026Source: AshbySource version: ashby-public-job-posting-v1

MeritLog read this listing from Perplexity's Ashby job board and last checked it on September 12, 2026.

Source: the employer's Ashby job board. Open the original listing for current details.

Job details

Work model
Hybrid
Salary
$200K - $350K
Location
San Francisco

What the role asks for

What you'd do

  • Build shared evaluation infrastructure that helps teams run reliable evals, analyze results, and make product and model decisions
  • Develop the platform for replaying and analyzing agent traces to reproduce production behavior and diagnose failures
  • Build and operate scalable systems for processing, storing, and monitoring interaction, trace, and evaluation data
  • Partner with data scientists, engineers, and product teams to turn answer-quality problems into evaluations, analyses, and product improvements
  • Operate in a small, high-impact team where your work directly shapes how Perplexity measures and improves Answer Quality

What they're asking for

  • 4+ years of software, data, or machine learning engineering experience shipping and operating production systemsExperience
  • Strong proficiency in Python and SQL, with solid fundamentals in system design, data modeling, and distributed systemsSkill
  • Experience building big-data systems, including distributed compute, large-scale storage, and high-volume pipelinesSkill
  • Demonstrated ownership of ambiguous technical projects from initial design through production operationSkill
  • Ability to work effectively with data scientists, engineers, and product partnersSkill
  • Experience building evaluation, experimentation, observability, or machine learning infrastructureSkillPreferred
  • Familiarity with LLM and agent systems, including tool use, execution traces, replay, and simulationSkillPreferred
  • Experience building on top of large-scale data processing platforms such as Databricks, Snowflake, or ClickHouseSkillPreferred

Parsed by MeritLog from the employer’s own posting. The full description follows below.

Job description

Perplexity serves tens of millions of users daily with reliable, high-quality answers grounded in an LLM-first search engine and specialized data sources. The Answer Quality team ensures that our prompts, tools, search systems, datasets, and models work together to create the best possible experience for our users. As our product and agent capabilities evolve, we need evaluation systems that are fast, reliable, production-faithful, and actionable. In this role, you will build and improve the technical foundations that support Answer Quality across Perplexity. This includes our shared evaluation infrastructure and the platform used to replay and analyze agent traces. You will work closely with data scientists, engineers, and product teams to identify quality problems, measure their impact, and turn evaluation findings into product improvements. RESPONSIBILITIES - Build shared evaluation infrastructure that helps teams run reliable evals, analyze results, and make product and model decisions - Develop the platform for replaying and analyzing agent traces to reproduce production behavior and diagnose failures - Build and operate scalable systems for processing, storing, and monitoring interaction, trace, and evaluation data - Partner with data scientists, engineers, and product teams to turn answer-quality problems into evaluations, analyses, and product improvements - Operate in a small, high-impact team where your work directly shapes how Perplexity measures and improves Answer Quality QUALIFICATIONS - 4+ years of software, data, or machine learning engineering experience shipping and operating production systems - Strong proficiency in Python and SQL, with solid fundamentals in system design, data modeling, and distributed systems - Experience building big-data systems, including distributed compute, large-scale storage, and high-volume pipelines - Demonstrated ownership of ambiguous technical projects from initial design through production operation - Ability to work effectively with data scientists, engineers, and product partners PREFERRED QUALIFICATIONS - Experience building evaluation, experimentation, observability, or machine learning infrastructure - Familiarity with LLM and agent systems, including tool use, execution traces, replay, and simulation - Experience building on top of large-scale data processing platforms such as Databricks, Snowflake, or ClickHouse

Keep exploring

More Data & Analytics roles

Search all jobs

Privacy choices

Analytics and advertising stay off unless you allow them. Private data stays out.

Read the privacy notice