Data Scientist - Agent Evaluations & Quality
Clera · Remote
This listing is no longer available.
MeritLog keeps this source-backed description for reference. Availability is not verified, and there is no application link here.
Source: Arbeitnow. Open the original listing for current details. Availability is not verified for this retained page.
Job details
- Work model
- Remote
- Salary
- Not listed by source
- Location
- remote
- Occupation
- Data Scientists(O*NET 15-2051.00)
Job description
About the Role This company is building an AI executive assistant that operates across email, calendars, meetings, and business software. As a Data Scientist - Agent Evaluations & Quality, you will own the measurement system that determines whether the assistant is genuinely improving in ambiguous, real-world environments. You'll partner directly with AI Agent Capabilities engineers to generate the evidence that shapes product decisions, model choices, and release quality. This is a high-ownership, deeply technical role at the intersection of applied data science, LLM evaluation, and product quality - ideal for someone who thrives on turning hard, open-ended quality questions into rigorous, actionable answers. What You'll Do • Architect and maintain automated evaluation pipelines that measure agent quality across product surfaces. • Translate agent capabilities into explicit pass, partial-pass, and failure criteria for complex multi-step tasks. • Build representative gold datasets and regression suites covering real workflows, edge cases, and adversarial scenarios. • Define meaningful metrics - task success, tool-selection accuracy, instruction adherence, factual consistency, latency, cost, and reliability. • Design deterministic and model-based graders, calibrate LLM-as-a-judge systems, and track grader agreement. • Compare models, prompts, and implementations using rigorous offline experiments and production evidence. • Analyze traces and production outcomes to identify root causes and build a practical failure taxonomy. • Turn production failures into regression cases and continuously close gaps in evaluation coverage. • Build dashboards and release-quality signals that make results actionable for engineering, product, and leadership. • Recommend improvements to capability engineers and verify that fixes raise quality without unacceptable regressions. What We're Looking For Required • 4+ years in Applied Data Science or Machine Learning roles, with a track record of building and delivering evaluation systems, automated data pipelines, or production ML infrastructure. • Experience designing and implementing automated evaluation frameworks, success criteria, and regression suites for complex AI/ML or agentic systems. • Production-grade proficiency in Python and SQL, with experience building and maintaining automated analytical pipelines on large datasets. • Applied statistical and experimental skills: significance testing, variance analysis, and sampling to evaluate non-deterministic AI/ML systems. • Experience developing labeled datasets, annotation guidelines, and quality-control processes for ground-truth data in dynamic product environments. • Solid understanding of LLM agent behaviors: tool use, multi-step execution, retrieval, and practical failure modes. • Demonstrated ability to analyze model traces, tool calls, and outputs to identify root causes across model, prompt, tool, and data layers. • Experience using production telemetry and observability data to monitor system quality, build dashboards, and analyze real-world user outcomes. Nice to Have • Hands-on experience with LLM-as-a-judge systems, model-based grading, or AI benchmarking platforms. • Experience shipping or operating production ML products, agentic systems, or customer-facing consumer software. • Experience reviewing and adapting public research benchmarks or academic evaluation methodologies to real-world product problems. What makes you a great fit • You're product-oriented - you prioritize metrics tied to real user outcomes, not just convenient measurements. • You drive ambiguous quality questions from evaluation design all the way into product decisions. • You write maintainable, production-quality code - not just ad-hoc notebooks. • You collaborate naturally with engineers and are comfortable digging into traces and system internals. Location This role is on-site. Visa sponsorship is not available for this position. Compensation & Benefits Compensation details were not provided for this listing. A competitive package commensurate with experience is expected at this stage of company growth. Find more English Speaking Jobs in Germany on Arbeitnow