Research Engineer - Data Quality & Evals

San Francisco, CA, US Mid Level Research Engineer

Interested in this Research Engineer role at Epsilon Labs?

Apply Now →

Skills & Technologies

Python

About This Role

AI job market dashboard showing open roles by category

About Us

------------

We're tackling one of healthcare's most critical challenges in medical imaging and diagnostics. Our company operates at the intersection of cutting\-edge AI and clinical practice, building technology that directly impacts patient outcomes. We've assembled one of the industry's most comprehensive and diverse medical imaging datasets and have a proven product\-market fit with a substantial customer pipeline already in place.

Role Overview

-----------------

We're seeking a Research Engineer to join our ML Research team, owning data quality and evaluation across our modeling efforts. On the data side, you'll build the filtering and curation systems that keep our VLM and classifier training sets clean \- catching label noise, misaligned image\-report pairs, duplicates, and quality issues at scale. On the evaluation side, you'll design how we measure radiology report generation, incorporating clinical accuracy, completeness, hallucination, and reporting style at once. The signals you build feed directly into how our foundation\-model and post\-training teams train and improve their models. This is a broad, dynamic role for someone with the agency to identify where the research team is bottlenecked and address it directly.

Key Responsibilities

------------------------

  • Build data filtering and curation pipelines that keep VLM and classifier training sets clean, detecting label noise, misaligned image\-report pairs, duplicates, corrupted studies, and low\-quality samples at scale.
  • Develop model\-based data quality signals (alignment scoring, automated flagging, active\-learning loops) to surface the ambiguous or high\-value cases worth human review.
  • Partner with radiologists and annotators to define quality criteria, adjudicate edge cases, and turn clinical judgment into reusable, scalable filters.
  • Design evaluation methodology for report generation that goes beyond surface\-level text overlap, measuring clinical accuracy through entity and relation extraction, hallucination and omission rates, and adherence to reporting style.
  • Build and maintain clinical benchmark sets, stratified by modality, pathology, and difficulty, with rigorous attention to train/eval contamination.
  • Develop and validate model\-based evaluators (LLM\-as\-judge, rubric grading) against radiologist judgment, and track how offline eval correlates with production and clinical outcomes.
  • Build continuous evaluation and regression testing so the team can measure every model change quickly and trust the result.
  • Work across the research stack (data, training, and evaluation) finding bottlenecks and shipping the tooling that lets research scientists move faster.

Qualifications

------------------

  • 2\+ years of industry or research experience in ML, data engineering, or a related area
  • Strong Python and solid software engineering fundamentals; comfortable building tooling and data pipelines from scratch
  • Strength in one or both of our core areas, with the willingness to grow into the other:

+ Data quality: dataset curation, filtering, deduplication, label\-noise detection, or data\-centric ML

+ Evaluation: designing metrics or eval harnesses for generative models, LLM\-as\-judge, or NLG / factuality evaluation

  • Demonstrated agency, i.e. a habit of identifying important problems and driving them to a result without waiting to be told
  • Comfort working in an ambiguous, fast\-moving research environment and collaborating closely with research scientists

Preferred Qualifications

----------------------------

  • Experience with medical imaging or clinical data (DICOM, radiology reports, clinical NLP)
  • Familiarity with vision\-language models or multimodal training
  • Experience building human\-in\-the\-loop annotation or review workflows, and reasoning about inter\-annotator agreement
  • Experience with clinical accuracy metrics for report generation (e.g., entity / relation extraction, RadGraph\-style scoring)
  • Experience with data pipeline and experiment tooling (Spark, Airflow, Databricks, or similar)
  • Publications or open\-source contributions in data\-centric ML, evaluation, or medical AI

Role Details

Company Epsilon Labs
Title Research Engineer - Data Quality & Evals
Location San Francisco, CA, US
Experience Mid Level
Salary Not disclosed
Remote No

About This Role

Research Engineers bridge the gap between research and production. They implement papers, build experiment infrastructure, optimize training pipelines, and make research prototypes production-ready. They're the engineers who make research work at scale.

The role sits at a unique intersection. You need to understand the math well enough to implement novel architectures correctly, and you need the engineering chops to make them run efficiently on distributed systems. When a research scientist has a breakthrough idea, you're the person who turns it from a notebook prototype into a training pipeline that runs on 256 GPUs.

Across the 4,317 AI roles we're tracking, Research Engineer positions make up 2% of the market. At Epsilon Labs, this role fits into their broader AI and engineering organization.

Research Engineer roles are growing as AI labs recognize that research velocity depends on engineering quality. The role is less competitive than Research Scientist (no PhD required), but the bar for engineering skill is very high. These roles are concentrated at major labs and well-funded startups.

What the Work Looks Like

A typical week involves: implementing a new attention mechanism from a recent paper, profiling and optimizing a training pipeline that's bottlenecked on data loading, building evaluation infrastructure for a new benchmark, debugging distributed training issues across a GPU cluster, and pair-programming with a research scientist on their latest experiment. The work is deeply technical.

Research Engineer roles are growing as AI labs recognize that research velocity depends on engineering quality. The role is less competitive than Research Scientist (no PhD required), but the bar for engineering skill is very high. These roles are concentrated at major labs and well-funded startups.

Skills Required

Python (52% of roles)

Strong software engineering fundamentals plus ML knowledge. Python, C++, and CUDA experience are common requirements. You'll need to read papers and turn ideas into working code. Distributed systems experience (especially distributed training) is highly valued. Performance optimization skills separate great candidates from good ones.

Experience with large-scale training infrastructure (FSDP, DeepSpeed, Megatron), GPU programming (CUDA, Triton), and the internals of ML frameworks (PyTorch internals, custom autograd functions) is what makes candidates stand out. The best research engineers can debug issues that span the full stack from GPU memory management to numerical precision to algorithmic correctness.

Strong postings mention the team's recent research, the infrastructure scale, and the specific technical challenges. They often list the research areas you'd support. Look for roles that emphasize both implementation quality and research understanding.

Compensation Benchmarks

Research Engineer roles pay a median of $272,100 based on 227 positions with disclosed compensation. Mid-level AI roles across all categories have a median of $194,400.

Across all AI roles, the market median is $215,000. Top-quartile compensation starts at $266,300. The 90th percentile reaches $320,790. For comparison, the highest-paying categories include AI Safety ($287,500) and AI Engineering Manager ($244,000). By seniority level: Entry: $110,000; Mid: $194,400; Senior: $227,400; Director: $274,554; VP: $241,000.

Epsilon Labs AI Hiring

Epsilon Labs has 1 open AI role right now. They're hiring across Research Engineer. Based in San Francisco, CA, US.

Location Context

AI roles in San Francisco pay a median of $265,000 across 1,335 tracked positions. That's 23% above the national median.

Career Path

Common paths into Research Engineer roles include Software Engineer, ML Engineer, Research Intern.

From here, career progression typically leads toward Senior Research Engineer, Research Scientist, ML Architect.

This is one of the best entry points into AI research without a PhD. Build a strong engineering portfolio with ML projects, contribute to open-source ML frameworks, and demonstrate that you can implement complex ideas correctly and efficiently. The transition to Research Scientist is possible with published first-author work, which some research engineer roles support.

What to Expect in Interviews

Technical screens test both engineering skill and research understanding. Expect coding rounds with performance-critical implementations (GPU optimization, efficient data loading). Be prepared to discuss papers relevant to the team's research area and explain how you'd implement key ideas. System design questions focus on training infrastructure: distributed training, experiment tracking, and compute resource management.

When evaluating opportunities: Strong postings mention the team's recent research, the infrastructure scale, and the specific technical challenges. They often list the research areas you'd support. Look for roles that emphasize both implementation quality and research understanding.

AI Hiring Overview

The AI job market has 4,317 open positions tracked in our dataset. By seniority: 138 entry-level, 2,071 mid-level, 1,655 senior, and 453 leadership roles (Director, VP, C-Level). Remote roles make up 15% of the market (635 positions). The remaining 3,657 roles require on-site or hybrid attendance.

The market median for AI roles is $215,000. Top-quartile compensation starts at $266,300. The 90th percentile reaches $320,790. Highest-paying categories: AI Safety ($287,500 median, 34 roles); Research Engineer ($272,100 median, 227 roles); AI Engineering Manager ($244,000 median, 23 roles).

Research Engineer roles are growing as AI labs recognize that research velocity depends on engineering quality. The role is less competitive than Research Scientist (no PhD required), but the bar for engineering skill is very high. These roles are concentrated at major labs and well-funded startups.

The AI Job Market Today

The AI job market spans 4,317 open positions across 15 role categories. The largest categories by volume: AI/ML Engineer (3,004), Data Scientist (345), AI Software Engineer (309). These three account for the majority of open positions, though smaller categories often have higher per-role compensation because of specialized skill requirements.

The seniority mix tells a story about where AI teams are in their maturity. Entry-level roles (138) are outnumbered by mid-level (2,071) and senior (1,655) positions, reflecting that most companies are past the 'build a team from scratch' phase and need experienced engineers who can ship production systems. Leadership roles (Director, VP, C-Level) total 453 positions, representing the bottleneck between technical execution and organizational strategy.

Remote work availability sits at 15% of all AI roles (635 positions), with 3,657 requiring on-site or hybrid attendance. The remote share has stabilized after the post-pandemic correction. Senior and specialized roles (Research Scientist, ML Architect) are more likely to be remote-eligible than entry-level positions, partly because experienced hires have more negotiating power and partly because these roles require less hands-on mentorship.

AI compensation is structured in clear tiers. The market median sits at $215,000. Top-quartile roles start at $266,300, and the 90th percentile reaches $320,790. These figures include base salary with disclosed compensation. Total compensation (including equity, bonuses, and sign-on) runs 20-40% higher at companies that offer those components.

Category matters for compensation. AI Safety roles lead at $287,500 median, while Prompt Engineer roles sit at $145,000. The spread between highest and lowest-paying categories reflects the premium on specialized technical skills versus broader analytical roles.

The most in-demand skills across all AI postings: Python (2,249 postings), Aws (1,224 postings), Azure (938 postings), Rag (915 postings), Gcp (660 postings), Pytorch (640 postings), Prompt Engineering (624 postings), Kubernetes (559 postings). Python dominates, appearing in the vast majority of role descriptions regardless of category. Cloud platform experience (AWS, GCP, Azure) is the second most common requirement. The newer entrants to the top skills list (RAG, vector databases, LLM APIs) reflect the shift from traditional ML toward generative AI applications.

AI Hiring Overview

The AI job market has 4,317 open positions tracked in our dataset. By seniority: 138 entry-level, 2,071 mid-level, 1,655 senior, and 453 leadership roles (Director, VP, C-Level). Remote roles make up 15% of the market (635 positions). The remaining 3,657 roles require on-site or hybrid attendance.

The market median for AI roles is $215,000. Top-quartile compensation starts at $266,300. The 90th percentile reaches $320,790. Highest-paying categories: AI Safety ($287,500 median, 34 roles); Research Engineer ($272,100 median, 227 roles); AI Engineering Manager ($244,000 median, 23 roles).

Frequently Asked Questions

Based on 227 roles with disclosed compensation, the median salary for Research Engineer positions is $272,100. Actual compensation varies by seniority, location, and company stage.
Strong software engineering fundamentals plus ML knowledge. Python, C++, and CUDA experience are common requirements. You'll need to read papers and turn ideas into working code. Distributed systems experience (especially distributed training) is highly valued. Performance optimization skills separate great candidates from good ones.
About 15% of the 4,317 AI roles we track offer remote work. Remote availability varies by company and seniority level, with senior and leadership roles more likely to offer location flexibility.
Epsilon Labs is among the companies actively hiring for AI and ML talent. Check our company profiles for detailed breakdowns of open roles, salary ranges, and hiring trends.
Common next steps from Research Engineer positions include Senior Research Engineer, Research Scientist, ML Architect. Progression depends on whether you lean toward technical depth, people management, or product strategy.

Get Weekly AI Career Intelligence

Salary data, skills demand, and market signals from 16,000+ AI job postings. Every Monday.