Senior Research Engineer, LLM Training & Post-Training

$165K - $310K New York, NY, US Senior Research Engineer

Interested in this Research Engineer role at Lightning AI?

Apply Now →

Skills & Technologies

Hugging FacePythonPytorchRlhfTransformers

About This Role

AI job market dashboard showing open roles by category

Who We Are

------------------

Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end\-to\-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with less friction.

Through our merger with Voltage Park, a neocloud and AI Factory, Lightning AI combines developer\-first software with cost\-efficient, large\-scale compute. Teams get the tools they need for experimentation, training, and production inference, with security, observability, and control built in.

We serve solo researchers, startups, and large enterprises. Lightning AI operates globally with offices in New York City, San Francisco, Seattle, and London, and is backed by Coatue, Index Ventures, Bain Capital Ventures, and Firstminute.

The Way We Work

-------------------

The people who thrive here are builders who move fast, communicate openly, take ownership, and continuously improve themselves, their teams, and our company. Here's what that looks like in practice:

  • Move with Urgency: We move quickly, make thoughtful decisions, and keep momentum. We value action over perfection and learn by shipping.
  • Take Ownership: We own outcomes, not just our individual work. We make decisions that move the company forward and follow through.
  • Communicate Openly: We communicate directly, seek to understand, and create clarity for others. Honest conversations help us move faster together.
  • Build Great Teams: We lead by example, empower others, and create healthy teams where people can do their best work.
  • Raise the Bar: We're always improving ourselves. We learn from feedback, consistently challenge ourselves to grow, and focus on the work that matters most.
  • Think Long\-Term: We design for what's next. We create scalable systems, simplify complexity, and use AI and automation to amplify our impact.

What We're Looking For

--------------------------

We're are looking for an experienced Senior Research Engineer who has built, trained, and optimized modern transformer\-based language models to join our Research Engineering function at Lightning.

This role will focus on advancing how large language models are trained, fine\-tuned, evaluated, and deployed across Lightning AI's platform and real\-world customer workloads. It will work across model training, post\-training, PyTorch, distributed systems, and AI systems engineering to improve model quality, training efficiency, and developer productivity while collaborating closely with researchers, infrastructure engineers, and customers.

We're looking for someone who enjoys turning cutting\-edge research into production systems. You have deep experience training and improving transformer\-based language models, strong software engineering fundamentals, and a passion for solving difficult problems across model training, evaluation, and AI systems. Rather than building applications on top of existing models, you're motivated by improving the models themselves and the systems that power them. Our work spans models that power the Lightning AI platform, customer\-specific model workloads, and research that translates into reusable training and platform capabilities.

*This role is hybrid with a minimum of 2 in\-office days per week in San Francisco, Seattle, NYC, or London, with fully remote work considered for candidates outside of our office hub locations. All employees participate in occasional team and company offsites.*

What You'll Do

------------------

  • Design, build, and optimize training and post\-training pipelines for large language models.
  • Improve model quality through supervised fine\-tuning, continued pretraining, preference optimization, reinforcement learning, evaluation, and experimentation.
  • Build and improve PyTorch\-based training infrastructure, tooling, and developer workflows.
  • Optimize distributed training across multi\-GPU environments by improving throughput, memory efficiency, scalability, and GPU utilization.
  • Investigate challenging model training issues, including convergence, instability, communication overhead, and performance bottlenecks.
  • Design evaluation methodologies, benchmark models, analyze failure modes, and guide model improvements through experimentation.
  • Collaborate directly with customers to understand real\-world workloads and translate those learnings into improvements across Lightning AI's research platform.
  • Partner closely with research, infrastructure, and platform engineering teams to build production\-ready AI systems.
  • Contribute to open\-source projects through new features, tooling improvements, documentation, and community engagement

What You'll Need

--------------------

### Required Qualifications

  • Significant experience training, fine\-tuning, evaluating, and optimizing transformer\-based language models using PyTorch.
  • Experience with modern LLM training and post\-training techniques such as continued pretraining, SFT, RLHF, preference optimization (DPO, PPO, GRPO), reward modeling, or similar approaches.
  • Strong understanding of distributed training and multi\-GPU systems, with experience improving training performance, scalability, or efficiency.
  • Strong software engineering fundamentals, including building production\-quality Python software and research tooling.
  • Experience designing experiments, evaluating model performance, and debugging complex training or optimization issues.
  • Excellent communication and collaboration skills, including the ability to work effectively across research, product, infrastructure, and customer\-facing engagements.
  • Comfortable working in fast\-moving, ambiguous environments where priorities evolve over time.
  • Master's degree, PhD, or equivalent industry experience in Machine Learning, AI, Computer Science, or a related field

### Ideal Experience

Experience with one or more of the following:

  • DeepSpeed, FSDP, Megatron\-LM, Hugging Face Transformers, TRL, PEFT, Lightning Fabric, or similar training frameworks.
  • CUDA, Triton, vLLM, SGLang, TensorRT, or other AI systems and performance optimization technologies.
  • GPU performance optimization, mixed precision, memory optimization, or distributed training optimization.
  • Open\-source contributions, research publications, or production AI platforms supporting large\-scale training or inference workloads.
  • Startup experience or experience working on highly cross\-functional engineering teams.

Compensation

----------------

Benefits and Perks

----------------------

We offer a comprehensive and competitive benefits package designed to support our employees' health, well\-being, and long\-term success:

  • Comprehensive Health Coverage: Medical, dental, and vision coverage for employees and eligible dependents.
  • Meaningful Equity: RSUs that give employees a stake in the company's long\-term success.
  • Retirement Savings: 401(k) matching (U.S.) and pension contributions (U.K.).
  • Flexible Time Off: Unlimited PTO, company holidays, and floating holidays to support work\-life balance.
  • Company\-Wide Winter Break: Two weeks of company closure each winter to disconnect and recharge.
  • Paid Parental \& Family Leave: Paid leave to support you and your family through life's important moments.
  • Professional Development: Annual learning and development allowance to support your professional growth.
  • Wellness Benefits: Wellness and work\-from\-home stipends to support your physical and mental well\-being.
  • Sabbatical Program: Four weeks of paid sabbatical leave after four years of service.
  • Flexible Work: Flexible schedules and a hybrid work model for our office\-based teams.
  • In\-Office Meals: Complimentary meals at our office hubs.

Benefits may vary by location, team, and role.

*At Lightning AI, we are committed to fostering an inclusive and diverse workplace. We believe that diverse teams drive innovation and create better products. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected characteristic. We are dedicated to building a culture where everyone can thrive and contribute to their fullest potential.*

Salary Context

This $165K-$310K range is above the median for Research Engineer roles in our dataset (median: $207K across 63 roles with salary data).

View full Research Engineer salary data →

Role Details

Company Lightning AI
Title Senior Research Engineer, LLM Training & Post-Training
Location New York, NY, US
Experience Senior
Salary $165K - $310K
Remote No

About This Role

Research Engineers bridge the gap between research and production. They implement papers, build experiment infrastructure, optimize training pipelines, and make research prototypes production-ready. They're the engineers who make research work at scale.

The role sits at a unique intersection. You need to understand the math well enough to implement novel architectures correctly, and you need the engineering chops to make them run efficiently on distributed systems. When a research scientist has a breakthrough idea, you're the person who turns it from a notebook prototype into a training pipeline that runs on 256 GPUs.

Across the 4,317 AI roles we're tracking, Research Engineer positions make up 2% of the market. At Lightning AI, this role fits into their broader AI and engineering organization.

Research Engineer roles are growing as AI labs recognize that research velocity depends on engineering quality. The role is less competitive than Research Scientist (no PhD required), but the bar for engineering skill is very high. These roles are concentrated at major labs and well-funded startups.

What the Work Looks Like

A typical week involves: implementing a new attention mechanism from a recent paper, profiling and optimizing a training pipeline that's bottlenecked on data loading, building evaluation infrastructure for a new benchmark, debugging distributed training issues across a GPU cluster, and pair-programming with a research scientist on their latest experiment. The work is deeply technical.

Research Engineer roles are growing as AI labs recognize that research velocity depends on engineering quality. The role is less competitive than Research Scientist (no PhD required), but the bar for engineering skill is very high. These roles are concentrated at major labs and well-funded startups.

Skills Required

Hugging Face (3% of roles) Python (52% of roles) Pytorch (15% of roles) Rlhf (1% of roles) Transformers (3% of roles)

Strong software engineering fundamentals plus ML knowledge. Python, C++, and CUDA experience are common requirements. You'll need to read papers and turn ideas into working code. Distributed systems experience (especially distributed training) is highly valued. Performance optimization skills separate great candidates from good ones.

Experience with large-scale training infrastructure (FSDP, DeepSpeed, Megatron), GPU programming (CUDA, Triton), and the internals of ML frameworks (PyTorch internals, custom autograd functions) is what makes candidates stand out. The best research engineers can debug issues that span the full stack from GPU memory management to numerical precision to algorithmic correctness.

Strong postings mention the team's recent research, the infrastructure scale, and the specific technical challenges. They often list the research areas you'd support. Look for roles that emphasize both implementation quality and research understanding.

Compensation Benchmarks

Research Engineer roles pay a median of $272,100 based on 227 positions with disclosed compensation. Senior-level AI roles across all categories have a median of $227,400. This role's midpoint ($237K) sits 13% below the category median. Disclosed range: $165K to $310K.

Across all AI roles, the market median is $215,000. Top-quartile compensation starts at $266,300. The 90th percentile reaches $320,790. For comparison, the highest-paying categories include AI Safety ($287,500) and AI Engineering Manager ($244,000). By seniority level: Entry: $110,000; Mid: $194,400; Senior: $227,400; Director: $274,554; VP: $241,000.

Lightning AI AI Hiring

Lightning AI has 2 open AI roles right now. They're hiring across Research Engineer, AI/ML Engineer. Based in New York, NY, US. Compensation range: $140K - $310K.

Location Context

AI roles in New York pay a median of $220,000 across 1,650 tracked positions.

Career Path

Common paths into Research Engineer roles include Software Engineer, ML Engineer, Research Intern.

From here, career progression typically leads toward Senior Research Engineer, Research Scientist, ML Architect.

This is one of the best entry points into AI research without a PhD. Build a strong engineering portfolio with ML projects, contribute to open-source ML frameworks, and demonstrate that you can implement complex ideas correctly and efficiently. The transition to Research Scientist is possible with published first-author work, which some research engineer roles support.

What to Expect in Interviews

Technical screens test both engineering skill and research understanding. Expect coding rounds with performance-critical implementations (GPU optimization, efficient data loading). Be prepared to discuss papers relevant to the team's research area and explain how you'd implement key ideas. System design questions focus on training infrastructure: distributed training, experiment tracking, and compute resource management.

When evaluating opportunities: Strong postings mention the team's recent research, the infrastructure scale, and the specific technical challenges. They often list the research areas you'd support. Look for roles that emphasize both implementation quality and research understanding.

AI Hiring Overview

The AI job market has 4,317 open positions tracked in our dataset. By seniority: 138 entry-level, 2,071 mid-level, 1,655 senior, and 453 leadership roles (Director, VP, C-Level). Remote roles make up 15% of the market (635 positions). The remaining 3,657 roles require on-site or hybrid attendance.

The market median for AI roles is $215,000. Top-quartile compensation starts at $266,300. The 90th percentile reaches $320,790. Highest-paying categories: AI Safety ($287,500 median, 34 roles); Research Engineer ($272,100 median, 227 roles); AI Engineering Manager ($244,000 median, 23 roles).

Research Engineer roles are growing as AI labs recognize that research velocity depends on engineering quality. The role is less competitive than Research Scientist (no PhD required), but the bar for engineering skill is very high. These roles are concentrated at major labs and well-funded startups.

The AI Job Market Today

The AI job market spans 4,317 open positions across 15 role categories. The largest categories by volume: AI/ML Engineer (3,004), Data Scientist (345), AI Software Engineer (309). These three account for the majority of open positions, though smaller categories often have higher per-role compensation because of specialized skill requirements.

The seniority mix tells a story about where AI teams are in their maturity. Entry-level roles (138) are outnumbered by mid-level (2,071) and senior (1,655) positions, reflecting that most companies are past the 'build a team from scratch' phase and need experienced engineers who can ship production systems. Leadership roles (Director, VP, C-Level) total 453 positions, representing the bottleneck between technical execution and organizational strategy.

Remote work availability sits at 15% of all AI roles (635 positions), with 3,657 requiring on-site or hybrid attendance. The remote share has stabilized after the post-pandemic correction. Senior and specialized roles (Research Scientist, ML Architect) are more likely to be remote-eligible than entry-level positions, partly because experienced hires have more negotiating power and partly because these roles require less hands-on mentorship.

AI compensation is structured in clear tiers. The market median sits at $215,000. Top-quartile roles start at $266,300, and the 90th percentile reaches $320,790. These figures include base salary with disclosed compensation. Total compensation (including equity, bonuses, and sign-on) runs 20-40% higher at companies that offer those components.

Category matters for compensation. AI Safety roles lead at $287,500 median, while Prompt Engineer roles sit at $145,000. The spread between highest and lowest-paying categories reflects the premium on specialized technical skills versus broader analytical roles.

The most in-demand skills across all AI postings: Python (2,249 postings), Aws (1,224 postings), Azure (938 postings), Rag (915 postings), Gcp (660 postings), Pytorch (640 postings), Prompt Engineering (624 postings), Kubernetes (559 postings). Python dominates, appearing in the vast majority of role descriptions regardless of category. Cloud platform experience (AWS, GCP, Azure) is the second most common requirement. The newer entrants to the top skills list (RAG, vector databases, LLM APIs) reflect the shift from traditional ML toward generative AI applications.

Frequently Asked Questions

Based on 227 roles with disclosed compensation, the median salary for Research Engineer positions is $272,100. Actual compensation varies by seniority, location, and company stage.
Strong software engineering fundamentals plus ML knowledge. Python, C++, and CUDA experience are common requirements. You'll need to read papers and turn ideas into working code. Distributed systems experience (especially distributed training) is highly valued. Performance optimization skills separate great candidates from good ones.
About 15% of the 4,317 AI roles we track offer remote work. Remote availability varies by company and seniority level, with senior and leadership roles more likely to offer location flexibility.
Lightning AI is among the companies actively hiring for AI and ML talent. Check our company profiles for detailed breakdowns of open roles, salary ranges, and hiring trends.
Common next steps from Research Engineer positions include Senior Research Engineer, Research Scientist, ML Architect. Progression depends on whether you lean toward technical depth, people management, or product strategy.

Get Weekly AI Career Intelligence

Salary data, skills demand, and market signals from 16,000+ AI job postings. Every Monday.