Interested in this MLOps Engineer role at Apple?
Apply Now →Skills & Technologies
About This Role
The Apple Services Engineering team (ASE) is one of the most exciting examples of Apple's long\-held passion for combining art and technology. These are the people who power the App Store, Apple TV, Apple Music, Apple Podcasts, and Apple Books \- at extensive scale, meeting high expectations to deliver a huge variety of entertainment in over 35 languages to more than 150 countries.
Within ASE, the Apple Data Platform SRE team keeps a massive, multi\-cloud platform running for thousands of internal engineers building the next generation of data and AI products at Apple. We sit at the intersection of infrastructure, automation, and customer success \- running incident response, providing hands\-on support to internal teams, and partnering with developers to make cutting\-edge services like Spark, Flink, Airflow, Ray, Notebooks and LLM\-based agent platforms reliable at scale.
Description
This is a rare opportunity to build deep expertise across one of the most technically diverse platforms at Apple \- while specialising in an area that's shaping the future of how Apple builds and operates AI. As an SRE on Apple Data Platform, you'll operate and support the team's full portfolio, from big data pipelines to multi\-cloud infrastructure, and grow into the team's go\-to expert for ML/AI platform services \- including Ray training and serving, LangGraph agent deployments, RAG architectures, embeddings platforms, and vector store platforms. You won't be building the models yourself, but you'll be the infrastructure backbone behind the teams who do \- keeping their services, pipelines, and platforms running flawlessly in production so they can focus on innovation.
We're looking for a self\-motivated engineer who thrives on ownership \- someone who wants a set of services to call their own, the autonomy to drive their reliability roadmap, and the collaborative instinct to keep that work aligned with the team's broader direction. If you love solving hard operational problems, enjoy being the trusted expert customers turn to, and want a front\-row seat to Apple's ML/AI infrastructure evolution, this role offers real room to grow your scope and impact over time.","responsibilities":"Operate, monitor, and triage production and non\-production environments across the ADP portfolio \- data processing, ML/AI, and multi\-cloud infrastructure.
Participate in a rotating on\-call schedule across supported services, including occasional weekday and weekend coverage.
Own the operational health of ML/AI platform services as SME \- driving reliability, support, and customer guidance for Ray, LangGraph agents, RAG pipelines, embeddings, and vector store platforms.
Provide Slack\-based support to internal customers; screen, triage, and resolve service related issues.
Partner with dev teams across time zones to onboard new services \- understanding architecture, then designing monitoring, alerting, and dashboards (Prometheus, Grafana, Splunk).
Build automation and self\-healing tooling that reduces manual toil and scales the team's operational capacity.
Identify, escalate, and resolve production issues to protect platform reliability and customer experience.
Collaborate with SRE and dev partner teams, engineering, and program management to align execution with team and org goals.
Preferred Qualifications
Hands\-on experience operating or supporting Ray (training/serving), LangGraph or similar agent orchestration frameworks, RAG architectures, embeddings platforms, or vector store platforms.
Familiarity with MCP\-based tooling and ML lifecycle/dataset management systems.
Experience with S3 and cloud storage/networking fundamentals.
Familiarity with observability tooling: Prometheus, Grafana, Splunk, PagerDuty.
Working knowledge of CI/CD pipelines and deployment workflows.
Deep understanding of one or more Big Data technologies (Spark, Flink, Airflow, Trino, Notebooks).
A track record of automating manual operations through scripting or tooling.
Intellectual curiosity and a drive to keep learning \- for yourself, your team, and the org.
Minimum Qualifications
Minimum Qualifications
- Bachelor's Degree in Computer Science, an engineering\-related field, or equivalent related experience.
1\-4 years in a Site Reliability Engineering, DevOps, or Infrastructure\-focused role.
Proficient in Python; working knowledge of Golang a plus.
Experience with Kubernetes and at least one major cloud provider (AWS or GCP).
Exposure to operating or supporting ML pipelines, model\-serving infrastructure, or LLM\-based systems in production.
Strong communication skills and composure under pressure during incidents.
Solid grounding in SRE principles, with prior on\-call or production\-support experience.
Role Details
About This Role
MLOps Engineers build the infrastructure that keeps ML models running in production. They own CI/CD pipelines for model deployment, monitoring for data drift and model degradation, and the tooling that lets data scientists ship faster. If ML Engineers build the models, MLOps Engineers build the roads those models travel on.
The job is fundamentally about reliability and velocity. Data scientists want to iterate fast. Product teams want stable predictions. Your job is to make both happen simultaneously. That means building deployment pipelines that catch regressions before they hit production, monitoring systems that alert on data drift before it degrades model performance, and self-service tooling that lets data scientists deploy without filing a ticket.
Across the 4,317 AI roles we're tracking, MLOps Engineer positions make up 1% of the market. At Apple, this role fits into their broader AI and engineering organization.
MLOps demand tracks closely with production ML adoption. As more companies move models from notebooks to production, the need for MLOps grows. The role is well-established at large tech companies and growing fast at mid-stage startups that are hitting the 'our models work in notebooks but break in production' phase.
What the Work Looks Like
A typical week involves: debugging a model deployment that's serving stale predictions, building a new monitoring dashboard for a feature team, writing Terraform for GPU-enabled inference clusters, reviewing pull requests for the ML platform's CI/CD pipeline, and meeting with data scientists to understand their pain points. You're the bridge between ML and infrastructure.
MLOps demand tracks closely with production ML adoption. As more companies move models from notebooks to production, the need for MLOps grows. The role is well-established at large tech companies and growing fast at mid-stage startups that are hitting the 'our models work in notebooks but break in production' phase.
Skills Required
Kubernetes, Docker, and cloud infrastructure are baseline. Most roles want experience with ML-specific tooling: MLflow, Kubeflow, Weights & Biases, or similar. Strong DevOps fundamentals matter more than ML theory. You need to understand model serving (TorchServe, Triton, vLLM), monitoring (Prometheus, Grafana), and infrastructure-as-code (Terraform, Pulumi).
GPU infrastructure knowledge is increasingly valuable as LLM inference becomes a major cost center. Understanding GPU scheduling, multi-node training setups, and inference optimization (quantization, batching, caching) puts you in the top tier. Experience with model registries and feature stores rounds out the profile.
Good MLOps postings specify their ML stack, infrastructure scale, and the problems they're solving (deployment velocity, cost optimization, monitoring gaps). Red flag: companies that want MLOps but don't have any models in production yet. You'll end up doing general DevOps instead.
Compensation Benchmarks
MLOps Engineer roles pay a median of $203,000 based on 85 positions with disclosed compensation. Mid-level AI roles across all categories have a median of $194,400.
Across all AI roles, the market median is $215,000. Top-quartile compensation starts at $266,300. The 90th percentile reaches $320,790. For comparison, the highest-paying categories include AI Safety ($287,500) and Research Engineer ($272,100). By seniority level: Entry: $110,000; Mid: $194,400; Senior: $227,400; Director: $274,554; VP: $241,000.
Apple AI Hiring
Apple has 57 open AI roles right now. They're hiring across AI/ML Engineer, AI Software Engineer, Research Scientist, AI Product Manager. Positions span Cupertino, CA, US, Sunnyvale, CA, US, San Diego, CA, US. Compensation range: $214K - $401K.
Location Context
AI roles in Austin pay a median of $214,343 across 143 tracked positions.
Career Path
Common paths into MLOps Engineer roles include DevOps Engineer, Platform Engineer, Data Engineer.
From here, career progression typically leads toward ML Platform Lead, Infrastructure Architect, Engineering Manager.
DevOps engineers with ML curiosity have the shortest path. You already understand deployment, monitoring, and infrastructure. Add ML-specific knowledge (model serving, data pipelines, experiment tracking) and you're competitive. The career ceiling is high: ML Platform Lead roles at top companies pay well because the infrastructure complexity is enormous.
What to Expect in Interviews
Interviews emphasize infrastructure and reliability. Expect questions about CI/CD for ML models, monitoring for data drift, and how you'd design a model serving platform that handles 10K requests per second. Coding rounds focus on Python and infrastructure-as-code (Terraform, Helm). Be ready to discuss tradeoffs between different model serving frameworks and how you'd handle rollback when a new model degrades performance.
When evaluating opportunities: Good MLOps postings specify their ML stack, infrastructure scale, and the problems they're solving (deployment velocity, cost optimization, monitoring gaps). Red flag: companies that want MLOps but don't have any models in production yet. You'll end up doing general DevOps instead.
AI Hiring Overview
The AI job market has 4,317 open positions tracked in our dataset. By seniority: 138 entry-level, 2,071 mid-level, 1,655 senior, and 453 leadership roles (Director, VP, C-Level). Remote roles make up 15% of the market (635 positions). The remaining 3,657 roles require on-site or hybrid attendance.
The market median for AI roles is $215,000. Top-quartile compensation starts at $266,300. The 90th percentile reaches $320,790. Highest-paying categories: AI Safety ($287,500 median, 34 roles); Research Engineer ($272,100 median, 227 roles); AI Engineering Manager ($244,000 median, 23 roles).
MLOps demand tracks closely with production ML adoption. As more companies move models from notebooks to production, the need for MLOps grows. The role is well-established at large tech companies and growing fast at mid-stage startups that are hitting the 'our models work in notebooks but break in production' phase.
The AI Job Market Today
The AI job market spans 4,317 open positions across 15 role categories. The largest categories by volume: AI/ML Engineer (3,004), Data Scientist (345), AI Software Engineer (309). These three account for the majority of open positions, though smaller categories often have higher per-role compensation because of specialized skill requirements.
The seniority mix tells a story about where AI teams are in their maturity. Entry-level roles (138) are outnumbered by mid-level (2,071) and senior (1,655) positions, reflecting that most companies are past the 'build a team from scratch' phase and need experienced engineers who can ship production systems. Leadership roles (Director, VP, C-Level) total 453 positions, representing the bottleneck between technical execution and organizational strategy.
Remote work availability sits at 15% of all AI roles (635 positions), with 3,657 requiring on-site or hybrid attendance. The remote share has stabilized after the post-pandemic correction. Senior and specialized roles (Research Scientist, ML Architect) are more likely to be remote-eligible than entry-level positions, partly because experienced hires have more negotiating power and partly because these roles require less hands-on mentorship.
AI compensation is structured in clear tiers. The market median sits at $215,000. Top-quartile roles start at $266,300, and the 90th percentile reaches $320,790. These figures include base salary with disclosed compensation. Total compensation (including equity, bonuses, and sign-on) runs 20-40% higher at companies that offer those components.
Category matters for compensation. AI Safety roles lead at $287,500 median, while Prompt Engineer roles sit at $145,000. The spread between highest and lowest-paying categories reflects the premium on specialized technical skills versus broader analytical roles.
The most in-demand skills across all AI postings: Python (2,249 postings), Aws (1,224 postings), Azure (938 postings), Rag (915 postings), Gcp (660 postings), Pytorch (640 postings), Prompt Engineering (624 postings), Kubernetes (559 postings). Python dominates, appearing in the vast majority of role descriptions regardless of category. Cloud platform experience (AWS, GCP, Azure) is the second most common requirement. The newer entrants to the top skills list (RAG, vector databases, LLM APIs) reflect the shift from traditional ML toward generative AI applications.
Frequently Asked Questions
Get Weekly AI Career Intelligence
Salary data, skills demand, and market signals from 16,000+ AI job postings. Every Monday.