Interested in this AI/ML Engineer role at Analysis Group?
Apply Now →Skills & Technologies
About This Role
Locations US\-MA\-Boston
Category
Information Technology
Overview
------------
Analysis Group is one of the largest international economics consulting firms, with more than 1,500 professionals across 15 offices in North America, Europe, and Asia. Since 1981, we have provided expertise in economics, finance, health care analytics, and strategy to top law firms, Fortune Global 500 companies, and government agencies worldwide. Our internal experts, together with our network of affiliated experts from academia, industry, and government, offer our clients exceptional breadth and depth of expertise.
The AI HPC Infrastructure Engineer owns the operation, performance, and growth of a hybrid high\-performance computing (HPC) and AI/GPU infrastructure environment. The engineer maintains the Linux\-based clustered computing platform that supports both traditional HPC/analytical workloads and large\-scale AI/ML training and inference, ensuring systems run efficiently, GPUs and other accelerators are current and well\-utilized, and operations are monitored, documented, and reported — including change management and performance statistics — across both domains. Essential Job Functions and Responsibilities* Maintain, tune, and manage the analytical and AI computing environment for researchers and data scientists, including Posit Workbench (RStudio Server Pro) environments.
- Optimize systems and infrastructure performance using parallelization technologies (MPI, OpenMP) and distributed/multi\-GPU training strategies (e.g., PyTorch Distributed, Horovod, DeepSpeed).
- Design, deploy, and maintain GPU\-accelerated compute infrastructure for large\-scale model training and inference.
- Manage GPU scheduling, multi\-tenancy, and utilization across SLURM and/or Kubernetes\-based environments.
- Administer the NVIDIA software stack — drivers, CUDA, cuDNN, NCCL — and coordinate firmware and health monitoring across GPU fleets.
- Tune and optimize LLM training and inference performance — including batching, quantization, KV\-cache utilization, parallelism strategies, and throughput/latency across GPU clusters.
- Build and maintain MLOps pipelines for model training, versioning, deployment, and monitoring (e.g., MLflow, Kubeflow).
- Manage container orchestration and runtimes (Docker, Kubernetes, Singularity/Apptainer) supporting both HPC jobs and ML workloads.
- Manage access authentication including PAM, LDAP integration, and single sign\-on.
- Design and develop scripts for system administration, automating tasks, monitoring, and usage reporting across HPC and AI resources.
- Manage high\-performance storage and data pipelines for AI training datasets and HPC workloads, primarily on GPFS (IBM Spectrum Scale).
- Troubleshoot, isolate, and resolve application, systems, and other technical problems (hardware, software, network, and GPU\-specific issues).
- Develop and implement backup and recovery programs.
- Research, deploy, and manage general infrastructure, including development of policies and procedures for both HPC and AI/ML environments.
- Migrate data from heterogeneous environments to Linux, on\-prem clusters, or cloud.
- Collaborate with data scientists and ML engineers to support the model development lifecycle and translate research needs into infrastructure requirements.
- Evaluate emerging AI hardware, accelerators, and cloud AI services, and recommend adoption where beneficial.
- Monitor performance, troubleshoot problem areas, and provide statistics and reports across compute, storage, and network.
- Create and maintain documentation related to system configuration, processes, change management, inventory, and service records.
- Ensure continuous network connectivity of all equipment.
- Conduct research and report on products, services, protocols, and standards to remain abreast of developments in HPC and AI infrastructure.
- Participate in a 24x7 on\-call rotation; troubleshoot and resolve issues remotely or onsite as necessary.
Qualifications* Bachelor's degree required; degree in computer science, electrical engineering, or a related field preferred.
- A minimum of 5 years of experience as a hands\-on Linux Systems Administrator in a research, HPC, or production setting.
- An ideal candidate will have 5 to 10 years of substantive relevant experience.
- Experience managing Posit Workbench (RStudio Server Pro), Python, and R environments; strong Posit Workbench administration experience is a significant plus.
- Experience with SLURM, Platform LSF, or other job schedulers required; experience scheduling GPU resources strongly preferred.
- Hands\-on experience with NVIDIA GPU infrastructure and software stack (CUDA, cuDNN, NCCL, NVIDIA GPU Operator) strongly preferred.
- Experience with Kubernetes and container orchestration for AI/ML workloads highly desired.
- Familiarity with ML/AI frameworks (PyTorch, TensorFlow) and distributed training patterns highly desired.
- Experience with MLOps tooling (MLflow, Kubeflow, Weights \& Biases, or similar) is a plus.
- Experience with Bright Cluster Manager is highly desired.
- Experience with Ansible is highly desired.
- Experience with containerization (Docker, Singularity/Apptainer) is highly desired.
- Proficiency with remote access technologies and tools such as RDP, SSH, and emulation software
- Hands\-on experience with GPFS (IBM Spectrum Scale) required.
- Demonstrated experience tuning LLM training and/or inference performance (e.g., batching, quantization, KV\-cache management, parallelism strategies) required.
- Experience with AI Gateways (e.g., LiteLLM, Kong AI Gateway, Portkey, or similar) is a very nice to have.
- Excellent hardware troubleshooting experience, including GPU\-specific diagnostics.
- Knowledge of applicable data privacy practices and laws.
- Strong interpersonal, written, and oral communication skills.
- Highly self\-motivated and directed, with keen attention to detail.
- Proven analytical and problem\-solving abilities.
- Strong customer service orientation.
- Experience working in a collaborative environment.
- An inclusive and growth\-oriented mindset, strong interpersonal skills, and an ability to work across functions.
- To the extent permitted by applicable law, eligible candidates must be authorized to work in the United States, without sponsorship or restriction, now and in the future.
Analysis Group embraces equal opportunity. We are committed to building teams that bring a variety of backgrounds, perspectives, and skills, as we believe that a strong and inclusive workforce directly supports our goal of providing the highest\-quality work. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, or any other class protected under applicable federal, state, or local law, and we encourage candidates of all backgrounds to apply.
Analysis Group offers competitive compensation and a comprehensive benefits package. The estimated salary range for this position is $150,000–$170,000\. Compensation offered will be based on a number of factors including work experience, education, and skill level. This role is eligible for a discretionary annual bonus that is determined in large part by individual performance.
\#LI\-Hybrid
Privacy Notice
------------------
For information about Analysis Group’s privacy practices, please refer to the applicable Analysis Group privacy policy.
- Equal Opportunity Employer/Protected Veterans/Individuals with Disabilities.
Salary Context
This $150K-$170K range is below the median for AI/ML Engineer roles in our dataset (median: $175K across 2162 roles with salary data).
View full AI/ML Engineer salary data →Role Details
About This Role
AI/ML Engineers build and deploy machine learning models in production. They work across the full ML lifecycle: data pipelines, model training, evaluation, and serving infrastructure. The role has evolved significantly over the past two years. Where ML Engineers once spent most of their time on model architecture, the job now tilts heavily toward inference optimization, cost management, and integrating LLM capabilities into existing systems. Companies want engineers who can ship production systems, and the experimenter-only role is fading fast.
Day-to-day, you're writing training pipelines, debugging data quality issues, setting up evaluation frameworks, and figuring out why your model performs differently in staging than it did on your dev set. The best ML engineers are obsessive about reproducibility and measurement. They instrument everything. They know that a model is only as good as the data feeding it and the infrastructure serving it.
Across the 4,317 AI roles we're tracking, AI/ML Engineer positions make up 70% of the market. At Analysis Group, this role fits into their broader AI and engineering organization.
Demand for AI/ML Engineers has been strong and consistent. Unlike some AI roles that spike with hype cycles, ML engineering is a foundational need. Every company deploying AI models needs people who can keep them running, and the gap between research prototypes and production systems keeps growing.
What the Work Looks Like
A typical week might include: debugging a data pipeline that's silently dropping 3% of training examples, running A/B tests on a new model version, writing documentation for a feature flag system that lets you roll back model deployments, and reviewing a junior engineer's PR for a new evaluation metric. Meetings tend to be cross-functional since ML touches product, engineering, and data teams.
Demand for AI/ML Engineers has been strong and consistent. Unlike some AI roles that spike with hype cycles, ML engineering is a foundational need. Every company deploying AI models needs people who can keep them running, and the gap between research prototypes and production systems keeps growing.
Skills Required
Python and PyTorch dominate the requirements. Most roles expect experience with cloud platforms (AWS, GCP, or Azure) and familiarity with ML frameworks like TensorFlow or JAX. RAG (Retrieval-Augmented Generation) has become a top-3 skill requirement as companies integrate LLMs into their products. Docker and Kubernetes show up in about a third of postings, reflecting the production focus of the role.
Beyond the core stack, employers increasingly want experience with experiment tracking tools (MLflow, Weights & Biases), feature stores, and vector databases. Fine-tuning experience is valuable but less common than you'd think from reading Twitter. Most production LLM work is RAG and prompt engineering, not fine-tuning. If you have both, you're in a strong position.
Companies that are serious about AI/ML hiring tend to post specific infrastructure details in the job description: the frameworks they use, their model serving stack, their data pipeline tools. Vague postings that just say 'ML experience required' without specifics are often companies that haven't figured out what they need yet.
Compensation Benchmarks
AI/ML Engineer roles pay a median of $214,900 based on 6,420 positions with disclosed compensation. Mid-level AI roles across all categories have a median of $194,400. This role's midpoint ($160K) sits 26% below the category median. Disclosed range: $150K to $170K.
Across all AI roles, the market median is $215,000. Top-quartile compensation starts at $266,300. The 90th percentile reaches $320,790. For comparison, the highest-paying categories include AI Safety ($287,500) and Research Engineer ($272,100). By seniority level: Entry: $110,000; Mid: $194,400; Senior: $227,400; Director: $274,554; VP: $241,000.
Analysis Group AI Hiring
Analysis Group has 1 open AI role right now. They're hiring across AI/ML Engineer. Based in Boston, MA, US. Compensation range: $170K - $170K.
Location Context
AI roles in Boston pay a median of $210,000 across 166 tracked positions.
Career Path
Common paths into AI/ML Engineer roles include Data Scientist, Software Engineer, Research Engineer.
From here, career progression typically leads toward ML Architect, AI Engineering Manager, Principal ML Engineer.
The fastest path into ML engineering is through software engineering with a self-directed ML education. A CS degree helps, but production engineering skills matter more than academic credentials. Build something that works, deploy it, and measure it. That portfolio project is worth more than a Coursera certificate. For career growth, the fork comes around the senior level: go deep on technical complexity (staff/principal track) or move into managing ML teams.
What to Expect in Interviews
Expect system design questions around ML pipelines: how you'd build a training pipeline for a specific use case, handle data drift, or design A/B testing infrastructure for model deployments. Coding rounds typically involve Python, with emphasis on data manipulation (pandas, numpy) and algorithm implementation. Take-home assignments often ask you to build an end-to-end ML pipeline from raw data to deployed model.
When evaluating opportunities: Companies that are serious about AI/ML hiring tend to post specific infrastructure details in the job description: the frameworks they use, their model serving stack, their data pipeline tools. Vague postings that just say 'ML experience required' without specifics are often companies that haven't figured out what they need yet.
AI Hiring Overview
The AI job market has 4,317 open positions tracked in our dataset. By seniority: 138 entry-level, 2,071 mid-level, 1,655 senior, and 453 leadership roles (Director, VP, C-Level). Remote roles make up 15% of the market (635 positions). The remaining 3,657 roles require on-site or hybrid attendance.
The market median for AI roles is $215,000. Top-quartile compensation starts at $266,300. The 90th percentile reaches $320,790. Highest-paying categories: AI Safety ($287,500 median, 34 roles); Research Engineer ($272,100 median, 227 roles); AI Engineering Manager ($244,000 median, 23 roles).
Demand for AI/ML Engineers has been strong and consistent. Unlike some AI roles that spike with hype cycles, ML engineering is a foundational need. Every company deploying AI models needs people who can keep them running, and the gap between research prototypes and production systems keeps growing.
The AI Job Market Today
The AI job market spans 4,317 open positions across 15 role categories. The largest categories by volume: AI/ML Engineer (3,004), Data Scientist (345), AI Software Engineer (309). These three account for the majority of open positions, though smaller categories often have higher per-role compensation because of specialized skill requirements.
The seniority mix tells a story about where AI teams are in their maturity. Entry-level roles (138) are outnumbered by mid-level (2,071) and senior (1,655) positions, reflecting that most companies are past the 'build a team from scratch' phase and need experienced engineers who can ship production systems. Leadership roles (Director, VP, C-Level) total 453 positions, representing the bottleneck between technical execution and organizational strategy.
Remote work availability sits at 15% of all AI roles (635 positions), with 3,657 requiring on-site or hybrid attendance. The remote share has stabilized after the post-pandemic correction. Senior and specialized roles (Research Scientist, ML Architect) are more likely to be remote-eligible than entry-level positions, partly because experienced hires have more negotiating power and partly because these roles require less hands-on mentorship.
AI compensation is structured in clear tiers. The market median sits at $215,000. Top-quartile roles start at $266,300, and the 90th percentile reaches $320,790. These figures include base salary with disclosed compensation. Total compensation (including equity, bonuses, and sign-on) runs 20-40% higher at companies that offer those components.
Category matters for compensation. AI Safety roles lead at $287,500 median, while Prompt Engineer roles sit at $145,000. The spread between highest and lowest-paying categories reflects the premium on specialized technical skills versus broader analytical roles.
The most in-demand skills across all AI postings: Python (2,249 postings), Aws (1,224 postings), Azure (938 postings), Rag (915 postings), Gcp (660 postings), Pytorch (640 postings), Prompt Engineering (624 postings), Kubernetes (559 postings). Python dominates, appearing in the vast majority of role descriptions regardless of category. Cloud platform experience (AWS, GCP, Azure) is the second most common requirement. The newer entrants to the top skills list (RAG, vector databases, LLM APIs) reflect the shift from traditional ML toward generative AI applications.
Frequently Asked Questions
Get Weekly AI Career Intelligence
Salary data, skills demand, and market signals from 16,000+ AI job postings. Every Monday.