Interested in this AI Engineering Manager role at Mastercard?
Apply Now →Skills & Technologies
About This Role
Our Purpose
*Mastercard powers economies and empowers people in 200\+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential.*
Title and Summary
Manager, AI Engineering (Tester )
Mastercard's Business \& Market Insights (B\&MI) group delivers unparalleled data\-driven intelligence and frontier AI solutions that help organizations make smarter, faster, and more impactful decisions. We are currently looking for a AI Tester for the Operational Intelligence Program within B\&MI. This is a highly specialized, hands\-on AI testing leadership position dedicated to ensuring our Generative AI, LLM, and agentic systems are accurate, safe, reliable, and enterprise\-ready. This role will lead AI quality engineering efforts — defining evaluation frameworks, red\-teaming strategies, and LLMOps quality gates — while fostering a culture of rigorous, first\-class AI testing across the program.
Roles and Responsibilities:
- Design and own end\-to\-end LLM evaluation frameworks — including automated prompt regression pipelines, output scoring, semantic benchmarking, and hallucination detection across model versions and prompt variations.
- Build comprehensive test suites for agentic AI systems — validating tool selection, inter\-agent coordination, task decomposition, goal completion, and failure handling across multi\-step reasoning workflows.
- Develop RAG pipeline evaluation frameworks assessing retrieval precision, chunk relevance, context faithfulness, answer grounding, and hallucination rates using tools like RAGAS, TruLens, and DeepEval.
- Lead structured red\-teaming and adversarial testing exercises targeting prompt injection, jailbreaks, data leakage, context poisoning, and model manipulation — building and maintaining an evolving adversarial test library.
- Execute fairness, bias, and Responsible AI audits — testing for demographic bias, sentiment skew, representation gaps, and validating explainability mechanisms, citations, and confidence score accuracy.
- Design and run inference performance benchmarks — measuring latency, throughput, token efficiency, and degradation under peak load — and enforce LLM quality gates within CI/CD pipelines on Databricks (AWS).
- Build production monitoring and drift detection pipelines tracking semantic output drift, embedding shifts, retrieval degradation, and anomalous agent behaviors using observability tooling (Grafana, Datadog, CloudWatch).
- Define the AI testing roadmap and quality standards for the program — establishing evaluation metrics, tooling choices, and documentation practices across all Gen AI workstreams.
- Partner with Gen AI engineers, ML engineers, and product stakeholders to embed quality from day one — reviewing prompt architectures, agent designs, and system workflows for testability and risk.
- Continuously research and adopt frontier evaluation benchmarks (RAGAS, MMLU, TruthfulQA, MT\-Bench) and emerging AI testing methodologies to keep quality practices at the cutting edge.
All About You:
- Master's/Bachelor's degree in Computer Science, AI/ML, or Software Engineering, with considerable hands\-on experience leading AI/ML quality engineering or LLM testing programs in production environments.
- Demonstrated expertise testing LLM and Gen AI systems — including prompt testing, output evaluation, hallucination detection, RAG pipeline assessment, and agentic workflow validation in real production settings.
- Deep hands\-on knowledge of AI evaluation frameworks and tooling: RAGAS, DeepEval, TruLens, LangSmith, PromptFlow, Weights \& Biases Evals, or equivalent platforms.
- Strong understanding of Gen AI failure modes — hallucination, prompt injection, retrieval grounding failures, context drift, agent loop failures — and proven methods to surface and document them systematically.
- Strong Python programming skills with the ability to independently build test automation scripts, evaluation pipelines, and API\-level integration tests; SQL proficiency required.
- Working knowledge of LLM ecosystems — OpenAI, Anthropic, Hugging Face, LangChain/LangGraph — sufficient to understand model behavior, prompt structure, and agent architecture deeply enough to test them rigorously.
- Familiarity with MLOps/LLMOps pipelines (MLflow, Databricks, SageMaker) and experience integrating automated quality gates into CI/CD workflows for AI systems.
- Experience with cloud AI infrastructure (AWS, Azure, or GCP) and observability tooling for monitoring live AI system behavior and output quality in production.
- Strong analytical, communication, and stakeholder management skills — with the ability to translate complex AI failure patterns into clear risk assessments and remediation recommendations for both technical and business audiences.
Mastercard is a merit\-based, inclusive, equal opportunity employer that considers applicants without regard to gender, gender identity, sexual orientation, race, ethnicity, disabled or veteran status, or any other characteristic protected by law. We hire the most qualified candidate for the role. In the US or Canada, if you require accommodations or assistance to complete the online application process or during the recruitment process, please contact reasonable\[email protected] and identify the type of accommodation or assistance you are requesting. Do not include any medical or health information in this email. The Reasonable Accommodations team will respond to your email promptly.Corporate Security Responsibility
All activities involving access to Mastercard assets, information, and networks comes with an inherent risk to the organization and, therefore, it is expected that every person working for, or on behalf of, Mastercard is responsible for information security and must:
- Abide by Mastercard’s security policies and practices;
- Ensure the confidentiality and integrity of the information being accessed;
- Report any suspected information security violation or breach, and
- Complete all periodic mandatory security trainings in accordance with Mastercard’s guidelines.
In line with Mastercard’s total compensation philosophy and assuming that the job will be performed in the US, the successful candidate will be offered a competitive base salary and may be eligible for an annual bonus or commissions depending on the role. The base salary offered may vary depending on multiple factors, including but not limited to location, job\-related knowledge, skills, and experience. Mastercard benefits for full time (and certain part time) employees generally include: insurance (including medical, prescription drug, dental, vision, disability, life insurance); flexible spending account and health savings account; paid leaves (including 16 weeks of new parent leave and up to 20 days of bereavement leave); 80 hours of Paid Sick and Safe Time, 25 days of vacation time and 5 personal days, pro\-rated based on date of hire; 10 annual paid U.S. observed holidays; 401k with a best\-in\-class company match; deferred compensation for eligible roles; fitness reimbursement or on\-site fitness facilities; eligibility for tuition reimbursement; and many more. Mastercard benefits for interns generally include: 56 hours of Paid Sick and Safe Time; jury duty leave; and on\-site fitness facilities in some locations.Pay Ranges
O'Fallon, Missouri: $140,000 \- $231,000 USD
Salary Context
This $140K-$231K range is above the median for AI Engineering Manager roles in our dataset (median: $185K across 13 roles with salary data).
Role Details
About This Role
This role sits at the intersection of AI and engineering, building systems that bring machine learning capabilities into production environments. The scope varies by company, but the common thread is applying AI technology to solve real business problems at scale. Most AI roles today require a combination of software engineering fundamentals and domain-specific ML knowledge, with the exact mix depending on the team's maturity and the product they're building.
The AI job market is evolving fast. New role categories emerge as companies figure out what they need to ship AI-powered products. What matters most is the ability to learn quickly, build working systems, and iterate based on real-world performance data. The specific title matters less than the skills you bring and the problems you can solve. Companies are past the experimentation phase and want engineers who can deliver production-quality systems that work reliably at scale.
Across the 4,317 AI roles we're tracking, AI Engineering Manager positions make up 0% of the market. At Mastercard, this role fits into their broader AI and engineering organization.
AI hiring keeps growing across industries. Companies in tech, finance, healthcare, and retail are all building AI teams. The strongest demand is for people who can bridge the gap between AI research and production engineering. The shift toward generative AI has created new role types (LLM Engineer, Prompt Engineer, AI Agent Developer) that didn't exist three years ago, while traditional roles (Data Scientist, ML Engineer) have evolved to incorporate LLM capabilities.
What the Work Looks Like
Day-to-day work involves a mix of building, debugging, and collaborating. You'll write code, review pull requests, participate in design discussions, and work with cross-functional teams (product, design, data) to define what AI features should do and how they should behave. Expect to spend time on both technical implementation and communication. Most AI teams operate in two-week sprint cycles, with regular demos and retrospectives. The ratio of heads-down coding to meetings and reviews varies by seniority, with senior roles spending more time on architecture decisions and mentorship.
AI hiring keeps growing across industries. Companies in tech, finance, healthcare, and retail are all building AI teams. The strongest demand is for people who can bridge the gap between AI research and production engineering. The shift toward generative AI has created new role types (LLM Engineer, Prompt Engineer, AI Agent Developer) that didn't exist three years ago, while traditional roles (Data Scientist, ML Engineer) have evolved to incorporate LLM capabilities.
Skills Required
Python and cloud platform experience are common requirements. Specific skill needs vary by company and focus area, but familiarity with ML frameworks, data pipelines, and API design covers the basics for most roles. RAG (Retrieval-Augmented Generation), vector databases, and LLM API integration are increasingly standard requirements across role types.
Beyond the core stack, communication skills matter more than many technical candidates realize. The ability to explain AI capabilities and limitations to non-technical stakeholders is a differentiator at every level. Technical writing, documentation, and clear thinking about tradeoffs are underrated skills in AI roles. Experience with evaluation methodology (how to measure whether an AI system is working well) is becoming a core requirement, especially for roles that involve LLM integration.
Look for job postings that specify the problems you'll work on, the tech stack, and the team structure. Vague postings that list every AI buzzword are often a sign the company hasn't figured out what they need. Strong postings describe the product context, the team you'd join, and the specific challenges you'd tackle.
Compensation Benchmarks
AI Engineering Manager roles pay a median of $244,000 based on 23 positions with disclosed compensation. Mid-level AI roles across all categories have a median of $194,400. This role's midpoint ($185K) sits 24% below the category median. Disclosed range: $140K to $231K.
Across all AI roles, the market median is $215,000. Top-quartile compensation starts at $266,300. The 90th percentile reaches $320,790. For comparison, the highest-paying categories include AI Safety ($287,500) and Research Engineer ($272,100). By seniority level: Entry: $110,000; Mid: $194,400; Senior: $227,400; Director: $274,554; VP: $241,000.
Mastercard AI Hiring
Mastercard has 6 open AI roles right now. They're hiring across Data Scientist, AI/ML Engineer, AI Engineering Manager. Positions span Salt Lake City, UT, US, New York, NY, US, O'Fallon, MO, US. Compensation range: $115K - $391K.
Location Context
Across all AI roles, 15% (635 positions) offer remote work, while 3,657 require on-site attendance. Top AI hiring metros: New York (1,650 roles, $220,000 median); San Francisco (1,335 roles, $265,000 median); Los Angeles (708 roles, $214,112 median).
Career Path
Common paths into AI Engineering Manager roles include Software Engineer, Data Scientist, Data Analyst.
From here, career progression typically leads toward Senior Engineer, AI Architect, Engineering Manager, Principal Engineer.
Focus on building things that work. A deployed project that solves a real problem is worth more than any certification. Contribute to open-source, build portfolio projects, and invest in fundamentals (software engineering, statistics, systems design) rather than chasing the latest framework. The AI field moves fast, but the engineers who succeed long-term are the ones with strong fundamentals who can adapt to new tools and paradigms as they emerge.
What to Expect in Interviews
AI interviews typically combine coding challenges (Python-focused), system design questions tailored to the role, and discussions about your experience with relevant tools and frameworks. Strong candidates demonstrate both technical depth and the ability to make pragmatic engineering tradeoffs. Prepare portfolio projects that demonstrate end-to-end capability rather than isolated skills.
When evaluating opportunities: Look for job postings that specify the problems you'll work on, the tech stack, and the team structure. Vague postings that list every AI buzzword are often a sign the company hasn't figured out what they need. Strong postings describe the product context, the team you'd join, and the specific challenges you'd tackle.
AI Hiring Overview
The AI job market has 4,317 open positions tracked in our dataset. By seniority: 138 entry-level, 2,071 mid-level, 1,655 senior, and 453 leadership roles (Director, VP, C-Level). Remote roles make up 15% of the market (635 positions). The remaining 3,657 roles require on-site or hybrid attendance.
The market median for AI roles is $215,000. Top-quartile compensation starts at $266,300. The 90th percentile reaches $320,790. Highest-paying categories: AI Safety ($287,500 median, 34 roles); Research Engineer ($272,100 median, 227 roles); AI Engineering Manager ($244,000 median, 23 roles).
AI hiring keeps growing across industries. Companies in tech, finance, healthcare, and retail are all building AI teams. The strongest demand is for people who can bridge the gap between AI research and production engineering. The shift toward generative AI has created new role types (LLM Engineer, Prompt Engineer, AI Agent Developer) that didn't exist three years ago, while traditional roles (Data Scientist, ML Engineer) have evolved to incorporate LLM capabilities.
The AI Job Market Today
The AI job market spans 4,317 open positions across 15 role categories. The largest categories by volume: AI/ML Engineer (3,004), Data Scientist (345), AI Software Engineer (309). These three account for the majority of open positions, though smaller categories often have higher per-role compensation because of specialized skill requirements.
The seniority mix tells a story about where AI teams are in their maturity. Entry-level roles (138) are outnumbered by mid-level (2,071) and senior (1,655) positions, reflecting that most companies are past the 'build a team from scratch' phase and need experienced engineers who can ship production systems. Leadership roles (Director, VP, C-Level) total 453 positions, representing the bottleneck between technical execution and organizational strategy.
Remote work availability sits at 15% of all AI roles (635 positions), with 3,657 requiring on-site or hybrid attendance. The remote share has stabilized after the post-pandemic correction. Senior and specialized roles (Research Scientist, ML Architect) are more likely to be remote-eligible than entry-level positions, partly because experienced hires have more negotiating power and partly because these roles require less hands-on mentorship.
AI compensation is structured in clear tiers. The market median sits at $215,000. Top-quartile roles start at $266,300, and the 90th percentile reaches $320,790. These figures include base salary with disclosed compensation. Total compensation (including equity, bonuses, and sign-on) runs 20-40% higher at companies that offer those components.
Category matters for compensation. AI Safety roles lead at $287,500 median, while Prompt Engineer roles sit at $145,000. The spread between highest and lowest-paying categories reflects the premium on specialized technical skills versus broader analytical roles.
The most in-demand skills across all AI postings: Python (2,249 postings), Aws (1,224 postings), Azure (938 postings), Rag (915 postings), Gcp (660 postings), Pytorch (640 postings), Prompt Engineering (624 postings), Kubernetes (559 postings). Python dominates, appearing in the vast majority of role descriptions regardless of category. Cloud platform experience (AWS, GCP, Azure) is the second most common requirement. The newer entrants to the top skills list (RAG, vector databases, LLM APIs) reflect the shift from traditional ML toward generative AI applications.
Frequently Asked Questions
Get Weekly AI Career Intelligence
Salary data, skills demand, and market signals from 16,000+ AI job postings. Every Monday.