Interested in this Data Engineer role at Capital One?
Apply Now →Skills & Technologies
About This Role
Lead Data Engineer (Python, AWS, Spark, Kafka, SQL, Snowflake, Databricks, GenAI)
Do you love building and pioneering in the technology space? Do you enjoy solving complex business problems in a fast\-paced, collaborative, inclusive, and iterative delivery environment?
At Capital One, you'll be part of a big group of makers, breakers, doers and disruptors, who solve real problems and meet real customer needs. We are seeking Data Engineers who arepassionate about data, data modeling, data quality, governance and permissible use, and building data pipelines with emerging technologies. As a Capital One Lead Data Engineer, you’ll have the opportunity to be on the forefront of driving a major transformation within Capital One.
The Marketing and Messaging team is responsible for delivering hyper\-personalized messages and experiences that will delight the customer, attract prospects and drive increasing business value. The team builds scalable platforms that deliver omnichannel messages in owned and paid Adtech channels.
What You’ll Do:
- Collaborate with and across Agile teams to design, develop, test, implement, and support technical solutions using data movement tools and technologies
- As a lead developer, work with a team of developers who have deep experience in data movement, distributed computing, and full\-stack systems
- Utilize programming languages like Python, SQL, and Open Source RDBMS and NoSQL databases and cloud\-based data warehousing services such as Snowflake and Databricks
- Optimize information systems for end\-users and downstream application consumers by using sound data design practices
- Share your passion for staying on top of tech trends, experimenting with and learning new technologies, participating in internal \& external technology communities, and mentoring other members of the engineering community
- Collaborate with digital product managers, and deliver robust cloud\-based solutions that drive powerful experiences to help millions of Americans achieve financial empowerment
- Perform unit tests and conduct reviews with other team members to make sure your code is rigorously designed, elegantly coded, and effectively tuned for performance
Basic Qualifications:
- Bachelor’s Degree
- At least 4 years of experience in application development (Internship experience does not apply)
- At least 2 years of experience in big data technologies
- At least 1 year experience with cloud computing (AWS, Microsoft Azure, Google Cloud)
Preferred Qualifications:
- Master's Degree
- 7\+ years of experience in application development including Python, SQL, Spark, ETL tools, or AWS Glue
- 4\+ years of experience with a public cloud (AWS, Microsoft Azure, Google Cloud)
- 4\+ years experience with Distributed data/computing tools (MapReduce, Hadoop, Hive, EMR, Kafka, Spark, or MySQL)
- 4\+ year experience working on real\-time data and streaming applications
- 4\+ years of experience with NoSQL implementation (Mongo, Cassandra)
- 4\+ years of data warehousing experience (Redshift or Snowflake)
- 4\+ years of experience with UNIX/Linux including basic commands and shell scripting
- 4\+ years of experience with data modeling for data warehousing
- 2\+ years of experience with Agile engineering practices
- Experience leveraging interactive AI tooling (Claude Code, GitHub Copilot) to accelerate software delivery
*At this time, Capital One will not sponsor a new applicant for employment authorization, or offer any immigration related support for this position (e.g. H1B, F\-1 OPT, F\-1 STEM OPT, F\-1 CPT, J\-1, TN, E\-3, and O\-1, or any other forms of work authorization that require immigration support from an employer).*
The minimum and maximum full\-time annual salaries for this role are listed below, by location. Please note that this salary information is solely for candidates hired to perform work within one of these locations, and refers to the amount Capital One is willing to pay at the time of this posting. Salaries for part\-time roles will be prorated based upon the agreed upon number of hours to be regularly worked.
McLean, VA: $197,300 \- $225,100 for Lead Data Engineer
New York, NY: $215,200 \- $245,600 for Lead Data Engineer
Richmond, VA: $179,400 \- $204,700 for Lead Data Engineer
Candidates hired to work in other locations will be subject to the pay range associated with that location, and the actual annualized salary amount offered to any candidate at the time of hire will be reflected solely in the candidate’s offer letter.
This role is also eligible to earn performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI). Incentives could be discretionary or non discretionary depending on the plan.
Capital One offers a comprehensive, competitive, and inclusive set of health, financial and other benefits that support your total well\-being. Learn more at the Capital One Careers website. Eligibility varies based on full or part\-time status, exempt or non\-exempt status, and management level.
This role is expected to accept applications for a minimum of 5 business days.
No agencies please. Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non\-discrimination in compliance with applicable federal, state, and local laws. Capital One promotes a drug\-free workplace. Capital One will consider for employment qualified applicants with a criminal history in a manner consistent with the requirements of applicable laws regarding criminal background inquiries, including, to the extent applicable, Article 23\-A of the New York Correction Law; San Francisco, California Police Code Article 49, Sections 4901\-4920; New York City’s Fair Chance Act; Philadelphia’s Fair Criminal Records Screening Act; and other applicable federal, state, and local laws and regulations regarding criminal background inquiries.
If you have visited our website in search of information on employment opportunities or to apply for a position, and you require an accommodation, please contact Capital One Recruiting at 1\-800\-304\-9102 or via email at [email protected]. All information you provide will be kept confidential and will be used only to the extent required to provide needed reasonable accommodations.
For technical support or questions about Capital One's recruiting process, please send an email to [email protected]
Capital One does not provide, endorse nor guarantee and is not liable for third\-party products, services, educational tools or other information available through this site.
Capital One Financial is made up of several different entities. Please note that any position posted in Canada is for Capital One Canada, any position posted in the United Kingdom is for Capital One Europe and any position posted in the Philippines is for Capital One Philippines Service Corp. (COPSSC).
Role Details
About This Role
Data Engineers build the pipelines that feed AI models. They design ETL workflows, manage data lakes, and ensure training and inference data is clean, timely, and accessible. Without good data engineering, AI projects fail. It's that simple.
The AI era has expanded the data engineer's scope far beyond batch ETL jobs. You're building real-time embedding pipelines for RAG systems, managing vector databases, ensuring training data quality at scale, and building the infrastructure that lets ML teams iterate on data as fast as they iterate on models. Data quality is the biggest predictor of model quality, and you're the person responsible for it.
Across the 4,317 AI roles we're tracking, Data Engineer positions make up 1% of the market. At Capital One, this role fits into their broader AI and engineering organization.
Data Engineer demand in AI contexts is strong and growing. Every company building AI needs clean, reliable data pipelines. The shift toward real-time AI applications (chatbots, recommendation engines, agent systems) means data engineering is more critical than ever. Companies are willing to pay premium salaries for data engineers with AI/ML pipeline experience.
What the Work Looks Like
A typical week includes: debugging a data pipeline that's producing stale embeddings for the RAG system, optimizing a Spark job that processes training data, building a data quality monitoring dashboard, meeting with the ML team to understand their next data requirements, and writing dbt models that transform raw event data into ML-ready features. The work is deeply technical and high-impact.
Data Engineer demand in AI contexts is strong and growing. Every company building AI needs clean, reliable data pipelines. The shift toward real-time AI applications (chatbots, recommendation engines, agent systems) means data engineering is more critical than ever. Companies are willing to pay premium salaries for data engineers with AI/ML pipeline experience.
Skills Required
SQL, Python, and distributed systems (Spark, Airflow, dbt) are core. Cloud data platforms (Snowflake, BigQuery, Redshift) are increasingly standard. Many AI-focused roles also want familiarity with vector databases and embedding pipelines. Understanding data modeling, pipeline orchestration, and data quality frameworks covers the essentials.
AI-specific data engineering skills include: building feature stores, managing training data versioning, implementing data lineage tracking, and building real-time embedding pipelines. Experience with streaming systems (Kafka, Flink) is valuable for real-time AI applications. Understanding ML data requirements (balanced datasets, data augmentation, evaluation set construction) makes you much more effective working with ML teams.
Strong postings specify the data stack, mention ML pipeline work, and describe the scale of data you'll be working with. Look for companies that understand the connection between data quality and model quality. Avoid roles that conflate data engineering with data analysis.
Compensation Benchmarks
Data Engineer roles pay a median of $185,000 based on 83 positions with disclosed compensation. Senior-level AI roles across all categories have a median of $227,400.
Across all AI roles, the market median is $215,000. Top-quartile compensation starts at $266,300. The 90th percentile reaches $320,790. For comparison, the highest-paying categories include AI Safety ($287,500) and Research Engineer ($272,100). By seniority level: Entry: $110,000; Mid: $194,400; Senior: $227,400; Director: $274,554; VP: $241,000.
Capital One AI Hiring
Capital One has 24 open AI roles right now. They're hiring across AI/ML Engineer, AI Product Manager, Data Scientist, AI Software Engineer. Positions span McLean, VA, US, New York, NY, US, Plano, TX, US.
Location Context
AI roles in New York pay a median of $220,000 across 1,650 tracked positions.
Career Path
Common paths into Data Engineer roles include Backend Engineer, Database Administrator, Analytics Engineer.
From here, career progression typically leads toward Senior Data Engineer, ML Engineer, Data Platform Lead.
Master SQL and Python first. Then learn a distributed processing framework (Spark or its modern alternatives) and a pipeline orchestrator (Airflow, Dagster, Prefect). Build a portfolio project that demonstrates end-to-end pipeline construction: ingest, transform, validate, serve. If you want to specialize in AI data engineering, add vector databases and embedding pipelines to your skill set.
What to Expect in Interviews
Expect SQL deep-dives (query optimization, partitioning strategies, data modeling), Python coding focused on data pipeline patterns, and system design questions about building scalable ETL workflows. Companies with ML teams will ask about feature stores, embedding pipelines, and training data management. Be ready to discuss data quality monitoring, pipeline orchestration, and how you'd handle schema evolution in a production data lake.
When evaluating opportunities: Strong postings specify the data stack, mention ML pipeline work, and describe the scale of data you'll be working with. Look for companies that understand the connection between data quality and model quality. Avoid roles that conflate data engineering with data analysis.
AI Hiring Overview
The AI job market has 4,317 open positions tracked in our dataset. By seniority: 138 entry-level, 2,071 mid-level, 1,655 senior, and 453 leadership roles (Director, VP, C-Level). Remote roles make up 15% of the market (635 positions). The remaining 3,657 roles require on-site or hybrid attendance.
The market median for AI roles is $215,000. Top-quartile compensation starts at $266,300. The 90th percentile reaches $320,790. Highest-paying categories: AI Safety ($287,500 median, 34 roles); Research Engineer ($272,100 median, 227 roles); AI Engineering Manager ($244,000 median, 23 roles).
Data Engineer demand in AI contexts is strong and growing. Every company building AI needs clean, reliable data pipelines. The shift toward real-time AI applications (chatbots, recommendation engines, agent systems) means data engineering is more critical than ever. Companies are willing to pay premium salaries for data engineers with AI/ML pipeline experience.
The AI Job Market Today
The AI job market spans 4,317 open positions across 15 role categories. The largest categories by volume: AI/ML Engineer (3,004), Data Scientist (345), AI Software Engineer (309). These three account for the majority of open positions, though smaller categories often have higher per-role compensation because of specialized skill requirements.
The seniority mix tells a story about where AI teams are in their maturity. Entry-level roles (138) are outnumbered by mid-level (2,071) and senior (1,655) positions, reflecting that most companies are past the 'build a team from scratch' phase and need experienced engineers who can ship production systems. Leadership roles (Director, VP, C-Level) total 453 positions, representing the bottleneck between technical execution and organizational strategy.
Remote work availability sits at 15% of all AI roles (635 positions), with 3,657 requiring on-site or hybrid attendance. The remote share has stabilized after the post-pandemic correction. Senior and specialized roles (Research Scientist, ML Architect) are more likely to be remote-eligible than entry-level positions, partly because experienced hires have more negotiating power and partly because these roles require less hands-on mentorship.
AI compensation is structured in clear tiers. The market median sits at $215,000. Top-quartile roles start at $266,300, and the 90th percentile reaches $320,790. These figures include base salary with disclosed compensation. Total compensation (including equity, bonuses, and sign-on) runs 20-40% higher at companies that offer those components.
Category matters for compensation. AI Safety roles lead at $287,500 median, while Prompt Engineer roles sit at $145,000. The spread between highest and lowest-paying categories reflects the premium on specialized technical skills versus broader analytical roles.
The most in-demand skills across all AI postings: Python (2,249 postings), Aws (1,224 postings), Azure (938 postings), Rag (915 postings), Gcp (660 postings), Pytorch (640 postings), Prompt Engineering (624 postings), Kubernetes (559 postings). Python dominates, appearing in the vast majority of role descriptions regardless of category. Cloud platform experience (AWS, GCP, Azure) is the second most common requirement. The newer entrants to the top skills list (RAG, vector databases, LLM APIs) reflect the shift from traditional ML toward generative AI applications.
Frequently Asked Questions
Get Weekly AI Career Intelligence
Salary data, skills demand, and market signals from 16,000+ AI job postings. Every Monday.