Sr Data Scientist II

Raleigh, NC, US Senior Data Scientist

Interested in this Data Scientist role at RELX Group?

Apply Now →

Skills & Technologies

EmbeddingsPythonRag

About This Role

AI job market dashboard showing open roles by category

Are you excited about shaping the next generation of AI\-powered legal technology through generative AI, retrieval systems, and production\-grade machine learning?

Do you enjoy building reliable, scalable applications that transform complex AI capabilities into impactful customer solutions?

About our Team

LexisNexis Legal \& Professional, which serves customers in more than 150 countries with 11,800 employees worldwide, is part of RELX ( http://www.relx.com ), a global provider of information\-based analytics and decision tools for professional and business customers. Our company has been a long\-time leader in deploying AI and advanced technologies to the legal market to improve productivity and transform the overall business and practice of law, deploying ethical and powerful generative AI solutions with a flexible, multi\-model approach that prioritizes using the best model from today’s top model creators for each individual legal use case. The company employs over 2,000 technologists, data scientists, and experts to develop, test, and validate solutions in line with RELX Responsible AI Principles ( https://stories.relx.com/responsible\-ai\-principles/index.html ).

About the Role

We are looking for a Senior Data Scientist II with deep expertise in Generative AI, Retrieval\-Augmented Generation (RAG), and agentic AI systems, combined with strong software\-engineering fundamentals and demonstrated ownership of production applications. This role will focus on improving LLM\-powered drafting and retrieval solutions through advanced search, embeddings, reranking, evaluation, and production\-grade ML components.

The successful candidate must be able to independently design, refactor, test, review, deploy, and support clean, reliable Python applications. This includes separating agent responsibilities, designing for failure, applying sound algorithmic reasoning, and establishing appropriate logging, monitoring, testing, and operational controls.

The ideal candidate has advanced Python proficiency, experience with OpenSearch or Solr, success working in monorepo environments, and a strong record of cross\-functional delivery.

Key Responsibilities

  • Architect modular agentic applications with clear separation among retrieval, prompt construction, model invocation, tool execution, state and history management, orchestration, validation, and response formatting.
  • Independently refactor complex or legacy Python code to improve correctness, readability, modularity, extensibility, testability, and runtime performance.
  • Own production readiness for AI components, including input validation, exception handling, timeout management, retries with backoff, fallback behavior, configuration management, and secure handling of credentials.
  • Establish observability for LLM and retrieval workflows through structured logging, metrics, distributed tracing, alerting, and actionable error reporting.
  • Design clear interfaces and data contracts between retrieval, orchestration, model, and downstream application components.
  • Write comprehensive unit, integration, regression, and end\-to\-end tests, including tests for failure modes, malformed model responses, empty retrieval results, and unavailable dependencies.
  • Review Python and agentic application code, identify architectural and operational risks, and provide actionable feedback aligned with production engineering standards.
  • Diagnose and optimize latency, memory usage, retrieval performance, token consumption, model cost, and application scalability.
  • Apply appropriate data structures, algorithms, and computational\-complexity analysis when designing and optimizing solutions.
  • Participate in production deployments, incident investigation, root\-cause analysis, remediation, and continuous reliability improvements.

Required Qualifications

  • Advanced Python proficiency demonstrated through independently designing, implementing, debugging, testing, reviewing, and refactoring production applications.
  • Strong command of Python fundamentals, standard data structures, common algorithms, object\-oriented and functional design principles, type annotations, and time and space complexity analysis.
  • Demonstrated ability to transform prototype or experimental code into modular, maintainable, observable, and production\-ready systems.
  • Strong understanding of software design principles, including separation of concerns, dependency injection, interface design, configuration management, and effective abstraction.
  • Experience implementing automated unit, integration, regression, and end\-to\-end testing using tools such as pytest, including appropriate mocking of external services.
  • Experience designing resilient distributed applications that account for timeouts, retries, rate limits, partial failures, malformed responses, idempotency, and graceful degradation.
  • Experience with production observability, including structured logging, metrics, tracing, alerting, and incident troubleshooting.
  • Demonstrated ability to conduct rigorous code reviews and identify correctness, maintainability, performance, security, testing, and operational risks.
  • Experience taking technical ownership of applications across their lifecycle, from design and experimentation through deployment, monitoring, incident response, and ongoing improvement.
  • Strong understanding of production LLM concerns, including structured output validation, context management, model and tool failures, prompt versioning, token and cost controls, security, and evaluation.

Preferred Qualifications

  • Experience with Python quality tooling such as pytest, ruff, mypy, profiling tools, and automated CI quality gates.
  • Experience defining typed schemas and validating LLM inputs and outputs using tools such as Pydantic.
  • Experience building evaluation frameworks for agentic systems, including task\-completion, retrieval\-quality, groundedness, hallucination, latency, reliability, and cost metrics.
  • Experience implementing model fallbacks, tool\-use controls, guardrails, human\-in\-the\-loop workflows, and auditability for AI applications.
  • Experience supporting production services and participating in incident response, root\-cause analysis, and post\-incident remediation.

The successful candidate will:

  • Demonstrate senior\-level Python proficiency and sound computer\-science fundamentals.
  • Treat correctness, maintainability, testing, resilience, security, and observability as core design requirements.
  • Recognize architectural issues and improve code beyond simply making it functional.
  • Independently review and refactor complex agentic application code.
  • Make clear engineering tradeoffs involving quality, latency, scalability, reliability, and cost.
  • Take end\-to\-end ownership from experimentation through production deployment and operational support.
  • Communicate technical decisions and code\-review feedback clearly and constructively.
  • Combine strong LLM and retrieval expertise with disciplined software\-engineering practices.

Work in a Way That Works for You

We promote a healthy work/life balance across the organisation. We offer an appealing working prospect for our people. With numerous wellbeing initiatives, shared parental leave, study assistance and sabbaticals, we will help you meet your immediate responsibilities and your long\-term goals.

Working Pattern

Working flexible hours \- flexing the times when you work in the day to help you fit everything in and work when you are the most productive.

About the Business

LexisNexis Legal \& Professional® provides legal, regulatory, and business information and analytics that help customers increase their productivity, improve decision\-making, achieve better outcomes, and advance the rule of law around the world. As a digital pioneer, the company was the first to bring legal and business information online with its Lexis® and Nexis® services. \#AIFluent

We know your well\-being and happiness are key to a long and successful career. We are delighted to offer country specific benefits. Click here to access benefits specific to your location.

We are committed to providing a fair and accessible hiring process. If you have a disability or other need that requires accommodation or adjustment, please let us know by completing our Applicant Request Support Form or please contact 1\-855\-833\-5120\.

Criminals may pose as recruiters asking for money or personal information. We never request money or banking details from job applicants. Learn more about spotting and avoiding scams here .

Please read our Candidate Privacy Policy .

We are an equal opportunity employer: qualified applicants are considered for and treated during employment without regard to race, color, creed, religion, sex, national origin, citizenship status, disability status, protected veteran status, age, marital status, sexual orientation, gender identity, genetic information, or any other characteristic protected by law.

*USA Job Seekers:*

EEO Know Your Rights .

Role Details

Company RELX Group
Title Sr Data Scientist II
Location Raleigh, NC, US
Category Data Scientist
Experience Senior
Salary Not disclosed
Remote No

About This Role

Data Scientists extract insights and build predictive models from data. In the AI era, many roles now include LLM-powered analytics, automated reporting, and integration with generative AI tools. The role has evolved from 'the person who runs SQL queries' to 'the person who builds AI-powered data products.'

Modern data science roles fall into two camps: analytics-focused (insights, dashboards, experimentation) and ML-focused (building predictive models, recommendation systems, NLP features). The best data scientists can operate in both modes. The AI shift means that even analytics-focused roles now involve building automated insight pipelines using LLMs, going well beyond one-off reports.

Across the 4,317 AI roles we're tracking, Data Scientist positions make up 8% of the market. At RELX Group, this role fits into their broader AI and engineering organization.

Data Scientist roles remain in high demand, though the definition keeps shifting. Companies increasingly want candidates who can bridge traditional statistics with modern ML and LLM capabilities. The 'pure insights' data scientist role is consolidating into analytics engineering, while the 'build models' data scientist role is merging with ML engineering.

What the Work Looks Like

A typical week includes: analyzing experiment results for a product feature launch, building a predictive model for customer churn, creating an automated reporting pipeline using LLM-powered summarization, presenting insights to stakeholders, and cleaning data (always cleaning data). The ratio of analysis to engineering varies by company, but expect both.

Data Scientist roles remain in high demand, though the definition keeps shifting. Companies increasingly want candidates who can bridge traditional statistics with modern ML and LLM capabilities. The 'pure insights' data scientist role is consolidating into analytics engineering, while the 'build models' data scientist role is merging with ML engineering.

Skills Required

Embeddings (7% of roles) Python (52% of roles) Rag (21% of roles)

Python, SQL, and statistical modeling are the foundation. Increasingly, roles want experience with LLMs for data analysis, automated insight generation, and building AI-powered data products. Familiarity with cloud data platforms (Snowflake, BigQuery, Databricks) and ML frameworks (scikit-learn, PyTorch) covers most job requirements.

Experimentation design and causal inference are underrated skills that separate strong candidates. Companies care about whether their product changes cause improvements, and can distinguish causation from correlation. A/B testing methodology, Bayesian statistics, and the ability to communicate uncertainty to non-technical stakeholders are high-value skills.

Good postings specify the data stack, the types of problems you'll work on, and the team structure. Look for companies that differentiate between analytics and ML data science. Vague 'data scientist' postings that list every skill under the sun usually mean the company doesn't know what they need.

Compensation Benchmarks

Data Scientist roles pay a median of $192,890 based on 789 positions with disclosed compensation. Senior-level AI roles across all categories have a median of $227,400.

Across all AI roles, the market median is $215,000. Top-quartile compensation starts at $266,300. The 90th percentile reaches $320,790. For comparison, the highest-paying categories include AI Safety ($287,500) and Research Engineer ($272,100). By seniority level: Entry: $110,000; Mid: $194,400; Senior: $227,400; Director: $274,554; VP: $241,000.

RELX Group AI Hiring

RELX Group has 2 open AI roles right now. They're hiring across AI/ML Engineer, Data Scientist. Positions span Alpharetta, GA, US, Raleigh, NC, US. Compensation range: $219K - $219K.

Location Context

Across all AI roles, 15% (635 positions) offer remote work, while 3,657 require on-site attendance. Top AI hiring metros: New York (1,650 roles, $220,000 median); San Francisco (1,335 roles, $265,000 median); Los Angeles (708 roles, $214,112 median).

Career Path

Common paths into Data Scientist roles include Data Analyst, Statistician, Quantitative Researcher.

From here, career progression typically leads toward Senior Data Scientist, ML Engineer, AI Product Manager.

Start with statistics and SQL. Build a real analysis project on public data that demonstrates insight generation alongside model building. The market values data scientists who can communicate findings clearly to business stakeholders. If you want to move toward ML engineering, invest in software engineering fundamentals and production deployment skills.

What to Expect in Interviews

Interviews combine statistics, coding, and business acumen. SQL is almost always tested, often with complex joins and window functions. Expect a case study round where you're given a business problem and asked to design an analysis plan. Coding rounds focus on pandas, statistical modeling, and visualization. The strongest differentiator is how well you communicate insights to non-technical stakeholders during presentation rounds.

When evaluating opportunities: Good postings specify the data stack, the types of problems you'll work on, and the team structure. Look for companies that differentiate between analytics and ML data science. Vague 'data scientist' postings that list every skill under the sun usually mean the company doesn't know what they need.

AI Hiring Overview

The AI job market has 4,317 open positions tracked in our dataset. By seniority: 138 entry-level, 2,071 mid-level, 1,655 senior, and 453 leadership roles (Director, VP, C-Level). Remote roles make up 15% of the market (635 positions). The remaining 3,657 roles require on-site or hybrid attendance.

The market median for AI roles is $215,000. Top-quartile compensation starts at $266,300. The 90th percentile reaches $320,790. Highest-paying categories: AI Safety ($287,500 median, 34 roles); Research Engineer ($272,100 median, 227 roles); AI Engineering Manager ($244,000 median, 23 roles).

Data Scientist roles remain in high demand, though the definition keeps shifting. Companies increasingly want candidates who can bridge traditional statistics with modern ML and LLM capabilities. The 'pure insights' data scientist role is consolidating into analytics engineering, while the 'build models' data scientist role is merging with ML engineering.

The AI Job Market Today

The AI job market spans 4,317 open positions across 15 role categories. The largest categories by volume: AI/ML Engineer (3,004), Data Scientist (345), AI Software Engineer (309). These three account for the majority of open positions, though smaller categories often have higher per-role compensation because of specialized skill requirements.

The seniority mix tells a story about where AI teams are in their maturity. Entry-level roles (138) are outnumbered by mid-level (2,071) and senior (1,655) positions, reflecting that most companies are past the 'build a team from scratch' phase and need experienced engineers who can ship production systems. Leadership roles (Director, VP, C-Level) total 453 positions, representing the bottleneck between technical execution and organizational strategy.

Remote work availability sits at 15% of all AI roles (635 positions), with 3,657 requiring on-site or hybrid attendance. The remote share has stabilized after the post-pandemic correction. Senior and specialized roles (Research Scientist, ML Architect) are more likely to be remote-eligible than entry-level positions, partly because experienced hires have more negotiating power and partly because these roles require less hands-on mentorship.

AI compensation is structured in clear tiers. The market median sits at $215,000. Top-quartile roles start at $266,300, and the 90th percentile reaches $320,790. These figures include base salary with disclosed compensation. Total compensation (including equity, bonuses, and sign-on) runs 20-40% higher at companies that offer those components.

Category matters for compensation. AI Safety roles lead at $287,500 median, while Prompt Engineer roles sit at $145,000. The spread between highest and lowest-paying categories reflects the premium on specialized technical skills versus broader analytical roles.

The most in-demand skills across all AI postings: Python (2,249 postings), Aws (1,224 postings), Azure (938 postings), Rag (915 postings), Gcp (660 postings), Pytorch (640 postings), Prompt Engineering (624 postings), Kubernetes (559 postings). Python dominates, appearing in the vast majority of role descriptions regardless of category. Cloud platform experience (AWS, GCP, Azure) is the second most common requirement. The newer entrants to the top skills list (RAG, vector databases, LLM APIs) reflect the shift from traditional ML toward generative AI applications.

Frequently Asked Questions

Based on 789 roles with disclosed compensation, the median salary for Data Scientist positions is $192,890. Actual compensation varies by seniority, location, and company stage.
Python, SQL, and statistical modeling are the foundation. Increasingly, roles want experience with LLMs for data analysis, automated insight generation, and building AI-powered data products. Familiarity with cloud data platforms (Snowflake, BigQuery, Databricks) and ML frameworks (scikit-learn, PyTorch) covers most job requirements.
About 15% of the 4,317 AI roles we track offer remote work. Remote availability varies by company and seniority level, with senior and leadership roles more likely to offer location flexibility.
RELX Group is among the companies actively hiring for AI and ML talent. Check our company profiles for detailed breakdowns of open roles, salary ranges, and hiring trends.
Common next steps from Data Scientist positions include Senior Data Scientist, ML Engineer, AI Product Manager. Progression depends on whether you lean toward technical depth, people management, or product strategy.

Get Weekly AI Career Intelligence

Salary data, skills demand, and market signals from 16,000+ AI job postings. Every Monday.