Interested in this Data Scientist role at Polaris I/O?
Apply Now →Skills & Technologies
About This Role
Lead Data Scientist
(Risk Modeling / Python / AWS / SageMaker)
Polaris I/O — Remote (U.S. only) \| Full\-time
About the company
Polaris I/O provides the world’s only integrated go\-to\-customer platform that combines transformation services, industry insights, and powerful technology to enable B2B companies to protect, retain, and grow large accounts. Our platforms take an outside\-in approach to improving commercial health, utilizing executive buyer insights to inform the orchestration of commercial activities that improve relationships, maximize growth, and scale beyond our engagement.
Role Summary
Polaris I/O is hiring a Staff Data Scientist to lead the design, validation, and ongoing health of the scoring models that power our risk and analytics platform. Our system ingests structured client data, news content, and public data sources to produce risk, threat, and anomaly scores across a range of business domains — with vendor and supply chain risk among the first use cases.
This is a foundational hire. You will be the first dedicated data scientist on the team and will own the model side of the platform end\-to\-end — from feature design and dataset preparation, through model authoring and validation, to deployment, monitoring, and iteration. You will partner closely with the Data Engineering team that builds and operates the underlying data platform, and with the CTO and product leadership on model strategy.
Our scoring architecture is intentionally model\-agnostic, so you will have meaningful influence over which models we build, how they are configured, and how we measure success. The platform is designed to support both formula\-based models and service\-based models deployed as endpoints.
Our stack includes:
- Python (pandas, scikit\-learn, statsmodels, NumPy)
- AWS — SageMaker, S3, Aurora PostgreSQL
- Jupyter / SageMaker Studio for exploratory and validation work
- YAML\-driven model and feature configuration
This role is hands\-on. You will write code, run validations, and own model quality from prototype through production.
What You’ll Do
Model Development
- Design, train, and validate scoring models across a range of risk domains
- Define features, weights, thresholds, lookback windows, and data handling rules in our model configuration format
- Build both formula\-based models and service\-based models, choosing the right approach for each problem
- Generate reproducible training, validation, and test datasets and document the splits clearly enough for any teammate to reproduce them
- Partner with engineering on the handoff from model development to production deployment
Model Validation \& Governance
- Establish validation standards and a governance framework for the platform — thresholds, cohort tests, and the criteria a model must meet before it ships
- Run validation studies with the right metrics for the model type (e.g., AUC\-ROC, precision/recall, calibration error, true/false positive rates, detection latency) and break results down by relevant cohorts
- Author and own recurring model health reports — covering drift, cohort performance, and score distribution
- Diagnose model failures in production and drive remediation
- Maintain reproducibility — every validation result should trace back to a specific dataset and a versioned model configuration
Model Lifecycle \& Operations
- Manage models through their full lifecycle, including planned replacements and re\-scoring of affected entities
- Define retraining cadence and triggers based on model health
- Produce explainability outputs (score rationale, summary explanations) suitable for client\-facing surfaces and audit
- Translate new modeling needs into requirements that the data engineering team can build into the underlying data pipeline
Cross\-Functional Leadership
- Serve as the data science point of contact for product, engineering, and platform teams
- Influence the long\-term direction of our model\-agnostic scoring architecture
- Mentor engineers and analysts on modeling fundamentals, validation rigor, and reproducibility
Required Job Qualifications
- 10\+ years of professional experience in data science, statistical modeling, or applied machine learning
- Demonstrated experience designing and shipping production models in at least one risk\-relevant domain: vendor risk, supply chain risk, threat assessment, fraud, anomaly detection, or comparable scoring/classification problems
- Expert\-level Python for data science work — pandas, scikit\-learn, statsmodels, NumPy
- Strong SQL skills, including the ability to write performant queries against PostgreSQL/Aurora for feature extraction and validation
- Deep expertise in at least one full statistical environment for exploratory analysis and validation: Python (Jupyter), R, SAS, or SPSS
- Production experience with model validation methodology — train/validation/test splits, cohort\-based validation, calibration, drift detection, and metric selection
- Experience deploying models as services (REST endpoints, SageMaker, or equivalent) and reasoning about latency, scaling, and operational behavior
- Working knowledge of AWS — S3 is required; familiarity with SageMaker, Aurora, and IAM is expected
- Strong written and verbal communication in English — this role includes model documentation, validation reports, governance frameworks, and collaboration with U.S.\-based teams and clients
Preferred Qualifications
- AWS SageMaker production experience (SageMaker Studio, model endpoints, model registry)
- Time series and anomaly detection methods — trend analysis, seasonality, spike detection, peer benchmarking
- Experience working with text\- or sentiment\-derived features (fine\-grained sentiment, emotion\-based sentiment, NLP\-extracted entities)
- Experience defining and operating formal model governance frameworks
- Familiarity with deep learning frameworks (PyTorch, TensorFlow)
- Experience working alongside Python/AWS engineering teams in an event\-driven microservice environment
Work Authorization Requirement: Candidates must be legally authorized to work in the United States on a full\-time, ongoing basis without the need for current or future employer sponsorship (for example, we are not able to sponsor employment visas now or in the future)Work Expectations
Work Hours: This is a full\-time role that requires you to be online and working during standard business hours, with occasional after\-hours support based on business needs.
Remote Work: This position is fully remote and all work must be performed within the United States. Candidates must be legally authorized to work in the United States. Regular working hours will align with the time zone.
Conflict of Interest: You are expected to devote your full professional time and attention to Polaris I/O and not take on other employment or contract work that interferes with your responsibilities, competes with the company, or creates a potential conflict of interest.
Compensation \& Benefits:
Polaris I/O offers a competitive salary, a comprehensive benefit package including medical, dental, and vision coverage, a generous PTO plan, and opportunities for professional growth and development.
Polaris I/O is an Equal Opportunity Employer. We consider all qualified applicants without regard to any characteristic protected by applicable federal, state, or local law.
Pay: $170,000\.00 \- $200,000\.00 per year
Benefits:
- Dental insurance
- Health insurance
- Vision insurance
Experience:
- in data science, modeling, or applied ML: 10 years (Preferred)
Work Location: Remote
Salary Context
This $170K-$200K range is above the median for Data Scientist roles in our dataset (median: $160K across 258 roles with salary data).
View full Data Scientist salary data →Role Details
About This Role
Data Scientists extract insights and build predictive models from data. In the AI era, many roles now include LLM-powered analytics, automated reporting, and integration with generative AI tools. The role has evolved from 'the person who runs SQL queries' to 'the person who builds AI-powered data products.'
Modern data science roles fall into two camps: analytics-focused (insights, dashboards, experimentation) and ML-focused (building predictive models, recommendation systems, NLP features). The best data scientists can operate in both modes. The AI shift means that even analytics-focused roles now involve building automated insight pipelines using LLMs, going well beyond one-off reports.
Across the 4,317 AI roles we're tracking, Data Scientist positions make up 8% of the market. At Polaris I/O, this role fits into their broader AI and engineering organization.
Data Scientist roles remain in high demand, though the definition keeps shifting. Companies increasingly want candidates who can bridge traditional statistics with modern ML and LLM capabilities. The 'pure insights' data scientist role is consolidating into analytics engineering, while the 'build models' data scientist role is merging with ML engineering.
What the Work Looks Like
A typical week includes: analyzing experiment results for a product feature launch, building a predictive model for customer churn, creating an automated reporting pipeline using LLM-powered summarization, presenting insights to stakeholders, and cleaning data (always cleaning data). The ratio of analysis to engineering varies by company, but expect both.
Data Scientist roles remain in high demand, though the definition keeps shifting. Companies increasingly want candidates who can bridge traditional statistics with modern ML and LLM capabilities. The 'pure insights' data scientist role is consolidating into analytics engineering, while the 'build models' data scientist role is merging with ML engineering.
Skills Required
Python, SQL, and statistical modeling are the foundation. Increasingly, roles want experience with LLMs for data analysis, automated insight generation, and building AI-powered data products. Familiarity with cloud data platforms (Snowflake, BigQuery, Databricks) and ML frameworks (scikit-learn, PyTorch) covers most job requirements.
Experimentation design and causal inference are underrated skills that separate strong candidates. Companies care about whether their product changes cause improvements, and can distinguish causation from correlation. A/B testing methodology, Bayesian statistics, and the ability to communicate uncertainty to non-technical stakeholders are high-value skills.
Good postings specify the data stack, the types of problems you'll work on, and the team structure. Look for companies that differentiate between analytics and ML data science. Vague 'data scientist' postings that list every skill under the sun usually mean the company doesn't know what they need.
Compensation Benchmarks
Data Scientist roles pay a median of $192,890 based on 789 positions with disclosed compensation. Senior-level AI roles across all categories have a median of $227,400. Disclosed range: $170K to $200K.
Across all AI roles, the market median is $215,000. Top-quartile compensation starts at $266,300. The 90th percentile reaches $320,790. For comparison, the highest-paying categories include AI Safety ($287,500) and Research Engineer ($272,100). By seniority level: Entry: $110,000; Mid: $194,400; Senior: $227,400; Director: $274,554; VP: $241,000.
Polaris I/O AI Hiring
Polaris I/O has 1 open AI role right now. They're hiring across Data Scientist. Based in US. Compensation range: $200K - $200K.
Location Context
AI roles in Austin pay a median of $214,343 across 143 tracked positions.
Career Path
Common paths into Data Scientist roles include Data Analyst, Statistician, Quantitative Researcher.
From here, career progression typically leads toward Senior Data Scientist, ML Engineer, AI Product Manager.
Start with statistics and SQL. Build a real analysis project on public data that demonstrates insight generation alongside model building. The market values data scientists who can communicate findings clearly to business stakeholders. If you want to move toward ML engineering, invest in software engineering fundamentals and production deployment skills.
What to Expect in Interviews
Interviews combine statistics, coding, and business acumen. SQL is almost always tested, often with complex joins and window functions. Expect a case study round where you're given a business problem and asked to design an analysis plan. Coding rounds focus on pandas, statistical modeling, and visualization. The strongest differentiator is how well you communicate insights to non-technical stakeholders during presentation rounds.
When evaluating opportunities: Good postings specify the data stack, the types of problems you'll work on, and the team structure. Look for companies that differentiate between analytics and ML data science. Vague 'data scientist' postings that list every skill under the sun usually mean the company doesn't know what they need.
AI Hiring Overview
The AI job market has 4,317 open positions tracked in our dataset. By seniority: 138 entry-level, 2,071 mid-level, 1,655 senior, and 453 leadership roles (Director, VP, C-Level). Remote roles make up 15% of the market (635 positions). The remaining 3,657 roles require on-site or hybrid attendance.
The market median for AI roles is $215,000. Top-quartile compensation starts at $266,300. The 90th percentile reaches $320,790. Highest-paying categories: AI Safety ($287,500 median, 34 roles); Research Engineer ($272,100 median, 227 roles); AI Engineering Manager ($244,000 median, 23 roles).
Data Scientist roles remain in high demand, though the definition keeps shifting. Companies increasingly want candidates who can bridge traditional statistics with modern ML and LLM capabilities. The 'pure insights' data scientist role is consolidating into analytics engineering, while the 'build models' data scientist role is merging with ML engineering.
The AI Job Market Today
The AI job market spans 4,317 open positions across 15 role categories. The largest categories by volume: AI/ML Engineer (3,004), Data Scientist (345), AI Software Engineer (309). These three account for the majority of open positions, though smaller categories often have higher per-role compensation because of specialized skill requirements.
The seniority mix tells a story about where AI teams are in their maturity. Entry-level roles (138) are outnumbered by mid-level (2,071) and senior (1,655) positions, reflecting that most companies are past the 'build a team from scratch' phase and need experienced engineers who can ship production systems. Leadership roles (Director, VP, C-Level) total 453 positions, representing the bottleneck between technical execution and organizational strategy.
Remote work availability sits at 15% of all AI roles (635 positions), with 3,657 requiring on-site or hybrid attendance. The remote share has stabilized after the post-pandemic correction. Senior and specialized roles (Research Scientist, ML Architect) are more likely to be remote-eligible than entry-level positions, partly because experienced hires have more negotiating power and partly because these roles require less hands-on mentorship.
AI compensation is structured in clear tiers. The market median sits at $215,000. Top-quartile roles start at $266,300, and the 90th percentile reaches $320,790. These figures include base salary with disclosed compensation. Total compensation (including equity, bonuses, and sign-on) runs 20-40% higher at companies that offer those components.
Category matters for compensation. AI Safety roles lead at $287,500 median, while Prompt Engineer roles sit at $145,000. The spread between highest and lowest-paying categories reflects the premium on specialized technical skills versus broader analytical roles.
The most in-demand skills across all AI postings: Python (2,249 postings), Aws (1,224 postings), Azure (938 postings), Rag (915 postings), Gcp (660 postings), Pytorch (640 postings), Prompt Engineering (624 postings), Kubernetes (559 postings). Python dominates, appearing in the vast majority of role descriptions regardless of category. Cloud platform experience (AWS, GCP, Azure) is the second most common requirement. The newer entrants to the top skills list (RAG, vector databases, LLM APIs) reflect the shift from traditional ML toward generative AI applications.
Frequently Asked Questions
Get Weekly AI Career Intelligence
Salary data, skills demand, and market signals from 16,000+ AI job postings. Every Monday.