Interested in this Data Scientist role at Replit?
Apply Now →Skills & Technologies
About This Role
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation.
About the Role
------------------
We're redefining how software is built and who gets to build it. Our mission is to achieve Autonomy for All: making programming accessible, collaborative, and powered by AI. Realizing that vision requires a platform that legitimate users can trust and adversarial actors cannot exploit.
We're hiring a Data Scientist to help build Replit's Trust \& Safety and Anti\-Abuse program from the ground up. You'll turn noisy behavioral, identity, payment, infrastructure, and content signals into the measurement systems, detections, and decisions that protect Replit's users, platform, and economics. You'll work closely with Engineering, Support, Legal, Security, Infrastructure, Money, and Growth to make abuse economically unviable while keeping friction low for legitimate users.
Replit sits at the frontier of AI\-native abuse. Our platform is a target for phishing and scam hosting, cryptomining, LLM token farming, card and coupon fraud, referral abuse, and increasingly, abuse driven by AI agents themselves. You'll help define how we identify, measure, and respond to these threats without compromising the experience of good users.
Who You Are
---------------
You're a data scientist who moves fast, goes deep, and thinks adversarially. You can spin up an analysis in hours that would take others days, not by cutting corners, but because you've built the intuition and technical toolkit to get to the right answer quickly. You dig past the top\-line abuse rate to understand selection effects, missing labels, policy changes, attacker adaptation, and the false positives hidden inside an aggregate metric.
You understand that Trust \& Safety data is imperfect and outcomes are high stakes. Ground truth is delayed, biased, and often incomplete; attackers react to defenses; and an apparently effective rule can quietly harm legitimate users. You pressure\-test your own work, quantify uncertainty, and distinguish correlation from evidence strong enough to justify enforcement.
You use AI agents and tools aggressively to multiply your output: writing code, exploring data, generating hypotheses, and prototyping investigations. But you treat every AI\-assisted output as a draft, not a deliverable. You know what good analysis looks like and won't ship anything that doesn't meet that bar.
You Will
------------
- Own the analytical foundation for Trust \& Safety, including abuse prevalence, fraud loss, false\-positive and false\-negative rates, time to detect, time to mitigate, appeal and reversal rates, and verification step\-up conversion.
- Build reliable datasets and dbt models that connect product events, account and identity signals, payment activity, infrastructure usage, content classifications, enforcement actions, appeals, and support outcomes.
- Develop and evaluate risk models, rules, and anomaly\-detection systems for threats such as phishing, scam hosting, cryptomining, token farming, payment fraud, promotional abuse, and AI\-agent exploitation.
- Design rigorous offline evaluations, shadow\-mode tests, holdouts, and controlled experiments to measure detection quality and the user impact of new policies, enforcement actions, and progressive verification.
- Define thresholds and decision frameworks that balance abuse reduction, economic loss, customer friction, and false positives across free, paid, and enterprise users.
- Investigate emerging abuse patterns, quantify their impact, identify coordinated behavior, and turn ambiguous signals into clear recommendations for product and engineering teams.
- Develop predictive models that estimate account, device, transaction, workspace, or deployment risk and embed those signals into detection, review, and escalation workflows.
- Partner with Support and Legal to improve case review, appeals, reason\-code quality, and feedback loops so human decisions become useful model and policy signals.
- Build monitoring that detects model drift, attacker adaptation, data\-quality failures, and unexpected harm to legitimate users.
- Communicate findings clearly to technical and non\-technical partners, including the tradeoffs, uncertainty, and evidence behind high\-impact decisions.
Examples of What You Could Do
---------------------------------
- Build a measurement framework for Replit's abuse surface, reconcile incomplete labels across automated detections, human review, appeals, chargebacks, and support cases, and establish a trustworthy baseline for the first time.
- Design and evaluate a risk\-scoring model for suspicious account clusters using identity, device, payment, graph, and product\-behavior signals, then define thresholds that materially reduce fraud while protecting legitimate users.
- Analyze a phishing detection rule that appears highly precise, uncover that it disproportionately bans paying users with legitimate brand references, and redesign its evaluation and review path to reduce false positives.
- Measure a progressive verification "ladder of trust," determining when to step users up to additional verification and quantifying the tradeoff between abuse prevented and legitimate\-user conversion lost.
- Detect coordinated token\-farming or promotional\-abuse networks by combining account\-linkage graphs, referral behavior, payment patterns, and infrastructure usage, then partner with Engineering to operationalize the findings.
- Evaluate a new enforcement policy in shadow mode, estimate its counterfactual impact, and recommend whether to launch, revise, or reject it before any users are affected.
Required Skills and Experience
----------------------------------
- 5\+ years of experience in data science, product analytics, fraud, risk, trust and safety, or a related field.
- Strong SQL and Python skills, with experience working with large behavioral datasets and building reliable data models or pipelines.
- Experience developing and evaluating predictive models, experiments, or decision systems, with sound judgment around uncertainty and tradeoffs.
- Ability to turn ambiguous data into clear recommendations and communicate them effectively across technical and non\-technical teams.
- Comfort working with imperfect labels, biased samples, and high\-impact decisions where false positives matter.
- You use AI tools extensively to increase your effectiveness while maintaining a high bar for analytical quality.
Preferred Qualifications
----------------------------
- Experience building or evaluating anti\-abuse, fraud, identity, security, spam, integrity, or content\-safety systems at scale.
- Built, shipped, and maintained ML models in production (classification, anomaly detection, or risk scoring), including feature engineering on behavioral and transaction data, threshold selection against precision/recall economics, and post\-launch monitoring
- Experience with graph analysis, entity resolution, coordinated\-behavior detection, reputation systems, anomaly detection, or risk scoring.
- Experience measuring false positives and enforcement harm, designing human\-review workflows, or using appeals and case outcomes as model feedback.
- Familiarity with progressive verification, KYC, account trust, or identity providers such as Prove, Persona, Socure, or Stripe Identity.
- Experience with causal inference methods such as difference\-in\-differences, propensity score methods, synthetic control, or uplift modeling.
- Experience with a modern data stack such as dbt, BigQuery, Snowflake, Fivetran, Amplitude, Mixpanel, or Segment.
- Experience at a consumer platform, developer tool, cloud provider, marketplace, fintech company, or other product with a meaningful adversarial surface.
Bonus Points
----------------
- You've built AI\-powered analytical tools, investigation systems, automated detections, or novel measurement approaches.
- You have experience with AI\-native abuse such as prompt injection, LLM token farming, model extraction, or agent\-driven abuse.
- You understand freemium, usage\-based, or promotional pricing models and the abuse incentives they create.
- You've worked directly with operational review teams and can translate analytical signals into practical playbooks, queues, and escalation paths.
*This is a full\-time role that can be held from our Foster City, CA office. The role has an in\-office requirement of Monday, Wednesday, and Friday.*
Full\-Time Employee Benefits Include:
Competitive Salary \& Equity
401(k) Program with a 4% match (*US Only*)
Health, Dental, Vision and Life Insurance
Short Term and Long Term Disability
Paid Parental, Medical, Caregiver Leave
Flexible Time Off (FTO) \+ Holidays
Commuter Benefits (*In\-Office Only*)
Monthly Wellness Stipend
Autonomous Work Environment
In Office Set\-Up Reimbursement (*In\-Office Only*)
Quarterly Team Gatherings
In Office Amenities (*In\-Office Only*)
Want to learn more about what we are up to?
- Meet the Replit Agent
- Replit: Make an app for that
- Replit Blog
- Amjad TED Talk
Interviewing \+ Culture at Replit
- Operating Principles
- Reasons not to work at Replit
To achieve our mission of making programming more accessible around the world, we need our team to be representative of the world. We welcome your unique perspective and experiences in shaping this product. We encourage people from all kinds of backgrounds to apply, including and especially candidates from underrepresented and non\-traditional backgrounds.
Compensation Range: $210K \- $310K
Salary Context
This $210K-$310K range is above the 75th percentile for Data Scientist roles in our dataset (median: $155K across 226 roles with salary data).
View full Data Scientist salary data →Role Details
About This Role
Data Scientists extract insights and build predictive models from data. In the AI era, many roles now include LLM-powered analytics, automated reporting, and integration with generative AI tools. The role has evolved from 'the person who runs SQL queries' to 'the person who builds AI-powered data products.'
Modern data science roles fall into two camps: analytics-focused (insights, dashboards, experimentation) and ML-focused (building predictive models, recommendation systems, NLP features). The best data scientists can operate in both modes. The AI shift means that even analytics-focused roles now involve building automated insight pipelines using LLMs, going well beyond one-off reports.
Across the 3,708 AI roles we're tracking, Data Scientist positions make up 8% of the market. At Replit, this role fits into their broader AI and engineering organization.
Data Scientist roles remain in high demand, though the definition keeps shifting. Companies increasingly want candidates who can bridge traditional statistics with modern ML and LLM capabilities. The 'pure insights' data scientist role is consolidating into analytics engineering, while the 'build models' data scientist role is merging with ML engineering.
What the Work Looks Like
A typical week includes: analyzing experiment results for a product feature launch, building a predictive model for customer churn, creating an automated reporting pipeline using LLM-powered summarization, presenting insights to stakeholders, and cleaning data (always cleaning data). The ratio of analysis to engineering varies by company, but expect both.
Data Scientist roles remain in high demand, though the definition keeps shifting. Companies increasingly want candidates who can bridge traditional statistics with modern ML and LLM capabilities. The 'pure insights' data scientist role is consolidating into analytics engineering, while the 'build models' data scientist role is merging with ML engineering.
Skills Required
Python, SQL, and statistical modeling are the foundation. Increasingly, roles want experience with LLMs for data analysis, automated insight generation, and building AI-powered data products. Familiarity with cloud data platforms (Snowflake, BigQuery, Databricks) and ML frameworks (scikit-learn, PyTorch) covers most job requirements.
Experimentation design and causal inference are underrated skills that separate strong candidates. Companies care about whether their product changes cause improvements, and can distinguish causation from correlation. A/B testing methodology, Bayesian statistics, and the ability to communicate uncertainty to non-technical stakeholders are high-value skills.
Good postings specify the data stack, the types of problems you'll work on, and the team structure. Look for companies that differentiate between analytics and ML data science. Vague 'data scientist' postings that list every skill under the sun usually mean the company doesn't know what they need.
Compensation Benchmarks
Data Scientist roles pay a median of $192,890 based on 463 positions with disclosed compensation. Mid-level AI roles across all categories have a median of $200,000. This role's midpoint ($260K) sits 35% above the category median. Disclosed range: $210K to $310K.
Across all AI roles, the market median is $217,500. Top-quartile compensation starts at $272,100. The 90th percentile reaches $325,000. For comparison, the highest-paying categories include AI Safety ($300,000) and Research Engineer ($280,000). By seniority level: Entry: $120,000; Mid: $200,000; Senior: $230,000; Director: $272,150; VP: $250,000.
Replit AI Hiring
Replit has 2 open AI roles right now. They're hiring across Data Scientist, AI/ML Engineer. Based in Foster City, CA, US. Compensation range: $275K - $310K.
Location Context
Across all AI roles, 14% (508 positions) offer remote work, while 3,180 require on-site attendance. Top AI hiring metros: New York (1,045 roles, $220,000 median); San Francisco (810 roles, $277,088 median); Los Angeles (397 roles, $215,000 median).
Career Path
Common paths into Data Scientist roles include Data Analyst, Statistician, Quantitative Researcher.
From here, career progression typically leads toward Senior Data Scientist, ML Engineer, AI Product Manager.
Start with statistics and SQL. Build a real analysis project on public data that demonstrates insight generation alongside model building. The market values data scientists who can communicate findings clearly to business stakeholders. If you want to move toward ML engineering, invest in software engineering fundamentals and production deployment skills.
What to Expect in Interviews
Interviews combine statistics, coding, and business acumen. SQL is almost always tested, often with complex joins and window functions. Expect a case study round where you're given a business problem and asked to design an analysis plan. Coding rounds focus on pandas, statistical modeling, and visualization. The strongest differentiator is how well you communicate insights to non-technical stakeholders during presentation rounds.
When evaluating opportunities: Good postings specify the data stack, the types of problems you'll work on, and the team structure. Look for companies that differentiate between analytics and ML data science. Vague 'data scientist' postings that list every skill under the sun usually mean the company doesn't know what they need.
AI Hiring Overview
The AI job market has 3,708 open positions tracked in our dataset. By seniority: 102 entry-level, 1,705 mid-level, 1,469 senior, and 432 leadership roles (Director, VP, C-Level). Remote roles make up 14% of the market (508 positions). The remaining 3,180 roles require on-site or hybrid attendance.
The market median for AI roles is $217,500. Top-quartile compensation starts at $272,100. The 90th percentile reaches $325,000. Highest-paying categories: AI Safety ($300,000 median, 21 roles); Research Engineer ($280,000 median, 147 roles); AI Architect ($254,798 median, 67 roles).
Data Scientist roles remain in high demand, though the definition keeps shifting. Companies increasingly want candidates who can bridge traditional statistics with modern ML and LLM capabilities. The 'pure insights' data scientist role is consolidating into analytics engineering, while the 'build models' data scientist role is merging with ML engineering.
The AI Job Market Today
The AI job market spans 3,708 open positions across 16 role categories. The largest categories by volume: AI/ML Engineer (2,605), Data Scientist (310), AI Software Engineer (259). These three account for the majority of open positions, though smaller categories often have higher per-role compensation because of specialized skill requirements.
The seniority mix tells a story about where AI teams are in their maturity. Entry-level roles (102) are outnumbered by mid-level (1,705) and senior (1,469) positions, reflecting that most companies are past the 'build a team from scratch' phase and need experienced engineers who can ship production systems. Leadership roles (Director, VP, C-Level) total 432 positions, representing the bottleneck between technical execution and organizational strategy.
Remote work availability sits at 14% of all AI roles (508 positions), with 3,180 requiring on-site or hybrid attendance. The remote share has stabilized after the post-pandemic correction. Senior and specialized roles (Research Scientist, ML Architect) are more likely to be remote-eligible than entry-level positions, partly because experienced hires have more negotiating power and partly because these roles require less hands-on mentorship.
AI compensation is structured in clear tiers. The market median sits at $217,500. Top-quartile roles start at $272,100, and the 90th percentile reaches $325,000. These figures include base salary with disclosed compensation. Total compensation (including equity, bonuses, and sign-on) runs 20-40% higher at companies that offer those components.
Category matters for compensation. AI Safety roles lead at $300,000 median, while Prompt Engineer roles sit at $140,000. The spread between highest and lowest-paying categories reflects the premium on specialized technical skills versus broader analytical roles.
The most in-demand skills across all AI postings: Python (1,890 postings), Aws (1,103 postings), Azure (877 postings), Rag (855 postings), Gcp (631 postings), Prompt Engineering (560 postings), Pytorch (545 postings), Claude (498 postings). Python dominates, appearing in the vast majority of role descriptions regardless of category. Cloud platform experience (AWS, GCP, Azure) is the second most common requirement. The newer entrants to the top skills list (RAG, vector databases, LLM APIs) reflect the shift from traditional ML toward generative AI applications.
Frequently Asked Questions
Get Weekly AI Career Intelligence
Salary data, skills demand, and market signals from 16,000+ AI job postings. Every Monday.