Interested in this MLOps Engineer role at TetraScience?
Apply Now →Skills & Technologies
About This Role
Who We Are
TetraScience is the Scientific Data and AI Cloud company. We are catalyzing the Scientific AI revolution by designing and industrializing AI\-native scientific data sets, which we bring to life in a growing suite of next gen lab data management solutions, scientific use cases, and AI\-enabled outcomes.
TetraScience is the category leader in this vital new market, generating more revenue than all other companies in the aggregate. In the last year alone, the world's dominant players in compute, cloud, data, and AI infrastructure have converged on TetraScience as the de facto standard, entering into co\-innovation and go\-to\-market partnerships.
In connection with your candidacy, you will be asked to carefully review the Tetra Way letter, authored directly by Patrick Grady, our co\-founder and CEO. This letter is designed to assist you in better understanding whether TetraScience is the right fit for you from a values and ethos perspective.
It is impossible to overstate the importance of this document and you are encouraged to take it literally and reflect on whether you are aligned with our unique approach to company and team building. If you join us, you will be expected to embody its contents each day.
The Role
We're looking for a Lead Software Platform Engineer working at the intersection of distributed systems and MLOps. You will own and scale our AI and data infrastructure as a product our customers build their science on, and that every other engineering team builds against. You will architect the cloud\-based services and MLOps infrastructure that enable production\-grade AI/ML workflows, working closely with Applied AI engineers, data engineers, and platform teams, and you will act as the technical design authority for how models, LLMs, and agents run in production. The work spans the model and prompt lifecycle, the evaluation and observability harness, the security and tenant boundaries around prompts and retrieval, and the cost and latency controls that keep production AI economically viable at scale.
This is a highly impactful seat for an engineer who has shipped AI/ML infrastructure as a multi\-tenant product rather than internal tooling. The model serving, MCP, and agent capabilities you build are consumed directly by scientists at the world's largest pharmaceutical companies, running against their own data under their compliance obligations. If that scope appeals to you, and you thrive on turning ambitious scalability and cost targets into concrete technical strategy inside a regulated environment, we'd love to talk to you.
What You Will Do* Own the technical architecture of the AI/ML platform: the service and API surface our customers use to run models and agents against their own scientific data, and that Applied AI and data engineering teams build against internally.
- Own the end\-to\-end model and prompt lifecycle across Databricks MLflow and AWS Bedrock, including registration, versioning, asset bundles, staged promotion, rollback, and multi\-model serving.
- Design the inference substrate for both real\-time and batch AI workloads, including routing, batching, caching, concurrency control, GPU and accelerator capacity planning, handling of large binary inputs such as instrument images, and graceful degradation under load.
- Integrate AI models and large language models (LLMs) into production systems using architectures like retrieval\-augmented generation (RAG), and architect the agentic layer above them: tool and function calling, MCP\-based tooling, and agent runtimes, deciding what belongs in the platform versus in the applications built on top of it.
- Design security into the AI platform rather than leaving it to the applications above it, partnering with our security team on guardrails, prompt\-injection and tool\-abuse defenses, PII and PHI handling, and hard tenant data boundaries across prompts, retrieval, and tool calls.
- Build the evaluation and quality infrastructure that makes AI shippable: offline and online eval harnesses, golden datasets, regression gates in CI, A/B and shadow deployment, and drift and hallucination detection in production.
- Establish observability for the AI platform, including monitoring, alerting, logging, and distributed tracing, and set the SLI, SLO, and SLA model for systems whose outputs are probabilistic.
- Design for reproducibility and lineage required in a validated pharma environment, with versioned data, code, prompts, and model artifacts, and an audit trail that can withstand customer and regulatory scrutiny.
- Contribute to the infrastructure\-as\-code and deployment automation for the AI platform (CloudFormation, AWS CDK), partnering with the team that owns production deployments to support multi\-tenant infrastructure, online upgrades, and on\-demand compute allocation.
- Own production readiness for the AI platform with Applied AI engineers, data engineers, and platform teams: the performance, reliability, and cost\-efficiency of models in production, plus incident response and runbooks.
- Act as SME and design authority across product and engineering. Lead design reviews, write the reference architectures and technical documentation others follow, and mentor senior and mid\-level engineers on distributed systems and AI engineering practice.
- Set technical direction on emerging AI infrastructure: evaluate new frameworks, serving runtimes, model providers, and data types, and make clear build\-versus\-buy decisions grounded in cost, risk, and scalability.
Requirements
- + 10\+ years of professional experience in software engineering and infrastructure engineering, with a proven track record of designing, building, and scaling distributed, cloud\-native systems in production.
+ Demonstrated experience as a technical leader or architect, accountable for the key decisions on system design, scalability, performance, and cost optimization.
+ Experience designing security into a multi\-tenant platform, including authorization boundaries between tenants and handling of sensitive data such as PII and PHI, with awareness of LLM\-specific risks like prompt injection and tool abuse.
+ Extensive experience building and maintaining AI/ML infrastructure in production, including model deployment and lifecycle management, delivered as a multi\-tenant product with external users rather than internal tooling for a single team. Candidates whose experience is limited to building pipelines for their own team will not be a fit.
+ Deep, hands\-on experience taking LLM\-based systems to production, including RAG architecture, retrieval and embedding design, prompt and model versioning, and tool or function calling. Not just prototyping with an SDK. We are looking for someone who has operated these systems under real traffic, real latency budgets, and real failure modes.
+ Expert\-level coding skills in TypeScript and Python building robust APIs and backend services, with the judgment to critically evaluate AI\-generated code for correctness, security, and architectural fit.
+ Production\-level experience with a model registry and serving stack, ideally Databricks MLflow, including model registration, versioning, asset bundles, and serving workflows.
+ Experience treating AI evaluation as a release gate rather than post\-hoc reporting, including eval harnesses, regression gates, and drift or quality monitoring for non\-deterministic systems.
+ Proficiency in API\-first design, including REST and OpenAPI specifications, designing APIs that are scalable, secure, versioned, and extensible.
+ Solid working knowledge of AWS and containerized workloads (e.g., Docker), and familiarity with infrastructure\-as\-code such as CloudFormation or CDK, with the ability to contribute to CI/CD pipelines and deployment automation.
+ Experience defining observability and SLI/SLO/SLA practice for production systems, including monitoring, alerting, and distributed tracing.
+ Ability to articulate ideas clearly to customers and cross\-functional teams, influence technical direction on teams you do not manage, and mentor other engineers.
Nice to Have* + Familiarity with emerging LLM frameworks for advanced prompt orchestration and programmatic LLM pipelines.
+ Experience running agentic or multi\-step orchestration in production, and with MCP as an integration surface for tooling and non\-human identities.
+ Understanding of LLM cost monitoring, latency optimization, and usage analytics in production environments, including per\-tenant cost attribution for tokens, GPU, and inference.
+ Experience with multimodal model inputs, including image and instrument data, and the throughput and cost implications of serving them at scale.
+ Experience with fine\-tuning, distillation, or model optimization such as quantization, batching, or KV\-cache strategy, to improve latency and cost.
+ Experience delivering ML or AI systems in a regulated or validated environment (GxP, 21 CFR Part 11, SOC 2\), including computer system validation and audit readiness.
+ Background in scientific, life sciences, or laboratory data domains.
Benefits
- 100% employer\-paid benefits for all eligible employees and immediate family members
- Unlimited paid time off (PTO)
- 401K
- Flexible working arrangements \- Remote work
- Company paid Life Insurance, LTD/STD
- A culture of continuous improvement where you can grow your career and get coaching
- The salary range for this position is $200K\-$270K USD. The salary range posted reflects our target baseline for this role. Final compensation is determined by a thorough evaluation of factors including the candidate’s specific experience, localized market data, and internal team equity.
We are not currently providing visa sponsorship for this position.
Salary Context
This $200K-$270K range is above the 75th percentile for MLOps Engineer roles in our dataset (median: $168K across 34 roles with salary data).
View full MLOps Engineer salary data →Role Details
About This Role
MLOps Engineers build the infrastructure that keeps ML models running in production. They own CI/CD pipelines for model deployment, monitoring for data drift and model degradation, and the tooling that lets data scientists ship faster. If ML Engineers build the models, MLOps Engineers build the roads those models travel on.
The job is fundamentally about reliability and velocity. Data scientists want to iterate fast. Product teams want stable predictions. Your job is to make both happen simultaneously. That means building deployment pipelines that catch regressions before they hit production, monitoring systems that alert on data drift before it degrades model performance, and self-service tooling that lets data scientists deploy without filing a ticket.
Across the 4,317 AI roles we're tracking, MLOps Engineer positions make up 1% of the market. At TetraScience, this role fits into their broader AI and engineering organization.
MLOps demand tracks closely with production ML adoption. As more companies move models from notebooks to production, the need for MLOps grows. The role is well-established at large tech companies and growing fast at mid-stage startups that are hitting the 'our models work in notebooks but break in production' phase.
What the Work Looks Like
A typical week involves: debugging a model deployment that's serving stale predictions, building a new monitoring dashboard for a feature team, writing Terraform for GPU-enabled inference clusters, reviewing pull requests for the ML platform's CI/CD pipeline, and meeting with data scientists to understand their pain points. You're the bridge between ML and infrastructure.
MLOps demand tracks closely with production ML adoption. As more companies move models from notebooks to production, the need for MLOps grows. The role is well-established at large tech companies and growing fast at mid-stage startups that are hitting the 'our models work in notebooks but break in production' phase.
Skills Required
Kubernetes, Docker, and cloud infrastructure are baseline. Most roles want experience with ML-specific tooling: MLflow, Kubeflow, Weights & Biases, or similar. Strong DevOps fundamentals matter more than ML theory. You need to understand model serving (TorchServe, Triton, vLLM), monitoring (Prometheus, Grafana), and infrastructure-as-code (Terraform, Pulumi).
GPU infrastructure knowledge is increasingly valuable as LLM inference becomes a major cost center. Understanding GPU scheduling, multi-node training setups, and inference optimization (quantization, batching, caching) puts you in the top tier. Experience with model registries and feature stores rounds out the profile.
Good MLOps postings specify their ML stack, infrastructure scale, and the problems they're solving (deployment velocity, cost optimization, monitoring gaps). Red flag: companies that want MLOps but don't have any models in production yet. You'll end up doing general DevOps instead.
Compensation Benchmarks
MLOps Engineer roles pay a median of $203,000 based on 85 positions with disclosed compensation. Senior-level AI roles across all categories have a median of $227,400. This role's midpoint ($235K) sits 16% above the category median. Disclosed range: $200K to $270K.
Across all AI roles, the market median is $215,000. Top-quartile compensation starts at $266,300. The 90th percentile reaches $320,790. For comparison, the highest-paying categories include AI Safety ($287,500) and Research Engineer ($272,100). By seniority level: Entry: $110,000; Mid: $194,400; Senior: $227,400; Director: $274,554; VP: $241,000.
TetraScience AI Hiring
TetraScience has 1 open AI role right now. They're hiring across MLOps Engineer. Based in US. Compensation range: $270K - $270K.
Location Context
AI roles in Austin pay a median of $214,343 across 143 tracked positions.
Career Path
Common paths into MLOps Engineer roles include DevOps Engineer, Platform Engineer, Data Engineer.
From here, career progression typically leads toward ML Platform Lead, Infrastructure Architect, Engineering Manager.
DevOps engineers with ML curiosity have the shortest path. You already understand deployment, monitoring, and infrastructure. Add ML-specific knowledge (model serving, data pipelines, experiment tracking) and you're competitive. The career ceiling is high: ML Platform Lead roles at top companies pay well because the infrastructure complexity is enormous.
What to Expect in Interviews
Interviews emphasize infrastructure and reliability. Expect questions about CI/CD for ML models, monitoring for data drift, and how you'd design a model serving platform that handles 10K requests per second. Coding rounds focus on Python and infrastructure-as-code (Terraform, Helm). Be ready to discuss tradeoffs between different model serving frameworks and how you'd handle rollback when a new model degrades performance.
When evaluating opportunities: Good MLOps postings specify their ML stack, infrastructure scale, and the problems they're solving (deployment velocity, cost optimization, monitoring gaps). Red flag: companies that want MLOps but don't have any models in production yet. You'll end up doing general DevOps instead.
AI Hiring Overview
The AI job market has 4,317 open positions tracked in our dataset. By seniority: 138 entry-level, 2,071 mid-level, 1,655 senior, and 453 leadership roles (Director, VP, C-Level). Remote roles make up 15% of the market (635 positions). The remaining 3,657 roles require on-site or hybrid attendance.
The market median for AI roles is $215,000. Top-quartile compensation starts at $266,300. The 90th percentile reaches $320,790. Highest-paying categories: AI Safety ($287,500 median, 34 roles); Research Engineer ($272,100 median, 227 roles); AI Engineering Manager ($244,000 median, 23 roles).
MLOps demand tracks closely with production ML adoption. As more companies move models from notebooks to production, the need for MLOps grows. The role is well-established at large tech companies and growing fast at mid-stage startups that are hitting the 'our models work in notebooks but break in production' phase.
The AI Job Market Today
The AI job market spans 4,317 open positions across 15 role categories. The largest categories by volume: AI/ML Engineer (3,004), Data Scientist (345), AI Software Engineer (309). These three account for the majority of open positions, though smaller categories often have higher per-role compensation because of specialized skill requirements.
The seniority mix tells a story about where AI teams are in their maturity. Entry-level roles (138) are outnumbered by mid-level (2,071) and senior (1,655) positions, reflecting that most companies are past the 'build a team from scratch' phase and need experienced engineers who can ship production systems. Leadership roles (Director, VP, C-Level) total 453 positions, representing the bottleneck between technical execution and organizational strategy.
Remote work availability sits at 15% of all AI roles (635 positions), with 3,657 requiring on-site or hybrid attendance. The remote share has stabilized after the post-pandemic correction. Senior and specialized roles (Research Scientist, ML Architect) are more likely to be remote-eligible than entry-level positions, partly because experienced hires have more negotiating power and partly because these roles require less hands-on mentorship.
AI compensation is structured in clear tiers. The market median sits at $215,000. Top-quartile roles start at $266,300, and the 90th percentile reaches $320,790. These figures include base salary with disclosed compensation. Total compensation (including equity, bonuses, and sign-on) runs 20-40% higher at companies that offer those components.
Category matters for compensation. AI Safety roles lead at $287,500 median, while Prompt Engineer roles sit at $145,000. The spread between highest and lowest-paying categories reflects the premium on specialized technical skills versus broader analytical roles.
The most in-demand skills across all AI postings: Python (2,249 postings), Aws (1,224 postings), Azure (938 postings), Rag (915 postings), Gcp (660 postings), Pytorch (640 postings), Prompt Engineering (624 postings), Kubernetes (559 postings). Python dominates, appearing in the vast majority of role descriptions regardless of category. Cloud platform experience (AWS, GCP, Azure) is the second most common requirement. The newer entrants to the top skills list (RAG, vector databases, LLM APIs) reflect the shift from traditional ML toward generative AI applications.
Frequently Asked Questions
Get Weekly AI Career Intelligence
Salary data, skills demand, and market signals from 16,000+ AI job postings. Every Monday.