Jobiglo

No results.

Senior/Staff ML Ops Engineer

waabi · Remote US & Canada

New Remote
Remote Senior 🇬🇧 English
Python Kubernetes GPU scheduling Helm AWS Terraform Pulumi PyTorch DDP FSDP Containers CI/CD Monorepo

Job description

About the role

You will design, build, and evolve Waabi's training infrastructure on Kubernetes, enabling reliable, large‑scale autonomous‑vehicle model training. The role focuses on creating developer‑friendly tools, shortening the training loop, and ensuring observability across the ML stack.

Key responsibilities

  • Develop GPU‑aware scheduling, autoscaling, and multi‑node distributed jobs on Kubernetes.
  • Design CLIs, SDKs, and job‑submission templates that make common workflows a single command.
  • Measure and reduce time‑to‑first‑run, edit‑to‑signal latency, and failure detection.
  • Evaluate and adopt best‑in‑class tooling, providing prototypes and migration paths.
  • Implement dataset versioning, sharding, and high‑throughput loading for multimodal sensor data.
  • Turn ad‑hoc Python scripts into robust, documented, observable libraries and services.
  • Build experiment dashboards, model registries, and lineage tracking from dataset to simulation results.
  • Ship CI/CD pipelines for models alongside code, ensuring consistent validation.
  • Create observability metrics for utilization, throughput, failure modes, and cost per experiment.
  • Provide documentation, onboarding guides, and office‑hour support for researchers.

Required profile

  • 5+ years of software or infrastructure engineering experience, preferably with ML or data‑intensive systems.
  • Hands‑on expertise with Kubernetes (GPU scheduling, autoscaling, Helm, networking).
  • Strong Python skills and experience designing user‑friendly APIs and CLIs.
  • Deep knowledge of AWS services (object storage, IAM, GPU compute, cost management) and IaC tools (Terraform, Pulumi).
  • Experience with distributed PyTorch training (DDP, FSDP) and model‑registry tooling.
  • Proficiency with containers, CI/CD, and large monorepos.
  • Ability to influence without authority and collaborate across teams.

Required skills

  • Python
  • Kubernetes
  • GPU scheduling & autoscaling
  • Helm
  • AWS (S3, IAM, EC2 GPU)
  • Terraform / Pulumi
  • PyTorch (DDP, FSDP)
  • Containers & CI/CD
  • Large monorepo management

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec waabi.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.
Source : ats:lever

Why are you reporting this job?

Thank you for your report. We will review this job.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

By continuing, you accept our terms of use.

Already have an account? Login

💬 Chat with us on Telegram Chat on WhatsApp

Published 16 hours ago

Expires 1 month from now

2 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

waabi

Remote US & Canada