Jobiglo

Aucun resultat.

Senior/Staff ML Ops Engineer

waabi · Remote US & Canada

Nouveau Remote
Remote Senior 🇬🇧 English
Python Kubernetes GPU scheduling Helm AWS Terraform Pulumi PyTorch DDP FSDP Containers CI/CD Monorepo

Description du poste

About the role

You will design, build, and evolve Waabi's training infrastructure on Kubernetes, enabling reliable, large‑scale autonomous‑vehicle model training. The role focuses on creating developer‑friendly tools, shortening the training loop, and ensuring observability across the ML stack.

Key responsibilities

  • Develop GPU‑aware scheduling, autoscaling, and multi‑node distributed jobs on Kubernetes.
  • Design CLIs, SDKs, and job‑submission templates that make common workflows a single command.
  • Measure and reduce time‑to‑first‑run, edit‑to‑signal latency, and failure detection.
  • Evaluate and adopt best‑in‑class tooling, providing prototypes and migration paths.
  • Implement dataset versioning, sharding, and high‑throughput loading for multimodal sensor data.
  • Turn ad‑hoc Python scripts into robust, documented, observable libraries and services.
  • Build experiment dashboards, model registries, and lineage tracking from dataset to simulation results.
  • Ship CI/CD pipelines for models alongside code, ensuring consistent validation.
  • Create observability metrics for utilization, throughput, failure modes, and cost per experiment.
  • Provide documentation, onboarding guides, and office‑hour support for researchers.

Required profile

  • 5+ years of software or infrastructure engineering experience, preferably with ML or data‑intensive systems.
  • Hands‑on expertise with Kubernetes (GPU scheduling, autoscaling, Helm, networking).
  • Strong Python skills and experience designing user‑friendly APIs and CLIs.
  • Deep knowledge of AWS services (object storage, IAM, GPU compute, cost management) and IaC tools (Terraform, Pulumi).
  • Experience with distributed PyTorch training (DDP, FSDP) and model‑registry tooling.
  • Proficiency with containers, CI/CD, and large monorepos.
  • Ability to influence without authority and collaborate across teams.

Required skills

  • Python
  • Kubernetes
  • GPU scheduling & autoscaling
  • Helm
  • AWS (S3, IAM, EC2 GPU)
  • Terraform / Pulumi
  • PyTorch (DDP, FSDP)
  • Containers & CI/CD
  • Large monorepo management

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec waabi.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.
Source : ats:lever

Pourquoi signalez-vous cette offre ?

Merci pour votre signalement. Nous allons examiner cette offre.

Postulez en 30 secondes

Entrez votre email pour postuler. Un compte sera cree automatiquement.

En continuant, vous acceptez nos conditions d'utilisation.

Deja un compte ? Connexion

💬 Contactez-nous sur Telegram Discuter sur WhatsApp

Publie il y a 19 heures

Expire dans 1 mois

3 vues · 0 interesses

Boostez vos chances

Importez votre CV : nous vous proposons les offres qui matchent votre profil.

Analyse de votre CV en cours...

waabi

Remote US & Canada