Jobiglo

Aucun resultat.

Software Engineer, GPU Inference

cerebras · United States and Canada

Nouveau
Senior 🇬🇧 English
C++ Python Multithreading Concurrency vLLM AMD ROCm HIP RCCL rocprofiler Linux Containers Kubernetes CI/CD

Description du poste

About the role

Cerebras is building a new generation of disaggregated AI inference systems that combine GPU‑accelerated prefill with ultra‑fast decode on its Wafer‑Scale Engine. We are looking for a Software Engineer to productionize and optimise the GPU serving stack, ensuring reliability, numerical correctness, observability and top‑tier performance.

Key responsibilities

  • Productionize and maintain the GPU inference stack, including API services, model‑serving workers, vLLM, PyTorch, ROCm and rack‑scale GPU infrastructure.
  • Establish operational practices for the AMD GPU fleet such as deployment, upgrades, health‑checking, capacity management and automated recovery.
  • Define and monitor service‑level indicators, improve fault isolation, graceful degradation and incident response for GPU‑backed inference.
  • Profile and optimise time‑to‑first‑token, throughput, tail latency, GPU utilisation and memory efficiency under production workloads.
  • Build validation and regression infrastructure to ensure numerical correctness and model quality across software and hardware releases.

Required profile

  • 5+ years of software engineering experience with ownership of complex production systems.
  • Hands‑on experience building or optimizing inference systems for large language or multimodal models on GPUs.
  • Strong C++ and Python programming skills, including multithreading, concurrency and performance‑critical code.
  • Experience with high‑performance model‑serving frameworks such as vLLM, Triton or similar.
  • Deep understanding of GPU execution, memory movement, kernel launches and profiling methodologies.
  • Proficiency with Linux, containers, Kubernetes (or comparable orchestration), CI/CD and latency‑sensitive services.

Required skills

  • C++
  • Python
  • Multithreading and concurrency
  • vLLM or equivalent model‑serving framework
  • AMD ROCm ecosystem (HIP, RCCL, rocprofiler)
  • Linux, containers and Kubernetes
  • Performance profiling and GPU optimisation
  • Distributed systems debugging

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec cerebras.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.
Source : ats:ashby

Pourquoi signalez-vous cette offre ?

Merci pour votre signalement. Nous allons examiner cette offre.

Postulez en 30 secondes

Entrez votre email pour postuler. Un compte sera cree automatiquement.

En continuant, vous acceptez nos conditions d'utilisation.

Deja un compte ? Connexion

💬 Contactez-nous sur Telegram Discuter sur WhatsApp

Publie il y a 6 heures

Expire dans 1 mois

4 vues · 0 interesses

Boostez vos chances

Importez votre CV : nous vous proposons les offres qui matchent votre profil.

Analyse de votre CV en cours...

cerebras

United States and Canada