Jobiglo

Aucun resultat.

Staff Software Engineer, GPU Inference

cerebras · Toronto

Nouveau
Senior 🇬🇧 English
C++ Python multithreading concurrency vLLM Triton Inference Server TensorRT-LLM PyTorch ROCm HIP Linux containers Kubernetes CI/CD profiling

Description du poste

About the role

Cerebras Systems is seeking a Staff Software Engineer to productionize and optimize its GPU inference stack. You will work on the end‑to‑end serving path that combines GPU‑accelerated prefill with ultra‑fast decode on the Wafer‑Scale Engine, ensuring reliability, performance and numerical correctness.

Key responsibilities

  • Design, build, deploy and maintain the complete GPU prefill path, spanning API services, model‑serving workers, vLLM, PyTorch, ROCm, GPU nodes and rack‑scale infrastructure.
  • Establish operational practices for the AMD GPU fleet, including deployment, upgrades, health‑checking, capacity management and failure recovery.
  • Define and monitor service‑level indicators for GPU‑backed inference, improve fault isolation, automated recovery and incident response.
  • Profile and optimise time‑to‑first‑token, throughput, tail latency, GPU utilisation and memory efficiency under production workloads.
  • Build validation and regression infrastructure for model quality, numerical accuracy and determinism across software and hardware releases.

Required profile

  • 8+ years of software engineering experience with ownership of complex production systems.
  • Hands‑on experience building or operating inference systems for large language or multimodal models on GPUs.
  • Strong communication and technical leadership skills, able to drive cross‑functional projects to completion.

Required skills

  • C++ and Python programming, including multithreading, concurrency and performance‑critical code.
  • Experience with high‑performance model‑serving frameworks such as vLLM, Triton or TensorRT‑LLM.
  • Deep understanding of GPU execution, ROCm/HIP, kernel launches, memory movement and profiling.
  • Proficiency with Linux, containers, Kubernetes, CI/CD pipelines and observability tools.

What we offer

  • Opportunity to build a breakthrough AI platform beyond traditional GPU constraints.
  • Work on one of the world’s fastest AI supercomputers and contribute to open‑source research.
  • Job stability combined with startup vitality and a non‑corporate culture.

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec cerebras.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.
Source : ats:ashby

Pourquoi signalez-vous cette offre ?

Merci pour votre signalement. Nous allons examiner cette offre.

Postulez en 30 secondes

Entrez votre email pour postuler. Un compte sera cree automatiquement.

En continuant, vous acceptez nos conditions d'utilisation.

Deja un compte ? Connexion

💬 Contactez-nous sur Telegram Discuter sur WhatsApp

Publie il y a 11 heures

Expire dans 1 mois

2 vues · 0 interesses

Boostez vos chances

Importez votre CV : nous vous proposons les offres qui matchent votre profil.

Analyse de votre CV en cours...

cerebras

Toronto