Jobiglo

Aucun resultat.

Engineering Manager, Kernel Reliability

cerebras · United States and Canada

Nouveau
Senior 🇬🇧 English
distributed programming message passing debuggers core dump handling code sanitizers diagnostic tool development incident response computer architecture

Description du poste

About the role

We are looking for a deeply technical, hands‑on engineering leader for our on‑field Kernel Reliability team. You will lead a high‑performing team to improve the reliability of advanced compute clusters and the underlying inference, training, and internal production services.

Key responsibilities

  • Provide hands‑on technical leadership, owning the technical vision and roadmap for kernel‑centric reliability of internal and customer‑facing systems.
  • Assist System and Cluster Operations teams in reducing downtime after failures by delivering tooling and manual interventions for failure analysis and diagnostics.
  • Collaborate with the Debug Team to enhance debug tools, accelerating failure analysis.
  • Work with software teams to improve the software stack, including kernels, for better on‑field debugging and failure analysis.
  • Partner with ASIC and hardware architecture teams to co‑design next‑generation architectures with reliability and debug‑friendliness in mind.
  • Lead, mentor, and grow a high‑caliber engineering team, fostering a culture of technical excellence and rapid execution.

Required profile

  • 6+ years of software engineering experience, including 3+ years leading teams in SW/HW reliability, debug, diagnostic, or failure‑analysis roles.
  • Strong background in operations and monitoring, with experience in incident response and post‑mortem analysis.
  • Demonstrated ability to recruit, retain, and mentor high‑performing engineers and collaborate cross‑functionally to deliver customer‑facing products.

Required skills

  • Parallel and distributed programming (message passing, multicore, GPU, embedded).
  • Debuggers, core‑dump handling, code sanitizers, and diagnostic tool development.
  • Experience debugging distributed and parallel applications (deadlocks, livelocks, race conditions).
  • Deep understanding of computer architectures (instruction pipelining, multithreading, networking).
  • Monitoring and reliability engineering (incident response, post‑mortem analysis).

What we offer

  • Opportunity to build a breakthrough AI platform beyond GPU constraints.
  • Publish and open‑source cutting‑edge AI research.
  • Work on one of the fastest AI supercomputers in the world.
  • Job stability combined with startup vitality.
  • A simple, non‑corporate culture that respects individual beliefs.

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec cerebras.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.
Source : ats:ashby

Pourquoi signalez-vous cette offre ?

Merci pour votre signalement. Nous allons examiner cette offre.

Postulez en 30 secondes

Entrez votre email pour postuler. Un compte sera cree automatiquement.

En continuant, vous acceptez nos conditions d'utilisation.

Deja un compte ? Connexion

💬 Contactez-nous sur Telegram Discuter sur WhatsApp

Publie il y a 7 heures

Expire dans 1 mois

3 vues · 0 interesses

Boostez vos chances

Importez votre CV : nous vous proposons les offres qui matchent votre profil.

Analyse de votre CV en cours...

cerebras

United States and Canada