Distributed Software Engineer
cerebras · Toronto
Job description
About the role
Cerebras Systems builds the world’s largest AI chip and delivers industry‑leading training and inference speeds. The Cluster engineering team owns the software that turns thousands of wafers, servers, and switches into a cloud that stays up, stays busy, and stays debuggable. You will stand clusters up from bare metal, schedule training and inference workloads, keep the fleet healthy, and make it observable to users, operators, and AI agents.
Key responsibilities
- Declarative, CRD‑driven automation of bare‑metal networking, OS, and application software across clusters of Cerebras systems, servers, and switches.
- Push‑button cluster install, upgrade, and security patching with real downtime budgets, gated by canaries.
- Kubernetes operators that schedule large inference workloads, handling resource locks, priority queues, network topology, and health‑aware placement.
- Development of gRPC control‑plane services, authorization, admission webhooks, and quota policy for a multi‑tenant fleet.
- Building metrics and log pipelines with purpose‑built exporters for wafer‑scale systems, servers (Redfish, IPMI), and network fabric (gNMI, sFlow) on Prometheus and Grafana, with SLOs and alerting.
Required profile
- 5+ years building and operating production distributed systems or infrastructure software.
- Strong production‑level Go and Python programming experience.
- Deep knowledge of Kubernetes controllers and operators, including CRDs, reconciliation, informers, admission webhooks, and RBAC.
- Excellent debugging skills across distributed systems, Linux, and networking.
- Self‑driving ability to create context, decide, and drive work across team boundaries in a fast‑moving, partially undocumented environment.
Required skills
- Go
- Python
- Kubernetes
- gRPC
- Prometheus
- Grafana
- PromQL
- Redfish
- IPMI
- gNMI
- sFlow
What we offer
- Build a breakthrough AI platform beyond the constraints of the GPU.
- Publish and open‑source cutting‑edge AI research.
- Work on one of the fastest AI supercomputers in the world.
- Enjoy job stability with startup vitality.
- Simple, non‑corporate work culture that respects individual beliefs.
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in Canada.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 1 day ago
Expires 1 month from now
4 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
cerebras
Toronto