Compute Server Platform Architect
cerebras · United States and Canada
Description du poste
About the role
As a Compute/Server Platform Architect on the Cluster Architecture Team, you will own the server‑side platform architecture that enables Cerebras CS‑3 AI clusters to deliver predictable performance, scalability and reliability for training and inference workloads.
Key responsibilities
- Own the architecture for all server roles in Cerebras clusters, defining server types, configurations and lifecycle strategy.
- Define and maintain server formulas, capacity planning and headroom policies.
- Specify platform configurations including CPU SKU, memory topology, PCIe topology, NIC selection and local NVMe policy.
- Translate software and runtime flows into measurable hardware requirements and communicate guardrails to software teams.
- Develop performance and scaling models, validate with micro‑benchmarks and identify bottlenecks.
- Define OS, BIOS, firmware and driver baselines for each server type.
- Stay current on emerging server technologies and run proof‑of‑concept evaluations.
- Lead technical vendor engagements, influence roadmaps and drive joint debugging.
- Define qualification and acceptance criteria and partner with TPMs for production rollout.
- Support bring‑up and deployment debugging, driving root‑cause analysis across firmware, drivers, OS and runtime.
Required profile
- PhD in Computer Science or Electrical/Computer Engineering with 8+ years industry experience, or Master’s/Bachelor’s with 10+ years.
- 5+ years experience in server platform architecture, systems performance engineering or large‑scale infrastructure design for AI/ML, HPC or performance‑sensitive distributed systems.
- Deep understanding of x86 server architecture, CPU microarchitecture, cache hierarchies, NUMA and memory bandwidth/latency trade‑offs.
- Strong Linux systems knowledge, including profiling, performance analysis and tuning.
- Experience with high‑performance I/O paths, NIC behavior, RDMA/RoCE and NVMe characteristics.
- Proven ability to create and validate capacity and performance models with rigorous benchmarking.
- Experience working directly with vendors to evaluate platforms and influence roadmaps.
- Excellent cross‑functional communication and ability to drive technical decisions.
Required skills
- x86 server architecture
- Linux systems
- CPU microarchitecture
- NUMA
- PCIe
- NVMe
- RDMA / RoCE
- C, C++, Python
What we offer
- Opportunity to work on a breakthrough AI platform beyond GPU constraints.
- Publish and open‑source cutting‑edge AI research.
- Work on one of the fastest AI supercomputers in the world.
- Job stability with startup vitality and a non‑corporate culture.
Questions fréquentes
Pourquoi signalez-vous cette offre ?
Aller plus loin
Salaires, guides et recherches au Canada.
Postulez en 30 secondes
Entrez votre email pour postuler. Un compte sera cree automatiquement.
En continuant, vous acceptez nos conditions d'utilisation.
Deja un compte ? Connexion
Publie il y a 14 heures
Expire dans 1 mois
4 vues · 0 interesses
Boostez vos chances
Importez votre CV : nous vous proposons les offres qui matchent votre profil.
Analyse de votre CV en cours...
cerebras
United States and Canada
Offres similaires
-
Principal AI Security Engineer
cerebras United States and Canada -
Software Engineer, Kernel Reliability
cerebras United States and Canada -
Technical Lead Manager, Infrastructure Hardware (Server and Network Systems)
cerebras United States and Canada -
Senior Backend Engineer (Remote, Canada)
shortstory Canada (Remote) -
Technical Support Engineer
sentry Toronto