Compute Server Platform Architect
cerebras · United States and Canada
Job description
About the role
As a Compute/Server Platform Architect on the Cluster Architecture Team, you will own the server‑side platform architecture that enables Cerebras CS‑3 AI clusters to deliver predictable performance, scalability and reliability for training and inference workloads.
Key responsibilities
- Own the architecture for all server roles in Cerebras clusters, defining server types, configurations and lifecycle strategy.
- Define and maintain server formulas, capacity planning and headroom policies.
- Specify platform configurations including CPU SKU, memory topology, PCIe topology, NIC selection and local NVMe policy.
- Translate software and runtime flows into measurable hardware requirements and communicate guardrails to software teams.
- Develop performance and scaling models, validate with micro‑benchmarks and identify bottlenecks.
- Define OS, BIOS, firmware and driver baselines for each server type.
- Stay current on emerging server technologies and run proof‑of‑concept evaluations.
- Lead technical vendor engagements, influence roadmaps and drive joint debugging.
- Define qualification and acceptance criteria and partner with TPMs for production rollout.
- Support bring‑up and deployment debugging, driving root‑cause analysis across firmware, drivers, OS and runtime.
Required profile
- PhD in Computer Science or Electrical/Computer Engineering with 8+ years industry experience, or Master’s/Bachelor’s with 10+ years.
- 5+ years experience in server platform architecture, systems performance engineering or large‑scale infrastructure design for AI/ML, HPC or performance‑sensitive distributed systems.
- Deep understanding of x86 server architecture, CPU microarchitecture, cache hierarchies, NUMA and memory bandwidth/latency trade‑offs.
- Strong Linux systems knowledge, including profiling, performance analysis and tuning.
- Experience with high‑performance I/O paths, NIC behavior, RDMA/RoCE and NVMe characteristics.
- Proven ability to create and validate capacity and performance models with rigorous benchmarking.
- Experience working directly with vendors to evaluate platforms and influence roadmaps.
- Excellent cross‑functional communication and ability to drive technical decisions.
Required skills
- x86 server architecture
- Linux systems
- CPU microarchitecture
- NUMA
- PCIe
- NVMe
- RDMA / RoCE
- C, C++, Python
What we offer
- Opportunity to work on a breakthrough AI platform beyond GPU constraints.
- Publish and open‑source cutting‑edge AI research.
- Work on one of the fastest AI supercomputers in the world.
- Job stability with startup vitality and a non‑corporate culture.
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in Canada.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 6 hours ago
Expires 1 month from now
3 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
cerebras
United States and Canada
Related job offers
-
Principal AI Security Engineer
cerebras United States and Canada -
Software Engineer, Kernel Reliability
cerebras United States and Canada -
Technical Lead Manager, Infrastructure Hardware (Server and Network Systems)
cerebras United States and Canada -
Senior Backend Engineer (Remote, Canada)
shortstory Canada (Remote) -
Technical Support Engineer
sentry Toronto