Jobiglo

No results.

Principal Architect – Hardware Efficient AI Foundation Model Training

huaweicanada · Markham

New
Permanent 🇬🇧 English
Distributed training and inference techniques Memory hierarchy and interconnect technologies

Job description

About the role

Huawei Canada’s Computing Data Application Acceleration Lab is seeking a Principal Architect to lead hardware‑efficient AI foundation model training. The role focuses on full‑stack innovations, software‑hardware co‑design, and optimizing data efficiency across storage and runtime layers.

Key responsibilities

  • Collaborate with internal and external partners to design foundational model architectures for LLM, code, and multimodal sub‑fields, driving breakthroughs in post‑training and continual training.
  • Define technical requirements for large‑scale distributed training and inference infrastructures, including parallelization strategies and operator fusion.
  • Analyze computational characteristics of emerging AI architectures to ensure accuracy, performance, and hardware evolution.

Required profile

  • Proven experience training and optimizing cutting‑edge AI models at scale (10B+ parameters).
  • Deep knowledge of modern AI architectures such as long‑sequence models, reinforcement learning, multimodal systems, and autonomous agents.
  • Strong understanding of AI algorithm mechanisms and their implementation.
  • Hands‑on expertise with AI frameworks (e.g., PyTorch, vLLM, SGLang) and mainstream distributed training/inference techniques.
  • Familiarity with AI chip architectures (GPU, NPU, TPU) and memory hierarchy/interconnect technologies.
  • PhD in AI architecture, computer architecture, or a related field is preferred.
  • Solid publication record in AI systems or chip design is an asset.

Required skills

  • Training and optimizing large AI models (10B+ parameters).
  • AI architectures: long‑sequence, reinforcement learning, multimodal, agents.
  • AI frameworks: PyTorch, vLLM, SGLang.
  • Distributed training and inference techniques.
  • AI chip architectures: GPU, NPU, TPU.
  • Memory hierarchy and interconnect technologies.

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec huaweicanada.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.
Le contrat proposé est un Permanent basé à Markham.
Source : ats:recruitee

Why are you reporting this job?

Thank you for your report. We will review this job.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

Apply now →

By continuing, you accept our terms of use.

Already have an account? Login

A question about this job?

Ask it here: you will get the full job summary by e-mail, right away.

💬 Chat with us on Telegram

Published 21 hours ago

Expires 1 month from now

3 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

huaweicanada

Markham