Jobiglo

No results.

This job is no longer available

This job expired on 13/09/2026. It no longer accepts applications.

Site Reliability Engineer (SRE) – Montreal

Keasis · Montréal

Senior 🇬🇧 English
Python Bash Kubernetes Docker Jenkins GitHub Actions GitLab CI Azure DevOps Terraform Ansible DNS TCP/IP Load balancers SSL Prometheus Grafana Splunk ELK Stack Datadog AppDynamics Logging Tracing Metrics collection Incident management Root cause analysis Problem management High availability Disaster recovery Capacity planning SQL Git

Job description

About the role

We are looking for a seasoned Site Reliability Engineer to join our Montreal technology team. The role focuses on enhancing the reliability, scalability, performance and availability of our enterprise applications and infrastructure.

Key responsibilities

  • Design, build, and maintain highly available and scalable production systems.
  • Implement SRE best practices such as SLIs, SLOs, and error budgets.
  • Automate operational tasks using scripting and Infrastructure as Code.
  • Monitor application and infrastructure health with modern observability tools.
  • Troubleshoot production incidents, conduct root cause analysis, and apply preventive measures.
  • Collaborate with development, infrastructure, and security teams to improve system reliability.
  • Develop CI/CD pipelines to enable efficient and reliable software delivery.
  • Optimize application performance, system capacity, and resource utilization.
  • Create operational runbooks and disaster recovery procedures.
  • Participate in on‑call support and production incident management.

Required profile

  • 7+ years of experience in Site Reliability Engineering, DevOps, or production support.
  • Strong Linux/Unix system administration background.
  • Hands‑on experience with major cloud platforms (AWS, Azure, or GCP).

Required skills

  • Python, Bash or Shell scripting
  • Kubernetes and Docker containerization
  • CI/CD tools (Jenkins, GitHub Actions, GitLab CI, Azure DevOps)
  • Infrastructure as Code tools (Terraform, Ansible)
  • Networking fundamentals (DNS, TCP/IP, load balancers, SSL)
  • Observability tools: Prometheus, Grafana, Splunk, ELK Stack, Datadog, AppDynamics
  • Logging, tracing and metrics collection
  • Incident management, root cause analysis, problem management
  • High availability, disaster recovery and capacity planning
  • SQL and database fundamentals
  • Git version control

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec Keasis.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

Why are you reporting this job?

Thank you for your report. We will review this job.

A question about this job?

Ask it here: you will get the full job summary by e-mail, right away.

💬 Chat with us on Telegram

Published 2 months ago

23 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

Keasis

Montréal