Who Can Apply
- Candidates must be legally authorized to work in Canada
Job Description
Insight Global is hiring a Site Reliability Engineer (SRE) for a leading Vancouver-based retailer.
This is a hands-on engineering role responsible for improving the reliability, availability, performance, and operational excellence of critical engineering platforms and cloud infrastructure. As part of the Enterprise Platform Engineering team, you will work closely with the Lead Platform Engineer, GitLab Platform Engineer, and Terraform Cloud Engineer to ensure platform stability, observability, incident management, and service resilience across the organization's technology ecosystem.
The ideal candidate combines strong software engineering and infrastructure expertise with a passion for automation, monitoring, and continuous improvement.
Responsibilities
- Design, implement, and maintain observability solutions, including monitoring, logging, alerting, dashboards, and service health metrics.
- Define and track Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets to drive platform reliability.
- Lead incident response, troubleshooting, root cause analysis, and post-incident reviews for production systems.
- Automate operational processes to reduce manual effort and improve system reliability.
- Partner with Platform Engineering teams to enhance the reliability and performance of GitLab SaaS, Terraform Cloud, cloud infrastructure, and CI/CD platforms.
- Build and maintain platform health dashboards, operational reporting, and reliability metrics.
- Identify and eliminate operational bottlenecks through automation, scalability improvements, and proactive monitoring.
- Support disaster recovery, resiliency, capacity planning, and business continuity initiatives.
- Participate in on-call rotations and drive continuous improvements in operational readiness.
- Leverage AI-assisted capabilities to improve monitoring, incident response, performance analysis, and operational workflows.
We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to HR@insightglobal.com.To learn more about how we collect, keep, and process your private information, please review Insight Global's Workforce Privacy Policy: https://insightglobal.com/workforce-privacy-policy/.
Required Skills & Experience
Required Skills & Experience
- 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Cloud Operations.
- Strong experience supporting large-scale production environments in AWS and Kubernetes.
- Experience implementing observability solutions using tools such as Grafana, Prometheus, Datadog, New Relic, Splunk, or OpenTelemetry.
- Strong understanding of incident management, root cause analysis, reliability engineering, and production support.
- Experience with Infrastructure-as-Code and automation tools such as Terraform.
- Experience supporting CI/CD platforms and modern cloud-native environments.
- Proficiency with scripting or automation using Python, Go, Bash, or similar languages.
- Strong understanding of high availability, resilience, disaster recovery, and scalability principles.
- Experience leveraging AI-assisted tools to improve operational efficiency and engineering workflows.
Preferred Experience
- Experience supporting GitLab SaaS, Terraform Cloud, or other enterprise engineering platforms.
- Experience defining and operating SLI/SLO frameworks.
- Experience with platform performance engineering and capacity planning.
- Experience working within large-scale enterprise platform engineering organizations.
Benefit packages for this role will start on the 1st day of employment and include medical, dental, and vision insurance, as well as HSA, FSA, and DCFSA account options, and 401k retirement account access with employer matching. Employees in this role are also entitled to paid sick leave and/or other paid time off as provided by applicable law.