Job Description
A large healthcare client of ours is seeking a Site Reliability Engineer to join a fast-paced operations team responsible for maintaining platform stability, managing critical incidents, and supporting enterprise applications. This individual will play a key role in incident response, service reliability, cross-functional coordination, and operational excellence.
The ideal candidate brings a strong operations mindset, excels under pressure, and can effectively lead technical incident management efforts while communicating with both technical and business stakeholders.
Responsibilities
Lead and coordinate response efforts for P1/P2 production incidents.
Serve as Incident Commander during major outages and service disruptions.
Drive incident triage and engage appropriate engineering, infrastructure, and business teams.
Monitor application and platform health to ensure system availability and reliability
Facilitate root cause analysis activities and support post-incident reviews
Manage escalation processes and ensure timely resolution of critical production issues.
Draft executive-level and customer-facing communications during outages and service-impacting events.
Partner with engineering, infrastructure, cloud, application support, and business teams to improve operational processes.
Participate in an on-call rotation supporting enterprise operations.
Identify opportunities for automation and operational efficiency improvements.
Support reliability initiatives focused on reducing incident frequency and improving service performance
payrate - $50-60
We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to HR@insightglobal.com.To learn more about how we collect, keep, and process your private information, please review Insight Global's Workforce Privacy Policy: https://insightglobal.com/workforce-privacy-policy/.
Required Skills & Experience
3–7 years of experience in Reliability Engineering, Production Support, Operations Engineering, or Incident Management
Hands-on experience managing high-severity production incidents (P1/P2).
Strong experience leading incident bridges and coordinating cross-functional response efforts
Excellent verbal and written communication skills with the ability to communicate effectively with technical teams and executive leadership.
Ability to thrive in a fast-paced operational environment.
Strong troubleshooting and problem-solving skills
Experience working within an on-call support model
Nice to Have Skills & Experience
Experience supporting cloud-based environments (AWS, Azure, or GCP)
Knowledge of observability and monitoring tools such as Dynatrace, Splunk, Datadog, AppDynamics, Prometheus, or Grafana
Familiarity with ITIL processes and Major Incident Management frameworks
Experience supporting large-scale enterprise applications within healthcare, retail, or highly regulated environments
Exposure to automation and scripting using Python, Bash, PowerShell, or similar technologies
Benefit packages for this role will start on the 1st day of employment and include medical, dental, and vision insurance, as well as HSA, FSA, and DCFSA account options, and 401k retirement account access with employer matching. Employees in this role are also entitled to paid sick leave and/or other paid time off as provided by applicable law.