We use cookies. Find out more about it here. By continuing to browse this site you are agreeing to our use of cookies.
#alert
Back to search results
New

Technical Operations Lead

First Citizens Bank
United States, North Carolina, Raleigh
100 East Tryon Road (Show on map)
Oct 02, 2026
Overview

We are seeking a highly skilled and motivated Senior Site Reliability Engineer (SRE) to join our Enterprise Observability and Site Reliability Engineering team. This role is responsible for improving the reliability, availability, performance, and operational excellence of critical enterprise platforms and applications.

The ideal candidate combines strong software engineering and infrastructure expertise with a passion for automation, observability, and operational excellence. You will partner closely with application development, cloud engineering, infrastructure, security, and production support teams to establish reliability standards, implement monitoring strategies, and drive proactive risk reduction across the technology landscape.


Responsibilities & Qualifications

Key Responsibilities

Reliability Engineering

  • Design, implement, and maintain highly available, resilient, and scalable technology platforms.
  • Manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets across critical services.
  • Lead reliability reviews and identify opportunities to improve system stability and performance.
  • Drive root cause analysis and corrective actions for high-severity incidents during the problem process.

Observability & Monitoring

  • Serve as a subject matter expert for enterprise observability platforms, including Dynatrace and related monitoring technologies.
  • Develop monitoring, alerting, synthetic testing, and dashboard strategies.
  • Improve visibility into application health, infrastructure performance, user experience, and business transaction monitoring.
  • Partner with engineering teams to embed observability practices throughout the software development lifecycle.

Automation & Platform Engineering

  • Identify and eliminate operational toil through automation.
  • Develop scripts, tooling, integrations, and self-service capabilities.
  • Enhance CI/CD processes to improve deployment reliability and operational efficiency.
  • Support infrastructure-as-code and automation-first operating models.

Incident Management & Operational Excellence

  • Participate in major incident response and problem management activities.
  • Establish and improve operational runbooks, standards, and best practices.
  • Drive reduction in Mean Time to Detect (MTTD) and Mean Time to Restore (MTTR).
  • Implement proactive monitoring and predictive alerting capabilities.

Leadership & Collaboration

  • Provide technical leadership and mentoring to engineers across the organization.
  • Influence reliability and observability strategy at the enterprise level.
  • Collaborate with application owners, cloud teams, security teams, and executive stakeholders.
  • Champion a culture of ownership, accountability, and continuous improvement.

Required Qualifications

  • Bachelor's degree in Computer Science, Engineering, Information Technology, or equivalent experience.
  • 7+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Infrastructure Engineering, or related disciplines.
  • Strong experience supporting mission-critical production environments.
  • Hands-on experience with observability and monitoring platforms such as Dynatrace, Splunk, Databahn or similar technologies.
  • Strong understanding of application performance monitoring (APM), distributed tracing, log analytics, and infrastructure monitoring.
  • Experience with cloud platforms such as AWS, Azure, or Google Cloud Platform.
  • Proficiency with scripting and automation.
  • Strong troubleshooting, analytical, and problem-solving skills.
  • Large financial institution experience in a complex environment.

Preferred Qualifications

  • Deep expertise with Dynatrace platform administration and implementation.
  • Experience in enterprise-scale financial services or highly regulated environments.
  • Knowledge of cloud-native observability patterns and OpenTelemetry.
  • Experience implementing SRE practices including SLOs, Error Budgets, reliability reviews, and operational readiness assessments.
  • Familiarity with enterprise event management and AIOps platforms.
  • Certifications in cloud platforms, Dynatrace, Kubernetes, or related technologies.
  • Experience leading large-scale monitoring transformation initiatives.

Additional Information

Benefits are an integral part of total rewards and First Citizens Bank is committed to providing a competitive, thoughtfully designed and quality benefits program to meet the needs of our associates. More information can be found at https://jobs.firstcitizens.com/benefits.

Applied = 0

(web-9db6c7984-nthgv)