SRE Lead
Chubb (Chubb) · Other
- Technology & Engineering
- Insurance & Actuarial
- Malaysia
- Regular · Full time
Key Objective:
- Lead the Site Reliability Engineering function to define and drive the organisation’s reliability engineering strategy — bridging software development and operations through engineering discipline, not manual process.
- Own the end-to-end reliability posture of production systems: define SLO/SLI frameworks, govern error budgets, and enforce production-readiness standards to protect business continuity.
- Build, mentor, and scale a high-performing SRE team that prioritises engineering over toil — automating manual work, embedding reliability into the SDLC, and driving down mean time to recovery through systematic improvement.
- Champion observability-led engineering through full-stack Dynatrace adoption, AIOps integration, and data-driven reliability decision-making at every layer of the stack.
- Serve as the primary reliability engineering partner to development and platform leadership, shaping architecture decisions, release policies, and automation strategy.
Key Responsibilities:
- Define and drive the SRE strategy and multi-year roadmap aligned to business priorities.
- Lead and develop the SRE team, including hiring, onboarding, performance management, career development, and succession planning.
- Own incident management, including severity classification, escalation, response SLAs, and leadership of major incidents.
- Champion blameless postmortems, root cause analysis, and implementation of systemic fixes.
- Establish and govern SLOs, SLIs, and error budgets, ensuring reliability targets are aligned to business needs.
- Drive resilience engineering, including chaos engineering, GameDays, production readiness reviews, and failure mode analysis.
- Reduce toil through automation, improved runbooks, and continuous operational improvement.
- Own observability and alerting standards, including monitoring strategy, dashboards, and alert quality.
- Partner with engineering, architecture, product, and leadership teams to embed reliability into design and delivery.
- Represent the SRE function in senior forums and provide reporting on reliability, risk, and operational performance.
Qualifications:
- Degree in Computer Science, Software Engineering, IT, or a related technical field.
- 8+ years’ experience in software engineering, platform reliability, or SRE, including 3+ years in people leadership.
- Strong hands-on coding ability in Python, Go, or similar, with experience building automation and self-healing solutions.
- Proven experience leading enterprise-scale SRE or platform reliability functions.
- Experience defining and operating RTO/RPO and SLI/SLO frameworks.
- Strong background in observability, production readiness, error budgets, and chaos engineering.
- Experience leading on-call models, incident response, and executive stakeholder engagement.
- Solid understanding of SDLC, Agile, and DevOps delivery models.
- ITIL Foundation is desirable.
Managerial & Soft Skills:
- Proven people leader with experience building and coaching high-performing teams.
- Strategic thinker who can turn business priorities into reliability roadmaps.
- Strong communicator who can explain technical risk in business terms.
- Calm and decisive during major incidents.
- Influential partner across engineering, product, and leadership teams.
- Strong advocate for developer experience and sustainable on-call practices.
- Data-driven and able to balance reliability, speed, and cost.
- Champions psychological safety, continuous learning, and operational excellence.
Technical Skills:
- Expert in Dynatrace, with experience in observability, monitoring, SLOs, tracing, and log management.
- Proficient in Grafana, Prometheus, Splunk, ELK, Azure Monitor, and Log Analytics.
- Strong knowledge of OpenTelemetry and telemetry pipeline design.
- Experience with ServiceNow, CI/CD tools, Kubernetes, Docker, Terraform, and Bicep.
- Familiar with Java, .NET, databases, APIs, Kafka, and cloud platforms, especially Azure.
- Experience with AIOps, AI-assisted triage, and automation tooling.
- Able to support reliability engineering through scripting, auto-remediation, and operational automation.
Desired:
- Experience in insurance or financial services.
- Dynatrace, Azure, ITIL 4, or Google Cloud/SRE-related certifications.
- Experience with chaos engineering, AIOps, MLOps, and FinOps.
- Strong analytical skills and experience working with large operational datasets.