Monitor and improve the reliability, availability, and performance of critical applications and platforms. Respond to incidents, maintain observability dashboards and runbooks, automate operational tasks, support root cause analysis and postmortems, and help implement SRE practices such as SLIs, SLOs, SLAs, and error budgets. Collaborate with engineering, cloud, infrastructure, and application teams across hybrid cloud and containerized environments.
Description and Requirements
ole Reasonability
MetLife is seeking a Site Reliability Engineer (SRE) to ensure the reliability, availability, and performance of critical applications and platforms.
The SRE Engineer will monitor production systems, respond to incidents, improve observability, maintain runbooks, and automate operational tasks. Working closely with engineering, cloud, and infrastructure teams, the role supports SRE practices, operational readiness, and service reliability through SLIs, SLOs, and Error Budget management.
Core Responsibilities
Skills and Experience
Minimum Qualifications
Preferred Exposure
About MetLife
Recognized on Fortune magazine's list of the "World's Most Admired Companies" and Fortune World's 25 Best Workplaces™, MetLife, through its subsidiaries and affiliates, is one of the world's leading financial services companies; providing insurance, annuities, employee benefits and asset management to individual and institutional customers. With operations in more than 40 markets, we hold leading positions in the United States, Latin America, Asia, Europe, and the Middle East.
As part of our New Frontier strategy, MetLife is building an AI-enabled, people-centered future. We're looking for people who bring curiosity, adaptability, and a growth mindset as we use AI to enhance how we serve customers, support communities, and evolve the way work gets done. At MetLife, AI is a responsible partner that supports human judgment, creativity, and continuous improvement while helping us build trust, inclusion, and long-term value.
Our purpose is simple - to help our colleagues, customers, communities, and the world at large create a more confident future. United by purpose and guided by our core values - Win Together, Do the Right Thing, Deliver Impact Over Activity, and Think Ahead - we're inspired to transform the next century in financial services. At MetLife, it's #AllTogetherPossible . Join us!
#BI-Hybrid
ole Reasonability
MetLife is seeking a Site Reliability Engineer (SRE) to ensure the reliability, availability, and performance of critical applications and platforms.
The SRE Engineer will monitor production systems, respond to incidents, improve observability, maintain runbooks, and automate operational tasks. Working closely with engineering, cloud, and infrastructure teams, the role supports SRE practices, operational readiness, and service reliability through SLIs, SLOs, and Error Budget management.
Core Responsibilities
- Monitoring: Monitor service health, dashboards, alerts, and key reliability indicators for assigned applications and platforms.
- Incident Response: Respond to alerts, support bridge calls, gather evidence, execute runbooks, communicate status, and escalate when required.
- Observability Support: Create and maintain dashboards, log queries, telemetry checks, alert validation, and actionable monitoring signals.
- Runbook Management: Document operational procedures, update recovery steps, validate readiness with service owners, and support knowledge sharing.
- Automation: Create scripts for repetitive checks, data collection, remediation, operational reporting, and toil reduction.
- Problem Follow-up: Support root cause analysis, postmortem documentation, and closure of assigned corrective/preventive action items.
- Continuous Improvement: Identify alert noise, toil, monitoring gaps, and preventive improvements for senior SRE review.
- SRE Alignment: Support adoption of SLOs, SLIs, SLAs, error budgets, operational readiness reviews, and production support standards.
- AI Readiness: Use or help improve AI-assisted tools for anomaly detection, incident correlation, root cause hints, and operational knowledge retrieval.
- Collaboration: Work with engineering, infrastructure, cloud, and application teams to align service performance with business goals.
Skills and Experience
- Foundations: Linux, networking fundamentals, application support, cloud fundamentals, production operations, and ITIL-style incident/change processes.
- Scripting: Python, PowerShell, Bash, or equivalent scripting for automation, diagnostics, evidence collection, and reporting.
- Tools: Git, CI/CD basics, ServiceNow or equivalent ticketing; exposure to Elastic/ELK, Grafana, Prometheus, Splunk, APM, and Azure Monitor preferred.
- Cloud & Containers: Azure services, Docker, Kubernetes, and hybrid cloud operations exposure; Terraform or infrastructure-as-code awareness preferred.
- Reliability: Basic understanding of SLIs, SLOs, SLAs, error budgets, alerting, incident response, postmortems, and operational runbooks.
- AI / AIOps Readiness: Ability to use AI-assisted investigation, anomaly detection, and correlation tools responsibly, with strong validation of evidence.
- Database: Hands-on SQL skills for operational diagnostics, data validation, and service health checks.
- Execution: Disciplined follow-through, evidence capture, documentation, escalation hygiene, and collaboration during incidents.
- Learning Mindset: Willingness to deepen skills in cloud, Kubernetes, observability, automation, resilience engineering, and secure operations.
Minimum Qualifications
- 3-7 years in production support, DevOps, infrastructure, cloud operations, or software engineering.
- Experience supporting business-critical systems and working in incident, problem, and change management processes.
- Ability to script and automate standard operational tasks using Python, PowerShell, Bash, or equivalent.
- Bachelor's degree in computer science, engineering, or equivalent practical experience.
- Exposure to regulated enterprise, insurance, banking, or financial services environments preferred.
- Business proficiency in English; Japanese language skills are a plus.
Preferred Exposure
- Hybrid cloud platforms including on-premises and Azure-hosted services.
- Observability platforms such as ELK/Elastic, Grafana, Prometheus, Splunk, Azure Monitor, and Azure Application Insights.
- GitHub, Azure DevOps, pipelines, repositories, and operational change controls.
- Kubernetes-based production services and containerized application support.
- SRE practices including operational readiness reviews, service health reviews, a
About MetLife
Recognized on Fortune magazine's list of the "World's Most Admired Companies" and Fortune World's 25 Best Workplaces™, MetLife, through its subsidiaries and affiliates, is one of the world's leading financial services companies; providing insurance, annuities, employee benefits and asset management to individual and institutional customers. With operations in more than 40 markets, we hold leading positions in the United States, Latin America, Asia, Europe, and the Middle East.
As part of our New Frontier strategy, MetLife is building an AI-enabled, people-centered future. We're looking for people who bring curiosity, adaptability, and a growth mindset as we use AI to enhance how we serve customers, support communities, and evolve the way work gets done. At MetLife, AI is a responsible partner that supports human judgment, creativity, and continuous improvement while helping us build trust, inclusion, and long-term value.
Our purpose is simple - to help our colleagues, customers, communities, and the world at large create a more confident future. United by purpose and guided by our core values - Win Together, Do the Right Thing, Deliver Impact Over Activity, and Think Ahead - we're inspired to transform the next century in financial services. At MetLife, it's #AllTogetherPossible . Join us!
#BI-Hybrid
MetLife Mahārāshtra, IND Office
MetLife Pune, IN Office

EON West LP-II, 2nd & 3rd Floor, Tower D, S.no. 106(P), 107(P), 108(P)/1, Maharashtra, India, 411057
Similar Jobs at MetLife
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Develops and maintains automated tests across web, API, and native applications. The role applies test methodologies, automation frameworks, CI/CD, test-driven development, Agile practices, and performance testing using JMeter and BlazeMeter. Candidates need experience in financial services or technology domains, proficiency in Java or TypeScript, and hands-on test automation framework experience.
Top Skills:
Azure DevopsBlazemeterCi/CdJavaJmeterTest Automation FrameworksTypescript
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Designs and builds enterprise Azure data platforms using Synapse, Databricks, ADLS Gen2, ADF, PySpark, and Spark SQL. Leads data modernization and migration initiatives, develops scalable pipelines, implements governance and data quality practices, optimizes performance and cloud costs, and mentors engineers. The role also requires advanced SQL, CI/CD, Azure DevOps, security controls, and experience with data warehouse and medallion architectures.
Top Skills:
Azure Data FactoryAzure Data Lake Storage Gen2Azure DatabricksAzure DevopsAzure Event HubsAzure Key VaultAzure Synapse AnalyticsCi/CdGitGitGithub CopilotInformatica PowercenterInfrastructure As CodeKafkaOraclePostgresPower BIPysparkRbacSpark SqlSQLSQL ServerTerraform
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Develop and maintain cloud-native full-stack applications using Java, Spring Boot, ReactJS, APIs, databases, and mobile technologies. Troubleshoot production issues, implement cloud and AI/data best practices, and collaborate with global teams on digital transformation. The role follows Agile, CI/CD, and DevOps practices.
Top Skills:
AgileAPIsAzure DevopsAzure Kubernetes ServiceCi/CdCosmos DbDevOpsJavaReact NativeReactSpring BootSQL ServerSwift
What you need to know about the Pune Tech Scene
Once a far-out concept, AI is now a tangible force reshaping industries and economies worldwide. While its adoption will automate some roles, AI has created more jobs than it has displaced, with an expected 97 million new roles to be created in the coming years. This is especially true in cities like Pune, which is emerging as a hub for companies eager to leverage this technology to develop solutions that simplify and improve lives in sectors such as education, healthcare, finance, e-commerce and more.


