METRO AG Logo

METRO AG

Sr. Devops Engineer

Posted 2 Days Ago
Be an Early Applicant
Hybrid
Kharadi, Pune, Maharashtra, IND
Senior level
Hybrid
Kharadi, Pune, Maharashtra, IND
Senior level
Build and maintain reliable, scalable cloud-native systems on GCP using Docker and Kubernetes. Define and monitor SLOs, SLAs, and SLIs; automate infrastructure with Terraform, Helm, and Kustomize; develop observability and alerting with Datadog and GCP tools; conduct incident reviews; integrate reliability practices into GitHub Actions CI/CD pipelines; and troubleshoot databases, networking, and Linux systems.
The summary above was generated by AI
Company Description

Company Description

Metro Global Solution Center (MGSC) is internal solution partner for METRO, a €31.6 Billion international wholesaler with operations in more than 30 countries. The store network comprises a total of 623 stores in 21 countries, of which 522 offer out-of-store delivery (OOS), and 94 dedicated depots. In 12 countries, METRO runs only the delivery business by its delivery companies (Food Service Distribution, FSD). 

HoReCa and Traders are core customer groups of METRO. The HoReCa section includes hotels, restaurants, catering companies as well as bars, cafés and canteen operators. The Traders section includes small grocery stores and kiosks. The majority of all customer groups are small and medium-sized enterprises as well as sole traders. METRO helps them manage their business challenges more effectively. 

MGSC, location wise is present in Pune (India), Düsseldorf (Germany) and Szczecin (Poland). We provide HR, Finance, IT, Strategy, Branding & Business operations support to 31 countries, speak 24+ languages and process over 18,000 transactions a day. We are setting tomorrow’s standards for customer focus, digital solutions, and sustainable business models. For over 10 years, we have been providing services and solutions from our two locations in Pune and Szczecin. This has allowed us to gain extensive experience in how we can best serve our internal customers with high quality and passion. We believe that we can add value, drive efficiency, and satisfy our customers.

Job Description

Job Description

Role Overview
We are seeking a Senior Site Reliability Engineer with strong experience in building and
maintaining scalable, resilient systems. The ideal candidate will have hands-on expertise in
cloud-native technologies, infrastructure as code, observability, and automation, with a
focus on Google Cloud Platform (GCP).
 

Key Responsibilities

  • Ensure the stability and reliability of cloud-native applications deployed on GCP, containerized with Docker and orchestrated via Kubernetes.
  • Define, implement, and monitor SLOs, SLAs, and SLIs to measure system performance and user experience.
  • Automate infrastructure provisioning using Terraform and manage Kubernetes configurations with Kustomize and Helm.
  • Develop and maintain monitoring and alerting systems using Datadog and GCP-native tools.
  • Conduct incident analysis and postmortems to drive continuous improvement.
  • Collaborate with development teams to integrate reliability practices into CI/CD pipelines using GitHub Actions.
  • Manage and troubleshoot database systems, particularly PostgreSQL and Cassandra.
  • Apply networking knowledge and Linux system administration skills to troubleshoot and optimize system connectivity and performance.

Work Experience & Skills

  • 7+ years of experience in Site Reliability Engineering.
  • Proven experience designing and operating elastic, resilient systems in cloud
  • environments.
  • Strong understanding of GCP, Kubernetes, and container orchestration.
  • Proficiency in infrastructure as code and configuration management tools (Terraform,
  • Helm, Kustomize).
  • Experience with monitoring and observability tools (Datadog, GCP Monitoring).
  • Solid scripting skills in bash and familiarity with automation frameworks.
  • Experience with CI/CD pipelines, especially using GitHub Actions.
  • Familiarity with networking fundamentals and troubleshooting.
  • Strong coding skills and ability to develop reliability-focused tooling.
  • Excellent communication skills in English (written and spoken)
  • Other Requirements

  • Strong problem-solving skills and a process-oriented mindset.
  • Ability to work independently and collaboratively in a fast-paced environment.
  • Passion for clean code, automation, and continuous improvement.
  • Nice-to-Have

  • Familiarity with monitoring tools (e.g., DataDog, Prometheus, GCP Monitoring).
  • Experience working in Agile/Scrum teams.

 

Qualifications

Education

  • Bachelor’s or Master’s degree in Computer Science, Software Engineering, or equivalent practical experience.

Similar Jobs

9 Days Ago
Hybrid
Pune, Maharashtra, IND
Senior level
Senior level
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Designs and operates Azure and AKS cloud platforms, CI/CD pipelines, Terraform infrastructure, and DevSecOps controls. Responsibilities include improving reliability, observability, incident response, container security, vulnerability remediation, developer self-service, cloud migrations, governance, and AI-assisted engineering workflows. The role partners with application and security teams to expand platform adoption, enforce secure SDLC practices, and support scalable enterprise operating models.
Top Skills: AksAzureAzure DevopsBashDockerEntra IdGithub ActionsGithub CopilotGithub EnterpriseKubernetesPowershellPythonRbacTerraform
13 Days Ago
Hybrid
Pune, Maharashtra, IND
Senior level
Senior level
Artificial Intelligence • Cloud • Sales • Security • Software • Cybersecurity • Data Privacy
Lead the design, operation, and standardization of global cloud infrastructure across AWS, Azure, and GCP. Own large-scale Kubernetes platforms, service mesh, GitOps, infrastructure as code, CI/CD, observability, security, compliance, and on-call excellence. Drive multi-account AWS architecture, mentor engineers, influence cross-functional technical direction, and develop AI-assisted and agentic operational automation using tools such as Amazon Bedrock.
Top Skills: Amazon BedrockAmazon EksArgocdAWSAws App MeshAws OrganizationsAzureAzure AksCi/CdFluxGCPGitopsGoGrafanaHelmIamIstioKargoKubernetesLinkerdLinuxMtlsOpensearchPrometheusPythonService Control PoliciesShell ScriptingTerraform
Yesterday
Remote or Hybrid
India
Senior level
Senior level
Artificial Intelligence • Cloud • Sales • Security • Software • Cybersecurity • Data Privacy
Design, build, and deploy event-driven automations to eliminate operational runbook toil. Operate and support production databases and streaming platforms, participate in on-call rotation, implement least-privilege IAM and auditability, and lead automation roadmap across data infrastructure.
Top Skills: AlertmanagerApache AirflowAws CloudtrailAws CloudwatchAws EventbridgeAws IamAws LambdaClaudeCursorDynamoDBElasticsearchGithub ActionsGitopsJavaScriptJenkinsKafkaKubernetesMySQLPostgresPrometheusPythonRedisTerraformTypescript

What you need to know about the Pune Tech Scene

Once a far-out concept, AI is now a tangible force reshaping industries and economies worldwide. While its adoption will automate some roles, AI has created more jobs than it has displaced, with an expected 97 million new roles to be created in the coming years. This is especially true in cities like Pune, which is emerging as a hub for companies eager to leverage this technology to develop solutions that simplify and improve lives in sectors such as education, healthcare, finance, e-commerce and more.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account