SID Global Solutions Logo

SID Global Solutions

Site Reliability Engineering (SRE) Lead

Posted One Month Ago
Be an Early Applicant
In-Office
Hyderabad, Telangana
Senior level
In-Office
Hyderabad, Telangana
Senior level
Lead SRE responsible for platform reliability, incident management, automation, observability, and production stability across cloud-native and Kubernetes environments. Mentor teams, run RCA, define SLIs/SLOs, automate operations, manage Apigee APIs, and support CI/CD and IaC initiatives to improve availability and performance.
The summary above was generated by AI

Job Title: Site Reliability Engineering (SRE) Lead


Location: Hyderabad / Mumbai
Employment Type: Full-Time

 

About SID Global Solutions

SID Global Solutions (SIDGS) is a leading Digital Engineering and Technology Services company specializing in Cloud, API Management, DevOps, Platform Engineering, Kubernetes, Microservices, and Digital Transformation. We are looking for an experienced Site Reliability Engineering (SRE) Lead to drive platform reliability, operational excellence, and production stability for mission-critical enterprise applications.

 

Job Summary

We are seeking a highly skilled SRE Lead with strong expertise in cloud infrastructure, Kubernetes, API Management, and production operations. The ideal candidate will be responsible for ensuring the availability, scalability, performance, and reliability of enterprise applications while leading incident management, observability, and automation initiatives.

In this role, you will work closely with Development, DevOps, Infrastructure, Platform Engineering, and Application Support teams to maintain highly available production environments, reduce operational risks, and improve service reliability through automation and proactive monitoring.

 

Key Responsibilities

Reliability Engineering

  • Define and implement Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets to improve platform reliability.
  • Review application architecture and infrastructure designs to ensure scalability, resilience, and high availability.
  • Drive reliability improvements across cloud-native applications and distributed systems.
  • Identify opportunities to eliminate operational bottlenecks through automation and process improvements.

Incident & Production Management

  • Lead Major Incident Management (P1/P2) activities and act as the primary technical escalation point for critical production issues.
  • Coordinate with cross-functional teams to restore services within defined SLAs.
  • Conduct Root Cause Analysis (RCA) and drive preventive and corrective actions to reduce recurring incidents.
  • Participate in Change Management and Release activities to ensure production stability.
  • Prepare incident reports and communicate status updates to business and technology stakeholders.

Automation & Platform Engineering

  • Develop automation scripts and self-healing solutions to improve operational efficiency.
  • Build reusable operational runbooks and standard operating procedures.
  • Automate routine operational tasks using Python, Bash, or similar scripting languages.
  • Support Infrastructure as Code (IaC) initiatives and CI/CD pipeline improvements.

Monitoring & Observability

  • Design and maintain enterprise monitoring and alerting solutions.
  • Create dashboards and alerts to proactively monitor application health, infrastructure, APIs, and Kubernetes environments.
  • Analyze performance trends and recommend improvements for system stability and capacity planning.
  • Ensure effective monitoring coverage across production environments.

API & Kubernetes Administration

  • Manage and troubleshoot Google Apigee API Gateway configurations, policies, and traffic routing.
  • Monitor Kubernetes clusters, workloads, namespaces, ingress controllers, and container health.
  • Optimize application performance and resource utilization within Kubernetes environments.
  • Support production deployments and post-release validation activities.

Leadership & Collaboration

  • Mentor SRE, DevOps, and Production Support engineers.
  • Establish operational best practices, troubleshooting guidelines, and technical documentation.
  • Collaborate with Development, QA, Infrastructure, Security, and Business teams to improve platform reliability.
  • Drive a culture of continuous improvement, automation, and operational excellence.

 

Required Technical Skills

  • Strong experience with Google Cloud Platform (GCP).
  • Hands-on expertise in Google Apigee API Management.
  • Experience managing production Kubernetes (K8s) environments.
  • Good understanding of NGINX, reverse proxy, and load balancing concepts.
  • Strong knowledge of Linux/Unix Administration.
  • Experience with Python, Bash, or Go scripting.
  • Familiarity with CI/CD pipelines, Git, and Jenkins.
  • Understanding of Infrastructure as Code (Terraform or equivalent is preferred).

Monitoring & Observability Tools

Experience with one or more of the following:

  • Datadog
  • Dynatrace
  • Prometheus
  • Grafana
  • Splunk
  • ELK Stack
  • AppDynamics

 

Required Qualifications

  • 7–10 years of experience in Site Reliability Engineering (SRE), DevOps, Platform Engineering, or Production Support.
  • Strong experience supporting enterprise production environments.
  • Hands-on experience with Kubernetes, GCP, and Google Apigee.
  • Good understanding of distributed systems, microservices architecture, and cloud-native applications.
  • Experience in Incident, Problem, Change, and Release Management.
  • Ability to troubleshoot complex production issues and coordinate cross-functional teams during critical incidents.

 

Preferred Qualifications

  • Experience in Banking, Financial Services, or other enterprise environments.
  • ITIL Foundation certification.
  • Google Cloud Professional Certification.
  • Kubernetes Certification (CKA/CKAD) is an added advantage.
  • Experience with Service Mesh technologies such as Istio is preferred.

 



Similar Jobs

45 Minutes Ago
In-Office
Senior level
Senior level
Big Data • Fintech • Information Technology • Insurance • Financial Services
Leads medium to large technical projects and small to medium programs from planning through delivery. Manages stakeholders, roadmaps, budgets, resources, risks, requirements, reporting, and project artifacts. Facilitates Agile and Kanban practices, removes delivery barriers, supports product owners, and drives process improvement. The role requires autonomous leadership across technology, financial services, cybersecurity, information risk, and third-party vendor environments.
Top Skills: AgileClarity PpmConfluenceJIRAKanbanScrumSdlc
2 Hours Ago
Hybrid
Expert/Leader
Expert/Leader
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Leads architecture, development, deployment, and monitoring of AI-native and cloud-native systems. Builds LLM integrations, agentic workflows, retrieval pipelines, evaluation frameworks, data pipelines, and ServiceNow applications. Provides technical leadership, mentorship, code reviews, customer issue resolution, and cross-functional collaboration while ensuring security, scalability, reliability, and responsible AI practices.
Top Skills: AgileAmazon BedrockAWSAzureAzure OpenaiCi/CdDockerFastapiGCPGlidescriptGoogle Vertex AiInfrastructure As CodeJavaJavaScriptKubernetesLangchainLlamaindexMicroservicesPgvectorPineconePythonRest ApisRetrieval-Augmented GenerationScrumSemantic KernelSemantic SearchServicenowSpring BootVector DatabasesWeaviate
2 Hours Ago
Hybrid
Senior level
Senior level
Fintech • Mobile • Payments • Software • Financial Services
Leads a KYC Operations team conducting customer due diligence and risk-based reviews. Responsibilities include coaching analysts, managing performance and KPIs, handling escalations, improving processes, maintaining quality assurance controls, and ensuring compliance with KYC/AML regulations. The role also requires stakeholder collaboration, root-cause analysis, regulatory expertise, and ownership of operational initiatives.
Top Skills: LookerSuperset

What you need to know about the Pune Tech Scene

Once a far-out concept, AI is now a tangible force reshaping industries and economies worldwide. While its adoption will automate some roles, AI has created more jobs than it has displaced, with an expected 97 million new roles to be created in the coming years. This is especially true in cities like Pune, which is emerging as a hub for companies eager to leverage this technology to develop solutions that simplify and improve lives in sectors such as education, healthcare, finance, e-commerce and more.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account