Centre for Computational Technologies Pvt. Ltd. (CCTech) Logo

Centre for Computational Technologies Pvt. Ltd. (CCTech)

Lead Site Reliability Engineer

Posted Yesterday
Be an Early Applicant
In-Office
Pune, Maharashtra, IND
Expert/Leader
In-Office
Pune, Maharashtra, IND
Expert/Leader
Lead reliability engineering for mission-critical AWS systems by designing scalable and resilient architectures, building infrastructure and deployment automation, improving observability, managing incidents, reducing operational toil, and driving reliability, cost, and operational maturity. The role partners with clients and cross-functional teams, influences technical direction, mentors engineers, and advances SRE practices.
The summary above was generated by AI
We are looking for Lead Site Reliability Engineers who combine deep reliability engineering expertise with strong ownership, communication, and systems thinking.
This is not a traditional operations-only role.
We are looking for techno-business leaders who can:
  • operate and improve mission-critical systems,
  • drive architectural and reliability decisions,
  • collaborate directly with clients and stakeholders,
  • identify platform and business improvement opportunities,
  • act as trusted technical anchors for both CCTech and customer teams.
Someone who understands production deeply, thinks in systems, values automation over toil, and continuously improves reliability, scalability, and operational maturity.
This role is highly hands-on while also requiring leadership, initiative, and the ability to influence technical direction across teams.

Key Responsibilities
  • Own reliability, uptime, and operational health of mission-critical cloud systems
  • Design systems for scalability, resilience, operability, and cost efficiency
Drive SRE best practices including:
  • Observability,incident prevention,postmortems,reliability engineering and automation-first operations
  • Lead architecture and operational decisions balancing reliability, scalability, maintainability, and cost.
  • Build and evolve Infrastructure-as-Code, CI/CD pipelines, deployment workflows, and recovery automation
  • Design and improve monitoring, logging, alerting, and observability frameworks
  • Lead critical incident investigations, root cause analysis, and long-term corrective actions
  • Reduce operational toil through automation, reusable tooling, and engineering discipline
  • Collaborate directly with product teams, platform teams, and client stakeholders to align technical direction with business needs
  • Act as a trusted technical partner during architecture discussions, solution reviews, brainstorming sessions, and platform evolution initiatives
  • Proactively identify opportunities for platform improvements, operational maturity, automation, reliability optimization, and cost reduction.
  • Mentor engineers, raise technical standards, and contribute to building high-performing teams
  • Help shape SRE culture, operational maturity, and engineering practices across projects


Requirements
  • 8–12 years of experience in SRE / DevOps / Cloud Engineering roles
Strong hands-on experience with:
  • AWS production systems
  • Infrastructure-as-Code (Terraform / CloudFormation)
  • CI/CD pipelines and deployment automation
  • Containerized environments (Docker / Kubernetes / ECS)
Proven experience in:
  • designing and operating reliable distributed systems,
  • handling production incidents at scale,
  • debugging complex system failures,
  • improving system reliability and operational maturity,
  • driving automation-first engineering practices
  • Strong programming/scripting ability in Python (preferred), Go.
  • Strong understanding of distributed systems, observability, scalability, performance optimization, and cost-aware architecture.
  • Experience working closely with stakeholders, customers, or cross-functional teams in technical discussions and solution alignment
  • Ability to independently drive initiatives, influence decisions, and take ownership beyond assigned tasks
  • Excellent communication skills with the ability to explain technical concepts clearly to both engineering and non-engineering stakeholders
Good to Have
  • Experience with SLOs, SLIs, error budgets, and reliability governance
  • Exposure to API platforms, Identity systems (OAuth2/OIDC), or platform engineering initiatives
  • Experience with chaos engineering, failure testing, or resilience validation
  • Exposure to regulated or enterprise-scale environments
  • Background in backend engineering before transitioning into SRE/DevOps
  • Experience contributing to technical proposals, architecture reviews, or client-facing solution discussions

Benefits
  • High ownership role with direct impact on mission-critical systems
  • Opportunity to shape platform reliability, operational maturity, and engineering direction
  • Exposure to advanced areas such as Digital Twin, AI/ML systems, cloud-native platforms, and large-scale distributed architectures
  • Work closely with global engineering organizations and strategic technology partners
  • Opportunity to grow into a technical leadership and client-facing advisory role within CCTech.


Similar Jobs

17 Days Ago
Hybrid
Pune, Maharashtra, IND
Senior level
Senior level
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Lead SRE responsible for designing, automating, and improving CI/CD pipelines, ensuring system availability and performance, managing incidents and postmortems, capacity planning, and mentoring engineers. Partner with development and product teams to shift reliability left, standardize support, and drive DevOps automation and operational excellence across distributed systems.
Top Skills: ArtifactoryBitbucketCC++ChefCi/CdDistributed SystemsGitGoItsmJavaJenkinsMavenPerlPythonRuby
7 Days Ago
In-Office
Pune, Maharashtra, IND
Expert/Leader
Expert/Leader
Biotech • Agriculture • Chemical
Leads DevOps and SRE practices for enterprise data and AI platforms, focusing on reliability, scalability, security, automation, and operational excellence. Responsibilities include designing CI/CD pipelines, managing cloud infrastructure, implementing Infrastructure as Code, defining SLOs and observability standards, leading incident response, improving MTTR, and mentoring engineers. The role also establishes operating models, governance, on-call processes, and platform reliability practices across data engineering, AI/ML, application, and product teams.
Top Skills: ArmAWSAzureAzure DevopsBashCi/CdCloudFormationCloudwatchDatadogDockerGCPGitGitlabGrafanaInfrastructure As CodeJenkinsKubernetesPrometheusPythonSlisSlosTerraform
Yesterday
In-Office
Pune, Maharashtra, IND
Mid level
Mid level
Artificial Intelligence • HR Tech • Professional Services • Software
Design, build, and operate scalable, resilient AWS production infrastructure using IaC and CI/CD. Automate load testing, validation, and recovery checks; implement monitoring, observability, and incident prevention. Modernize legacy systems, contribute to internal developer platforms, and enable secure, compliant (FedRAMP-aligned) environments with chaos engineering and cost-optimized architectures.
Top Skills: AWSAws FisBashCloudFormationCloudwatchEcsEventbridgeGithub ActionsGrafanaGremlinJenkinsKubernetesLambdaNode.jsPrometheusPythonRdsTerraform

What you need to know about the Pune Tech Scene

Once a far-out concept, AI is now a tangible force reshaping industries and economies worldwide. While its adoption will automate some roles, AI has created more jobs than it has displaced, with an expected 97 million new roles to be created in the coming years. This is especially true in cities like Pune, which is emerging as a hub for companies eager to leverage this technology to develop solutions that simplify and improve lives in sectors such as education, healthcare, finance, e-commerce and more.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account