This role involves managing and automating infrastructure, deploying applications, ensuring system availability on cloud platforms, and collaborating with cross-functional teams.
Position: Site Reliability Engineer
Job Summary
As a Site Reliability Engineer, you will play a critical role in ensuring the availability and performance of our customer-facing platform. You will work closely with DevOps, DBA, and Development teams to provision and maintain infrastructure, deploy and monitor our applications, and automate workflows. Your contributions will have a direct impact on customer satisfaction and overall experience.
Responsibilities and Deliverables
- Manage, monitor, and maintain highly available systems (Windows and Linux)
- Analyze metrics and trends to ensure rapid scalability.
- Address routine service requests while identifying ways to automate and simplify.
- Create infrastructure as code using Terraform, ARM Templates, Cloud Formation.
- Maintain data backups and disaster recovery plans.
- Design and deploy CI/CD pipelines using GitHub Actions, Octopus, Ansible, Jenkins, Azure DevOps.
- Adhere to security best practices through all stages of the software development lifecycle
- Follow and champion ITIL best practices and standards.
- Become a resource for emerging and existing cloud technologies with a focus on AWS.
Organizational Alignment
- Reports to the Senior SRE Manager
- This role involves close collaboration with DevOps, DBA, and security teams.
Technical Proficiencies
- Hands-on experience with AWS is a must-have.
- Proficiency analyzing application, IIS, system, security logs and CloudTrail events
- Practical experience with CI/CD tools such as GitHub Actions, Jenkins, Octopus
- Experience with observability tools such as New Relic, Application Insights, AppDynamics, or DataDog.
- Experience maintaining and administering Windows, Linux, and Kubernetes.
- Experience in automation using scripting languages such as Bash, PowerShell, or Python.
- Configuration management experience using Ansible, Terraform, Azure Automation Run book or similar.
- Experience with SQL Server database maintenance and administration is preferred.
- Good Understanding of networking (VNET, subnet, private link, VNET peering).
- Familiarity with cloud concepts including certificates, Oauth, AzureAD, ASE, ASP, AKS, Azure Apps, Load Balancers, Application Gateway, Firewall, Load Balancer, API Management, SQL Server, Databases on Azure
Experience
- 5+ years of experience in SRE or System Administration role
- Demonstrated ability building and supporting high availability Windows/Linux servers, with emphasis on the WISA stack (Windows/IIS/SQL Server/ASP.net)
- 3+ years of experience with CI/CD tools
- 3+ years of experience working with cloud technologies including AWS, Azure.
- 1+ years of experience working with container technology including Docker and Kubernetes.
- Comfortable using Scrum, Kanban, or Lean methodologies.
Education
- Bachelor’s Degree or College Diploma in Computer Science, Information Systems, or equivalent experience.
Similar Jobs
Information Technology • Legal Tech • Analytics
Build and operate secure, scalable, highly available cloud and hybrid platforms. Lead platform engineering, internal developer platforms, infrastructure-as-code, DevOps, and DevSecOps initiatives. Define SLIs, SLOs, and error budgets; lead incident response and postmortems; maintain disaster recovery plans; automate operational workflows; and implement observability solutions. Partner with engineering, security, and platform teams on reliability improvements, technical roadmaps, and architecture decisions.
Top Skills:
AWSAzureCi/CdCloud ComputingConfiguration ManagementDevsecopsDockerInfrastructure As CodeKubernetesLinuxObservabilitySecrets Management
Information Technology
Develop Python applications, APIs, automation tools, and platform capabilities while supporting CI/CD, Docker, Kubernetes, cloud and on-premise deployments. Implement observability, monitoring, alerting, and SRE practices including SLIs, SLOs, error budgets, incident response, disaster recovery, and root cause analysis. Troubleshoot production systems, automate recurring operational work, maintain documentation, and collaborate across engineering, security, platform, product, operations, and external teams.
Top Skills:
AWSAzureAzure DevopsClaude CodeConfluenceDnsDockerGCPGitGithub ActionsGithub CopilotGitlab CiGrafanaHttp/HttpsJenkinsJIRAJSONKubernetesLinux/UnixPower BIPrometheusPythonRest ApisServicenowTcp/IpTerraformWindows
Cloud • Information Technology • Consulting • App development
Own and improve high-scale backend systems and TB-scale data pipelines. Ensure microservice availability, performance, resilience, and reliability. Build deployment, monitoring, and incident-response automation while contributing scalable backend features. Take end-to-end ownership of production systems, debug complex issues, and collaborate with founders and global teams in a high-ambiguity startup environment.
Top Skills:
.Net CoreAWSC#Distributed SystemsDynamoDBKinesisMicroservicesNode.jsPythonTypescript
What you need to know about the Pune Tech Scene
Once a far-out concept, AI is now a tangible force reshaping industries and economies worldwide. While its adoption will automate some roles, AI has created more jobs than it has displaced, with an expected 97 million new roles to be created in the coming years. This is especially true in cities like Pune, which is emerging as a hub for companies eager to leverage this technology to develop solutions that simplify and improve lives in sectors such as education, healthcare, finance, e-commerce and more.



