Photon Logo

Photon

Site Reliability Engineer - IN

Reposted 9 Days Ago
Be an Early Applicant
Remote
Hiring Remotely in India
Mid level
Remote
Hiring Remotely in India
Mid level
Maintain and monitor production systems to ensure uptime and performance. Build automation and infrastructure tooling, analyze metrics, support large-scale distributed applications, partner with development teams on testing and releases, and drive reliability improvements and capacity planning.
The summary above was generated by AI

SRE Engineer is responsible for ensuring website uptime, optimizing performance, and maintaining security of the production application. This role involves monitoring site reliability, addressing technical issues, automating maintenance tasks, and collaborating with cross-functional teams to meet business objectives. 

Responsibilities 

Run the production environment by monitoring availability and taking a holistic view of system health 

Build software and systems to manage platform infrastructure and applications 

Improve reliability, quality, and time-to-market of our suite of software solutions 

Measure and optimize system performance, with an eye toward pushing our capabilities forward, getting ahead of customer needs, and innovating for continual improvement 

Provide primary operational support and engineering for multiple large-scale distributed software applications 

Gather and analyze metrics from operating systems as well as applications to assist in performance tuning and fault finding 

Partner with development teams to improve services through rigorous testing and release procedures 

Participate in system design consulting, platform management, and capacity planning 

Create sustainable systems and services through automation and uplifts 

Balance feature development speed and reliability with well-defined service-level objectives 

Required skills and qualifications 

Bachelor’s degree (or equivalent) in computer science or related discipline 

Experience in SRE, DevOps, or similar roles. 

Expertise in monitoring tools, infrastructure management, and automation. 

Strong problem-solving skills and a collaborative mindset. 

Proactive approach to identifying problems, performance bottlenecks, and areas for improvement 

Similar Jobs

Yesterday
In-Office or Remote
India
Senior level
Senior level
Cloud • Security • Software • Cybersecurity
Oversee, scale, and optimize high-density AI hardware infrastructure across regional data centers. Build Python automation and infrastructure-as-code tooling, integrate incident workflows, develop telemetry pipelines and monitoring dashboards, and improve reliability across private cloud, bare-metal, and virtualized environments. Lead on-call incident response, runbooks, post-mortems, service rollouts, vendor coordination, and field technician activities while driving uptime, performance, and operational readiness.
Top Skills: Ai-Based Anomaly DetectionApi IntegrationsBare-Metal InfrastructureBgpGrafanaInfrastructure As CodeIpv4Ipv6LlmsLokiOpentelemetryPagerdutyPrivate CloudPrometheusPythonSlackTelemetry PipelinesVirtualization
Yesterday
Remote
India
Mid level
Mid level
Information Technology
Develop Python applications, APIs, automation tools, and platform capabilities while supporting CI/CD, Docker, Kubernetes, cloud and on-premise deployments. Implement observability, monitoring, alerting, and SRE practices including SLIs, SLOs, error budgets, incident response, disaster recovery, and root cause analysis. Troubleshoot production systems, automate recurring operational work, maintain documentation, and collaborate across engineering, security, platform, product, operations, and external teams.
Top Skills: AWSAzureAzure DevopsClaude CodeConfluenceDnsDockerGCPGitGithub ActionsGithub CopilotGitlab CiGrafanaHttp/HttpsJenkinsJIRAJSONKubernetesLinux/UnixPower BIPrometheusPythonRest ApisServicenowTcp/IpTerraformWindows
2 Days Ago
Remote
IND
Senior level
Senior level
Cloud • Information Technology • Consulting • App development
Own and improve high-scale backend systems and TB-scale data pipelines. Ensure microservice availability, performance, resilience, and reliability. Build deployment, monitoring, and incident-response automation while contributing scalable backend features. Take end-to-end ownership of production systems, debug complex issues, and collaborate with founders and global teams in a high-ambiguity startup environment.
Top Skills: .Net CoreAWSC#Distributed SystemsDynamoDBKinesisMicroservicesNode.jsPythonTypescript

What you need to know about the Pune Tech Scene

Once a far-out concept, AI is now a tangible force reshaping industries and economies worldwide. While its adoption will automate some roles, AI has created more jobs than it has displaced, with an expected 97 million new roles to be created in the coming years. This is especially true in cities like Pune, which is emerging as a hub for companies eager to leverage this technology to develop solutions that simplify and improve lives in sectors such as education, healthcare, finance, e-commerce and more.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account