JPMorganChase Logo

JPMorganChase

Site Reliability Engineer III - Python, Grafana, Splunk, AWS, Jenkins

Posted 3 Hours Ago
Be an Early Applicant
Hybrid
Hyderabad, Telangana
Mid level
Hybrid
Hyderabad, Telangana
Mid level
Builds and improves reliable, scalable applications and cloud infrastructure through automation, monitoring, infrastructure as code, and CI/CD. The role configures systems, implements observability and SLO-based alerting, investigates incidents, applies authorized AI capabilities to operational workflows, and collaborates with engineering teams to resolve issues and reduce recurring toil. It also guides peers in adopting SRE practices and iteratively improving application availability, reliability, and scalability.
The summary above was generated by AI

There’s nothing more exciting than being at the center of a rapidly growing field in technology and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. 
As a Site Reliability Engineer III at JPMorgan Chase within the Employee Platforms Team, you will solve complex and broad business problems with simple and straightforward solutions. Through code and cloud infrastructure, you will configure, maintain, monitor, and optimize applications and their associated infrastructure to independently decompose and iteratively improve on existing solutions. You are a significant contributor to your team by sharing your knowledge of end-to-end operations, availability, reliability, and scalability of your application or platform. 
Job responsibilities

  • Guides and assists others in the areas of building appropriate level designs and gaining consensus from peers where appropriate, supporting adoption of site reliability engineering best practices within your team
  • Collaborates with other software engineers and teams to design, develop, test, and implement deployment and reliability approaches using automated continuous integration and continuous delivery pipelines
  • Uses enterprise-authorized AI capabilities within the work environment to accelerate incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
  • Implements infrastructure, configuration, and network as code for the applications and platforms in your remit
  • Collaborates with technical experts, key stakeholders, and team members to resolve complex problems and proactively address issues using service level indicators and objectives before they impact customers
  • Applies enterprise-authorized AI capabilities within the work environment to identify patterns in operational signals that indicate reliability risk or recurring toil, prioritizing reuse-first improvements tied to SLO outcomes.
  • Familiar with availability, reliability, scalability, and solutions in their applications and works with partners to improve these outcomes iteratively
  • Proactively recognizes road blocks and identifies improvements to solve business problems, including exploring new technologies where appropriate

Required qualifications, capabilities, and skills 

  • Formal training or certification on site reliability engineering concepts and 3+ years applied experience 
  • Proficient in site reliability culture and principles and familiarity with how to implement site reliability within an application or platform
  • Proficient in at least one programming language such as Python, Java/Spring Boot, and .Net
  • Working knowledge of using enterprise-authorized AI capabilities within the work environment to support SRE workflows with strong validation habits and awareness of data sensitivity
  • Experience in Python, Grafana , Prometheus, Dynatrace, Datadog and Splunk, AWS, CICD experience, Jenkins, Github, Terraform
  • Ability to validate AI-assisted operational recommendations before applying changes, escalating when uncertain and following data sensitivity requirements
  • Proficient knowledge of software applications and technical processes within a given technical discipline (e.g., Cloud, AI, Android, etc.)
  • Experience in observability such as white and black box monitoring, service level objective alerting, and telemetry collection
  • Experience with continuous integration and continuous delivery tooling


Preferred  qualifications, capabilities, and skills 

  • Familiarity with AI coding assistant tools
  • Familiarity with container and container orchestration and troubleshooting common networking technologies and issues

Similar Jobs

An Hour Ago
Hybrid
Junior
Junior
Big Data • Marketing Tech • Sales • Software • Analytics • Big Data Analytics
Manage regional billing operations, including client invoicing, manual invoices, discrepancy resolution, client onboarding, credit applications, contract review, revenue recognition, month-end reviews, reporting, and special projects. The role requires strong bookkeeping knowledge, communication, organization, analytical ability, attention to detail, and proficiency with accounting, spreadsheet, office, and CRM tools. This position is based in Hyderabad with in-office work three days per week and North American working hours.
Top Skills: ExcelMS OfficeOracle NetsuitePipedriveQuickbooksSage IntacctSalesforce
3 Hours Ago
In-Office
Senior level
Senior level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Owns the technical product lifecycle from definition through release and end of life. Serves as Technical Product Owner for agile scrum teams, prioritizing and refining backlogs, defining requirements and user stories, supporting PI planning, and coordinating dependencies. Partners with engineering, design, scrum teams, and business leaders while using market and customer insights to guide product strategy and performance.
Top Skills: AgileAhaRallyScrumSoftware Development Life Cycle (Sdlc)
3 Hours Ago
In-Office
Senior level
Senior level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Owns technical product management for two agile development teams, prioritizing and refining backlogs, defining requirements and user stories, managing dependencies, supporting PI planning, and communicating product vision, benefits, and personas. The role also follows and improves product development processes while partnering with Strategic Product Management and technical product management leaders.
Top Skills: AgileAhaRallyScrumSoftware Development Life Cycle (Sdlc)

What you need to know about the Pune Tech Scene

Once a far-out concept, AI is now a tangible force reshaping industries and economies worldwide. While its adoption will automate some roles, AI has created more jobs than it has displaced, with an expected 97 million new roles to be created in the coming years. This is especially true in cities like Pune, which is emerging as a hub for companies eager to leverage this technology to develop solutions that simplify and improve lives in sectors such as education, healthcare, finance, e-commerce and more.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account