Jefferies Logo

Jefferies

Associate - SRE - Platform Engineering

Posted 2 Days Ago
Be an Early Applicant
In-Office
2 Locations
Mid level
In-Office
2 Locations
Mid level
Build and operate reliable, scalable platforms supporting post-trade processing. Responsibilities include monitoring, incident response, troubleshooting, observability, capacity planning, automation, production support, and release and change management. The role develops dashboards and alerts with Grafana, Prometheus, and OpenTelemetry; supports Kafka-based messaging and distributed systems; analyzes logs, metrics, and traces; and improves resilience by reducing operational toil.
The summary above was generated by AI

Associate Platform Reliability Engineer (SRE)

Location: Mumbai / Pune 

Role Overview

We are seeking a highly motivated Associate Platform Reliability Engineer (SRE) to join our global Platform Reliability Engineering team. This is a hands-on engineering role focused on the reliability, scalability, and operational excellence of critical front-to-back platforms supporting post-trade processing.

The ideal candidate will have a strong software engineering foundation, production support experience, and a passion for automation, observability, and reliability engineering. You will work closely with development, infrastructure, and business teams to improve system resilience, enhance operational visibility, reduce manual intervention, and deliver highly available services.

 

Key Responsibilities

  • Proactively monitor platform health and drive improvements in reliability, performance, availability, and operational efficiency.
  • Perform incident triage, troubleshooting, communication, and post-incident reviews to minimize business impact and prevent recurrence.
  • Collaborate with engineering, infrastructure, and business stakeholders to design and implement scalable and resilient solutions.
  • Build and enhance deployment and observability capabilities, including dashboards, alerts, and service health monitoring using Grafana, Prometheus, and OpenTelemetry.
  • Analyze logs, metrics, and distributed traces to proactively identify system bottlenecks and reliability issues.
  • Support enterprise messaging and event-driven architectures, including Kafka-based platforms and integrations.
  • Apply capacity planning and availability management best practices to improve platform resilience.
  • Develop automation to reduce operational toil, minimize manual intervention, and improve service efficiency.
  • Participate in production support, problem management, release management, and change management activities.

Required Qualifications

  • Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related discipline.
  • 3+ years of experience in Site Reliability Engineering (SRE), Platform Reliability Engineering (PRE), DevOps, Production Support, or Application Support.
  • Strong programming and scripting experience in one or more languages such as Python, Go, or Java.
  • Solid understanding of software engineering principles, data structures, algorithms, and system design.
  • Strong working knowledge of Linux/Unix and Windows Server environments.
  • Experience with modern monitoring and observability practices and tooling.
  • Strong understanding of observability and reliability concepts, including metrics, logs, traces, SLIs, SLOs, and alerting.
  • Good understanding of event-driven architectures and enterprise messaging platforms such as Kafka and MQ.
  • Experience troubleshooting distributed production systems, including APIs, middleware components, and message flows.
  • Understanding incident management, problem management, root cause analysis, and operational support processes.
  • Familiarity with source control, CI/CD pipelines, Infrastructure as Code (IaC), and DevOps practices.
  • Strong verbal and written communication skills with the ability to engage both technical and business stakeholders.
  • Self-motivated, detail-oriented, and capable of working independently in a fast-paced environment.

 

Preferred Qualifications

Observability & Monitoring: Grafana, Prometheus, OpenTelemetry and Loki

DevOps & Automation: Git, Ansible and CI/CD Frameworks

Container & Platform Technologies: Docker, Kubernetes

Data & Messaging Platforms: Kafka, Redis, MQ

Cloud Technologies: AWS

About Us

Jefferies is a leading global, full-service investment banking and capital markets firm that provides advisory, sales and trading, research, and wealth and asset management services. With more than 40 offices around the world, we offer insights and expertise to investors, companies, and governments.

At Jefferies, we believe that diversity fosters creativity, innovation and thought leadership through the infusion of new ideas and perspectives. We have made a commitment to building a culture that provides opportunities for all employees regardless of our differences and supports a workforce that is reflective of the communities where we work and live. As a result, we are able to pool our collective insights and intelligence to provide fresh and innovative thinking for our clients.

Jefferies is an equal employment opportunity employer, and takes affirmative action to ensure that all qualified applicants will receive consideration for employment without regard to race, creed, color, national origin, ancestry, religion, gender, pregnancy, age, physical or mental disability, marital status, sexual orientation, gender identity or expression, veteran or military status, genetic information, reproductive health decisions, or any other factor protected by applicable law. We are committed to hiring the most qualified applicants and complying with all federal, state, and local equal employment opportunity laws. As part of this commitment, Jefferies will extend reasonable accommodations to individuals with disabilities, as required by applicable law.

Similar Jobs

7 Days Ago
In-Office
2 Locations
Mid level
Mid level
Financial Services
Build and operate reliable, scalable post-trade platforms through monitoring, incident response, troubleshooting, automation, observability, and capacity planning. Develop dashboards, alerts, and service health monitoring with Grafana, Prometheus, and OpenTelemetry. Support Kafka-based event-driven systems, production releases, change management, and root-cause analysis while collaborating with engineering, infrastructure, and business stakeholders to improve resilience and reduce operational toil.
Top Skills: AnsibleAWSCi/CdDockerGitGoGrafanaInfrastructure As CodeJavaKafkaKubernetesLinuxLokiMqOpentelemetryPrometheusPythonRedisUnixWindows Server
An Hour Ago
Remote or Hybrid
India
Senior level
Senior level
Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Facilitate Scrum ceremonies, coach teams on Agile and Jira practices, track delivery metrics, remove impediments, and support product backlog management. The role also assists with project planning, risk management, stakeholder communications, and change management while developing toward full project ownership. It additionally requires information security, compliance, risk management, and Identity and Access Management experience, along with security standards implementation and awareness training.
Top Skills: AgileAtlassian JiraCi/CdDevOpsIdentity And Access ManagementJqlScrum
Senior level
Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Coordinate testing, data migration, and cutover activities for Mondelēz’s S/4HANA AMEA implementation. Develop test strategies, plans, schedules, and reporting; track defects and lead triage. Build and maintain master cutover plans, coordinate rehearsals, monitor readiness, and manage contingency planning. Serve as the central data and testing contact across markets, ensuring data governance, validation, documentation, stakeholder alignment, and go-live readiness.
Top Skills: AsanaJIRASap S/4Hana

What you need to know about the Pune Tech Scene

Once a far-out concept, AI is now a tangible force reshaping industries and economies worldwide. While its adoption will automate some roles, AI has created more jobs than it has displaced, with an expected 97 million new roles to be created in the coming years. This is especially true in cities like Pune, which is emerging as a hub for companies eager to leverage this technology to develop solutions that simplify and improve lives in sectors such as education, healthcare, finance, e-commerce and more.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account