Lead and build a Site Reliability Engineering function: provide technical leadership, drive incident management and MTTR SLAs, ensure alert coverage, automate reliability and performance improvements, manage project execution, and mentor a local SRE team to maintain availability and performance of mission-critical banking platforms.
About Us
Build the future of banking.
Zeta is a next-generation banking technology company providing cloud-native, fully stackable processing and core banking platforms for issuers. With a focus on scalability, compliance, and innovation, Zeta empowers financial institutions to modernize their technology infrastructure and deliver secure, seamless digital banking experiences.
Our impact runs at real-world scale. Today, over 25 million cards are live on Zeta-powered platforms across 7 countries, supported by a passionate team of 1,700+ Zetanauts across India, the US, EMEA, and Asia. Backed by SoftBank Vision Fund, Mastercard, and other reputed strategic investors, we reached a valuation of $2 billion in 2025.
Our focus is on establishing product lines that focus on key outcomes by addressing real customer pain points, modernizing legacy systems, and strengthening core fundamentals. As a result, our systems and platforms support a wide range of banking and payments capabilities, including:
1. Tachyon, our cloud-native banking stack built for population-scale systems
2. Cipher, our unified authentication platform for secure, high-volume banking environments
3. Digital Credit as a Service, enabling banks to launch credit lines on UPI
4. Elena, our intelligent and conversational AI platform for banking.
5. Pixel, India’s first digital-native credit card, launched in partnership with HDFC Bank, for whom we also revamped their PayZapp mobile app: Winner of the Celent Model Bank Award for Payments Innovation 2024.
6. Sparrow, the leading card experience for non-prime cardholders in the US
…and more across cards, payments, lending, and core banking.
We are an engineering-first organization that values ownership, bias for action, and long-term thinking. Together, we solve some of the hardest problems in banking tech. Our culture is built around trust, collaboration, and creating the conditions for you to drive impact proportionate to your potential. Reinforcing our commitment to creating an inclusive and supportive workplace, we have been consistently recognized as a Great Place to Work.
If you want to build cutting-edge banking tech that enables banks to serve millions reliably, securely, and at a population scale, Zeta is your playground.If you would like to learn more about how we have grown and evolved over the years, watch our journey here. You can also explore our website and follow us on LinkedIn, Instagram,YouTube, and X.
Responsibilities
- Establish a SRE site and help build an effective, inclusive SRE team.
- Provide technical leadership for the local team and work closely with partner team technical leads and cloud leadership.
- Provide guidance to other team members on managing availability and performance of mission critical services, on building automation to prevent problem recurrence, and building automated responses for non-exceptional service conditions.
- Manage execution of project priorities, deadlines, and deliverables.
- Lead Incident Management during Incidents.
- Responsible for driving MTTR as per the Incident SLA.
- Responsible for having 100% coverage for various alerts covering Application, Infrasture, Security, Flows etc
Skills
- Experience designing, analyzing, and troubleshooting large-scale distributed systems.
- Experience in MySQL or Postgres SQL in database.
- Hands-on experience on operating with k8s and any cloud.
- Excellent communication skills and a sense of ownership, with a systematic problem-solving approach
Experience and Qualifications:
- Bachelor’s/Master’s degree in engineering (computer science, information systems)
- 6-10 years of experience in distributed systems, storage systems, or databases, algorithms and data structures and/or Unix/Linux systems internals (e.g., filesystems, system calls) and administration
Zeta is an equal opportunity employer.
At Zeta, we are committed to equal employment opportunities regardless of job history, disability, gender identity, religion, race, marital/parental status, or another special status. We are proud to be an equitable workplace that welcomes individuals from all walks of life if they fit the roles and responsibilities.
Similar Jobs
Biotech • Pharmaceutical
Lead reliability and platform engineering for enterprise AI and machine learning ecosystems across multi-cloud environments. Design and operate ML infrastructure, observability, CI/CD, Infrastructure as Code, self-healing systems, ChatOps, and LLM, RAG, and AI Agent platforms. Establish reliability standards, automate remediation, assess technical debt, guide modernization, and mentor engineering teams. Partner with data science, AI engineering, and platform teams to deliver secure, scalable, production-ready solutions.
Top Skills:
Ai AgentsAksAmazon Sagemaker AiAWSAws CdkAzureBashChatopsCi/CdDatabricksDatadogDataikuEksEvidently AiFeastGCPGkeGoGoogle Vertex AiGrafanaInfrastructure As CodeKubeflowKubernetesLangsmithLlmsMlflowOpentelemetryPrometheusPulumiPythonRagRagasSlmsTerraformWeights & Biases
Artificial Intelligence • Cloud • Information Technology • Automation
Lead SRE responsible for platform reliability, incident management, automation, observability, and production stability across cloud-native and Kubernetes environments. Mentor teams, run RCA, define SLIs/SLOs, automate operations, manage Apigee APIs, and support CI/CD and IaC initiatives to improve availability and performance.
Top Skills:
AppdynamicsBashCi/CdDatadogDynatraceElk StackGitGoGoogle ApigeeGoogle Cloud Platform (Gcp)Google Kubernetes (K8S)GrafanaInfrastructure As Code (Iac)IstioJenkinsKubernetesLinux/UnixNginxPrometheusPythonSplunkTerraform
Software
Lead reliability, scalability, and performance of Azure-hosted platforms; drive SRE practices, automation, observability (SLOs/SLIs), incident response, disaster recovery, and onboarding. Mentor engineers, run solutions workshops, influence architecture, and provide on-call support.
Top Skills:
AnsibleAPIsArmBicepCi/CdDockerItilKubernetesAzureObservabilityScriptingSimcorp DimensionSlisSlosSQLTerraform
What you need to know about the Pune Tech Scene
Once a far-out concept, AI is now a tangible force reshaping industries and economies worldwide. While its adoption will automate some roles, AI has created more jobs than it has displaced, with an expected 97 million new roles to be created in the coming years. This is especially true in cities like Pune, which is emerging as a hub for companies eager to leverage this technology to develop solutions that simplify and improve lives in sectors such as education, healthcare, finance, e-commerce and more.



