Weekday, Inc. Logo

Weekday, Inc.

Site Reliability Engineer

Posted 2 Days Ago
Be an Early Applicant
In-Office
Pune, Mahārāshtra, IND
Senior level
In-Office
Pune, Mahārāshtra, IND
Senior level
Manage, deploy, and optimize Apache Kafka clusters and large-scale streaming platforms. Monitor performance, troubleshoot production incidents, automate operational tasks, implement capacity planning and disaster recovery, ensure security and compliance, create runbooks, and participate in on-call incident response to maintain highly available, scalable platform infrastructure.
The summary above was generated by AI

๐—ง๐—ต๐—ถ๐˜€ ๐—ฟ๐—ผ๐—น๐—ฒ ๐—ถ๐˜€ ๐—ณ๐—ผ๐—ฟ ๐—ผ๐—ป๐—ฒ ๐—ผ๐—ณ ๐˜๐—ต๐—ฒ ๐—ช๐—ฒ๐—ฒ๐—ธ๐—ฑ๐—ฎ๐˜†'๐˜€ ๐—ฐ๐—น๐—ถ๐—ฒ๐—ป๐˜๐˜€

๐—ฆ๐—ฎ๐—น๐—ฎ๐—ฟ๐˜† ๐—ฟ๐—ฎ๐—ป๐—ด๐—ฒ: ๐—ฅ๐˜€ ๐Ÿญ๐Ÿฎ๐Ÿฌ๐Ÿฌ๐Ÿฌ๐Ÿฌ๐Ÿฌ - ๐—ฅ๐˜€ ๐Ÿฎ๐Ÿด๐Ÿฌ๐Ÿฌ๐Ÿฌ๐Ÿฌ๐Ÿฌ (๐—ถ๐—ฒ ๐—œ๐—ก๐—ฅ ๐Ÿญ๐Ÿฎ-๐Ÿฎ๐Ÿด ๐—Ÿ๐—ฃ๐—”)

Experience: 6+ yrs

Location: Hyderabad, Bengaluru, Pune, Chennai, Tamil Nadu, India, Mumbai, Maharashtra, India

Job Type: Full-time

We are seeking an experienced Kafka Platform Engineer with strong expertise in distributed systems, large-scale messaging platforms, and production operations. This role is ideal for professionals who are passionate about building, maintaining, and optimizing highly available streaming infrastructure while ensuring reliability, scalability, and operational excellence across enterprise environments.

As a Kafka Platform Engineer, you will be responsible for managing mission-critical messaging platforms, improving platform performance, automating operational processes, and supporting production environments. You will collaborate with infrastructure, application, DevOps, and engineering teams to deliver resilient streaming solutions, troubleshoot complex production issues, and continuously enhance platform reliability through automation, monitoring, and best practices.


RequirementsKey Responsibilities
  • Design, deploy, manage, and optimize Apache Kafka clusters and large-scale messaging or streaming platforms.
  • Monitor platform health, system performance, and resource utilization using modern monitoring, logging, and alerting tools.
  • Troubleshoot production incidents, identify root causes, and implement long-term solutions to improve system stability.
  • Automate operational tasks, deployments, and maintenance activities using scripting and infrastructure automation techniques.
  • Collaborate with application and infrastructure teams to support messaging architecture, integrations, and production workloads.
  • Implement performance tuning, capacity planning, and scalability improvements for distributed systems.
  • Maintain high availability, fault tolerance, and disaster recovery strategies for messaging infrastructure.
  • Ensure platform security, system compliance, and operational best practices across production environments.
  • Develop operational documentation, runbooks, and knowledge-sharing resources to improve support efficiency.
  • Participate in on-call support, incident response, and continuous improvement initiatives to maintain service reliability.
What Makes You a Great Fit
  • 6+ years of experience managing distributed systems, production infrastructure, or platform engineering environments.
  • Strong hands-on experience with Apache Kafka or large-scale messaging and event streaming platforms.
  • Deep understanding of distributed systems architecture, scalability, fault tolerance, and production operations.
  • Experience with monitoring, logging, alerting, and observability tools for enterprise infrastructure.
  • Proficiency in at least one scripting or programming language such as PythonBash, or Java.
  • Strong knowledge of Linux system administration, networking fundamentals, and troubleshooting methodologies.
  • Experience automating operational workflows and improving platform reliability through scripting and infrastructure automation.
  • Excellent analytical, debugging, and problem-solving skills with a proactive operational mindset.
  • Strong communication and collaboration skills with the ability to work effectively across cross-functional engineering teams.
  • Passion for building reliable, secure, and highly available platform infrastructure while continuously improving operational excellence.

Similar Jobs

Yesterday
In-Office
2 Locations
Mid level
Mid level
Financial Services
Operate and support AWS EKS Kubernetes platform, handle incident response and on-call duties, automate infrastructure with Terraform and GitOps, improve observability and SLOs, troubleshoot cloud-native issues, create runbooks, and collaborate with architecture, networking, and security teams to reduce toil and ensure platform reliability.
Top Skills: ArgocdAutoscalingAws Ec2Aws IamAws VpcBambooBitbucketChatgptClaudeConfluenceDnsEksGitGithub ActionsGithub CopilotGitlabGoGrafanaIngress ControllersJenkinsJIRAKubernetesOpentelemetry (Otel)PrometheusPythonRbacService MeshShell ScriptingSplunkStorage ClassesTerraformTerraform Enterprise
8 Days Ago
Hybrid
Pune, Mahārāshtra, IND
Mid level
Mid level
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Operate and improve production multi-cloud platform services: deploy and maintain Kubernetes clusters, build CI/CD and GitOps pipelines, create self-service control planes, manage capacity and DR, implement observability and security, own on-call/incidents, automate repetitive work, and enable consuming engineering teams.
Top Skills: AIArgo WorkflowsArgocdCrossplaneFluxGithub ActionsGitopsGoGrafanaIstioJaegerJenkinsKubernetesLinkerdOpentelemetryPrometheusPulumiTektonTemporalTerraform
5 Days Ago
In-Office or Remote
2 Locations
Senior level
Senior level
Cloud • Enterprise Web • Hardware • Information Technology • Internet of Things • Robotics • Semiconductor
Build, automate, and operate a global cloud platform: develop automation in Go/Python, manage large-scale EKS clusters (Karpenter), author Terraform and Helm IaC, lead incident response and post-mortems, define SLIs/SLOs, implement observability (Datadog/Prometheus/Grafana) and PagerDuty on-call, and develop secure self-service tools to meet SOC2. Night-shift role based in Ahmedabad, India.
Top Skills: Amazon EksAWSCachingDatadogDynamoDBGoGrafanaHelmKafkaKarpenterKubernetesMskPagerdutyPrometheusPythonTerraform

What you need to know about the Pune Tech Scene

Once a far-out concept, AI is now a tangible force reshaping industries and economies worldwide. While its adoption will automate some roles, AI has created more jobs than it has displaced, with an expected 97 million new roles to be created in the coming years. This is especially true in cities like Pune, which is emerging as a hub for companies eager to leverage this technology to develop solutions that simplify and improve lives in sectors such as education, healthcare, finance, e-commerce and more.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account