Weekday, Inc. Logo

Weekday, Inc.

Observability Engineer

Posted 2 Days Ago
Be an Early Applicant
In-Office
Pune, Mahārāshtra, IND
Senior level
In-Office
Pune, Mahārāshtra, IND
Senior level
Lead design, implementation, and scaling of enterprise observability platforms (metrics, logs, traces). Migrate legacy monitoring to cloud-native solutions, manage observability on Kubernetes/OpenShift, build dashboards/alerts, develop Helm charts, automate workflows with Python/Bash, troubleshoot production issues, and provide architectural and team leadership to improve reliability and incident response.
The summary above was generated by AI

This role is for one of the Weekday's clients

Salary range: Rs 500000 - Rs 1700000 (ie INR 5 - 17 LPA)

Min Experience: 6+ years

Location: Chennai, Bangalore, Hyderabad, Mumbai, Pune
JobType: full-time

We are looking for a highly skilled Senior Observability Engineer to design, implement, and scale enterprise-grade observability platforms that provide deep visibility into distributed systems. This role is ideal for professionals passionate about improving system reliability, performance, and operational excellence through modern observability practices.

As a key member of the engineering team, you will lead the design and deployment of comprehensive monitoring, logging, and tracing solutions while driving the migration from traditional monitoring platforms to a modern observability ecosystem. You will collaborate closely with platform, infrastructure, and application teams to establish best practices and improve the overall health and resilience of mission-critical systems.


RequirementsKey Responsibilities
  • Design, develop, and manage scalable end-to-end observability solutions covering metrics, logs, traces, and alerting across enterprise environments.
  • Lead the migration from legacy monitoring platforms to modern observability frameworks and cloud-native monitoring solutions.
  • Deploy, administer, and optimize observability platforms running on Kubernetes or OpenShift environments.
  • Build reusable dashboards, alerts, and monitoring standards to improve operational visibility and incident response.
  • Develop and maintain Helm charts for deployment and lifecycle management of observability components.
  • Implement automation for deployment, configuration management, and operational workflows using Python or Bash scripting.
  • Collaborate with engineering and application teams to define observability standards and integrate monitoring into development workflows.
  • Analyze system performance, identify bottlenecks, and recommend improvements that enhance platform reliability and scalability.
  • Provide technical leadership, architectural guidance, and strategic recommendations for observability initiatives.
  • Support production operations by troubleshooting complex monitoring and infrastructure issues.
  • Contribute to continuous improvement initiatives and drive adoption of observability best practices across engineering teams.
RequirementsMust-Have Skills
  • Strong hands-on experience with OpenTelemetry for instrumentation and telemetry collection.
  • Expertise in the Grafana Enterprise Stack, including Mimir, Loki, and Tempo.
  • Experience administering and scaling ITRS Geneos in enterprise environments.
  • Strong knowledge of Prometheus and PromQL.
  • Hands-on experience with Grafana, including dashboard creation, alerting, and data source management.
  • Experience administering OpenShift or Kubernetes clusters.
  • Expertise in developing and managing Helm Charts for Kubernetes deployments.
  • Experience designing, deploying, and scaling enterprise observability platforms.
Good-to-Have Skills
  • Experience with Google Cloud Observability or other cloud-native monitoring solutions.
  • Automation and scripting experience using Python or Bash.
  • Familiarity with CI/CD pipelines and enterprise deployment processes.
  • Experience implementing observability in cloud-native or microservices-based architectures.
Preferred Qualifications
  • Bachelor's or Master's degree in Computer Science, Information Technology, or a related field.
  • 8–15 years of experience in Platform Engineering, Site Reliability Engineering (SRE), DevOps, Infrastructure Engineering, or Observability Engineering.
  • Proven experience implementing observability solutions at enterprise scale.
  • Strong understanding of distributed systems, container platforms, and cloud-native technologies.
Soft Skills
  • Strong analytical and problem-solving abilities.
  • Strategic thinking with the ability to influence technical direction.
  • Excellent communication and stakeholder management skills.
  • Ability to collaborate effectively across cross-functional teams.
  • Strong leadership, mentoring, and relationship-building capabilities.
  • Service-oriented mindset with a focus on operational excellence.
  • Ability to manage multiple initiatives in a fast-paced environment.

Similar Jobs

19 Days Ago
In-Office or Remote
India
Senior level
Senior level
Cloud • Information Technology • Software • Infrastructure as a Service (IaaS)
Build ingestion pipelines for logs and metrics, scalable alerting engines, and observability APIs. Interface with product teams and develop microservices using Golang and Rust.
Top Skills: AnsibleGoGraphQLGrpcRustTerraformTypescript
19 Days Ago
In-Office or Remote
India
Senior level
Senior level
Software
The Senior Infra Engineer will build and maintain ingestion pipelines, scalable alerting engines, and observability APIs, while ensuring resilience and scalability in infrastructure. They will work with tools like Golang, Rust, Terraform, and Ansible, documenting requirements and interfacing with product teams.
Top Skills: AnsibleGoGraphQLGrpcRustTerraformTypescript
5 Days Ago
In-Office
Pune, Mahārāshtra, IND
Senior level
Senior level
Internet of Things • Semiconductor
The role involves administrating and configuring monitoring tools, designing and maintaining monitoring solutions, and developing documentation and workflows. It requires extensive experience in monitoring and observability.
Top Skills: AnsibleApache SupersetAWSBashEcsElk/Efk StackJmxJSONKibanaLambdaMskN8NOpensearchOpsgeniePerlPythonRestS3SagemakerSnmpSoap ApisSyslogTcp/IpUptrendsWebhooksZabbixZenoss

What you need to know about the Pune Tech Scene

Once a far-out concept, AI is now a tangible force reshaping industries and economies worldwide. While its adoption will automate some roles, AI has created more jobs than it has displaced, with an expected 97 million new roles to be created in the coming years. This is especially true in cities like Pune, which is emerging as a hub for companies eager to leverage this technology to develop solutions that simplify and improve lives in sectors such as education, healthcare, finance, e-commerce and more.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account