SigNoz Jobs

Sr Site Reliability Engineer

SigNoz

Sr Site Reliability Engineer

Reposted 11 Days Ago

Be an Early Applicant

Remote

Hiring Remotely in India

Senior level

Remote

Hiring Remotely in India

Senior level

Own reliability, scalability, and operability of a petabyte-scale observability SaaS. Improve SLOs/SLIs, incident response, and on-call practices; scale and tune ingest pipelines and ClickHouse; manage Kubernetes clusters, autoscaling, multi-tenancy, and upgrades; build infra-as-code, CI/CD, capacity planning, and observability for the platform.

The summary above was generated by AI

About SigNoz

SigNoz is an open-source observability platform that helps modern engineering teams monitor, debug, and optimize their applications with deep visibility into metrics, traces, and logs — all in one place. We're built natively on OpenTelemetry and offer both self-hosted and cloud options, so teams can run observability the way they want, without vendor lock-in.

We are growing fast and building core developer infra products. And we are not fooling around:

27,000+ GitHub stars
800+ customers
7,000+ members in our Slack community

Role: Sr Site Reliability Engineer (SRE)

We're looking for an SRE to own the reliability, scalability, and operability of the SigNoz cloud platform. You'll keep a petabyte-scale observability system fast and dependable — making sure the people who trust us to watch their systems can always trust ours. The platform team handles infra, scalability of SaaS, ingest pipelines, staging environments, automation, and the operational backbone of the product.

This is a deeply hands-on role for someone who understands what actually breaks in production at scale — and enjoys fixing it for good.

What we're looking for

Kubernetes at scale — not just "I've deployed to k8s," but real fluency with the nuances and gotchas: resource tuning, autoscaling behavior, networking, stateful workloads, upgrades, and the failure modes that only show up under load
Working knowledge of ClickHouse — operating it, tuning queries, and understanding its behavior at scale — is a strong plus
Knowledge of Golang is a plus (most of our stack and tooling is in Go)
Familiarity with OpenTelemetry and running large-scale data ingest pipelines is a plus

What you'll work on

You'll work with a high-caliber team across areas like:

Reliability of the SigNoz cloud platform: SLOs/SLIs, error budgets, incident response, and on-call practices that don't burn people out
Scaling the ingest path — making it robust to bursts while maintaining data freshness
SaaS auto-scalability and capacity planning across a petabyte-scale system
Operating and tuning ClickHouse and the data layer for performance and cost
Kubernetes infrastructure: cluster operations, upgrades, multi-tenancy, and the automation that keeps it boring
Observability of SigNoz itself — we dogfood our own product, so you'll help make it world-class
Infrastructure-as-code, CI/CD, and the tooling that lets a small team operate big systems

What will make you successful

5–8 years in SRE, infrastructure, or platform/backend roles operating production systems at scale
Deep, practical Kubernetes experience — you know where the bodies are buried
Strong grasp of distributed systems failure modes, performance debugging, and capacity planning
Comfortable in code (Go preferred) — you automate and fix things, not just configure them
Loves open source — ideally with prior contributions to OSS projects (any size)
Comfortable in a high-ownership, fast-moving, remote-first environment
Strong communication — can write clear runbooks and tech docs and explain trade-offs

Nice-to-haves

Past experience on platform/infra/SRE teams of Series B+ startups
Hands-on experience operating ClickHouse, Kafka, or similar high-throughput data systems
Experience in observability (monitoring / logging / tracing) and with OpenTelemetry

Why you'll love working at SigNoz

Work on a globally used open-source project that engineers actually love
Huge scope and ownership — your work directly shapes how teams adopt SigNoz
Collaborate with a high-caliber team who just can't stop shipping
Remote-first, async-friendly culture
Opportunity to help define the future of open-source observability

Similar Jobs

LSEG (London Stock Exchange Group)

Senior Site Reliability Engineer

9 Days Ago

Remote

Shri Bhrigukshetra, BLR, Uttar Pradesh, IND

Senior level

Fintech • Analytics

Design, build, and maintain reliable infrastructure platforms; drive automation with Python, Ansible, and Terraform; implement CI/CD and IaC; enhance observability, monitoring, and alerting; lead incident response, RCA, and service restoration; apply ITIL practices for incident, change, and problem management.

Top Skills: AlertingAnsibleAWSAzureCi/CdClickhouseCriblGrafanaInfrastructure As Code (Iac)MonitoringObservabilityOpentelemetryPythonTerraform

NVIDIA

Senior Site Reliability Engineer

16 Days Ago

In-Office or Remote

India

Senior level

Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse

Operate and improve the reliability, availability, and performance of large-scale GeForce NOW services. Participate in incident triage and on-call rotations, build automation and tooling, enhance observability (metrics/logs/traces), drive SLO/SRI practices, run postmortems, and design/operate Kubernetes-based services across cloud and datacenter environments.

Top Skills: AWSAzureBashContainerizationElk/OpensearchGCPGoGrafanaKubernetesMicroservicesOpentelemetryPrometheusPython

Akamai Technologies

Senior Site Reliability Engineer

17 Days Ago

In-Office or Remote

India

Senior level

Cloud • Security • Software • Cybersecurity

Design, implement, and maintain reliable, scalable infrastructure for large distributed content delivery systems. Define and measure SLIs/SLOs, monitor availability and performance, troubleshoot incidents, and implement corrective actions. Develop automation to reduce manual work, participate in design reviews, and collaborate with product and engineering teams to improve system reliability and performance.

Top Skills: AdbmsBashCloud ComputingDatadogGrafanaJavaScriptOracle SqlPrometheusPythonUnix/Linux

What you need to know about the Pune Tech Scene

Once a far-out concept, AI is now a tangible force reshaping industries and economies worldwide. While its adoption will automate some roles, AI has created more jobs than it has displaced, with an expected 97 million new roles to be created in the coming years. This is especially true in cities like Pune, which is emerging as a hub for companies eager to leverage this technology to develop solutions that simplify and improve lives in sectors such as education, healthcare, finance, e-commerce and more.