CloudRaft Logo

CloudRaft

Senior SRE

Posted 9 Hours Ago
Be an Early Applicant
Remote
Hiring Remotely in India
Senior level
Remote
Hiring Remotely in India
Senior level
Design, build, operate, and scale cloud-native infrastructure and Kubernetes clusters across major clouds. Implement CI/CD, observability, IaC, SLIs/SLOs, automate with Go/Python, participate in 24x7 on-call, and contribute to open-source and technical knowledge sharing.
The summary above was generated by AI
About CloudRaft
CloudRaft is a premier cloud-native consulting and engineering company that helps ambitious startups and digital-first organizations build, scale, and operate mission-critical platforms. We partner with innovators at the forefront of artificial intelligence, developer productivity, observability, digital commerce, and enterprise software—enabling them to accelerate growth with resilient, scalable, and production-ready cloud infrastructure.

Our experience spans organizations developing AI safety and governance platforms, AI Cloud, AI agent ecosystems, developer tooling, observability solutions, digital health products, customer engagement platforms, and technology-driven franchise networks. By combining deep expertise in Platform Engineering, Kubernetes, DevOps, Observability, and Cloud Native technologies, CloudRaft helps high-growth companies move faster, operate more reliably, and focus on building category-defining products.


Job Description
We are looking for passionate Site Reliability Engineers (SREs) to join our growing team. In this role, you will take end-to-end ownership of designing, building, operating, and scaling mission-critical infrastructure for our partners. You will be responsible for ensuring reliability, performance, security, and operational excellence while driving automation, improving system efficiency, and implementing innovative solutions. Working at the intersection of software engineering and operations, you will help create resilient platforms that enable fast-growing organizations to scale with confidence.


Responsibilities
  • Manage and maintain Kubernetes clusters across cloud platforms, including OpenShift, Amazon EKS, Azure AKS, and Google GKE.
  • Implement and manage CI/CD pipelines using tools such as Jenkins, GitHub Actions, Argo CD, or GitLab CI/CD.
  • Design and maintain observability stacks with tools including Prometheus, Grafana, Loki, OpenTelemetry, and related technologies. Be part of the team who support open source projects like Prometheus, Thanos, Mimir, CloudNativePG, Istio and more.
  • Optimize system performance and resolve production issues. Be part of the on call roster to provide 24x7 coverage for the critical production systems.
  • Implement SRE principles, including Service Level Indicators (SLIs) and Service Level Objectives (SLOs), to uphold system reliability.
  • Automate infrastructure and operational tasks using programming languages such as Go or Python, and Infrastructure as Code (IaC) tools like Terraform.
  • Apply agentic AIto automate the SDLC lifecycle, AIOps and automation.
  • Learn about emerging technologies, including AI, GPU Infrastructure
  • Contribute to knowledge sharing through technical writing and presentations.

Qualifications
  • Bachelor’s degree in Computer Science, Information Technology, or a related field.
  • 5+ years of experience in SRE, Platform Engineering, or DevOps Engineer.
  • Strong expertise in Kubernetes, cloud-native technologies, on-premise and major cloud platforms (AWS, Azure, GCP).
  • Proficiency in programming languages such as Python or Go or Node.js.
  • Familiarity with CI/CD tools and modern deployment practices.
  • Proficiency in one or more open source observability stacks and Infrastructure as Code (Terraform/Pulumi).
  • CKA/CKAD Certified (Brownie points!)
  • Excellent problem-solving abilities and communication skills.
  • Inclination toward open-source contributions is advantageous.

Benefits : 
- Competitive salary
- Premium health insurance and various health & wellness benefits from a leading insurance provider through Plum
- Opportunity to work on the latest AI stack and GPU infrastructure
- Collaborative and supportive work environment full of learning
- Chance to take a front seat where you lead and deliver

Similar Jobs

5 Days Ago
In-Office or Remote
Senior level
Senior level
Software
Build and maintain reliable, automated cloud-native infrastructure. Manage Kubernetes lifecycles, implement IaC with Terraform, develop automation in Go/Python, implement monitoring and SLI/SLOs, lead moderate incidents, optimize performance, and mentor junior engineers.
Top Skills: GoGrafanaKubernetesLinuxPrometheusPythonTerraform
10 Days Ago
In-Office or Remote
2 Locations
Senior level
Senior level
Cloud • Enterprise Web • Hardware • Information Technology • Internet of Things • Robotics • Semiconductor
Build, automate, and operate a global cloud platform: develop automation in Go/Python, manage large-scale EKS clusters (Karpenter), author Terraform and Helm IaC, lead incident response and post-mortems, define SLIs/SLOs, implement observability (Datadog/Prometheus/Grafana) and PagerDuty on-call, and develop secure self-service tools to meet SOC2. Night-shift role based in Ahmedabad, India.
Top Skills: Amazon EksAWSCachingDatadogDynamoDBGoGrafanaHelmKafkaKarpenterKubernetesMskPagerdutyPrometheusPythonTerraform
11 Days Ago
In-Office or Remote
Senior level
Senior level
Internet of Things • Mobile • Retail
Lead platform reliability and DevOps automation: implement CI/CD with GitHub Actions, automate JFrog/Helm and image migrations, enable microservices deployments, and operate observability and logging stacks. Provide Tier 3 troubleshooting and incident leadership, manage cloud infrastructure governance, capacity and DR planning, cost/license governance, and maintain SOPs and reliability best practices.
Top Skills: AirflowAlertmanagerApache FlinkAws MskAzure Container Registry (Acr)Azure Event HubAzure Kubernetes Service (Aks)Azure MonitorConfluent CloudConfluent KafkaFluentbitGithub ActionsGrafanaHelmJavaJfrogKubernetesOpensearchPostgresPrometheusPythonReactSpring BootThanos

What you need to know about the Pune Tech Scene

Once a far-out concept, AI is now a tangible force reshaping industries and economies worldwide. While its adoption will automate some roles, AI has created more jobs than it has displaced, with an expected 97 million new roles to be created in the coming years. This is especially true in cities like Pune, which is emerging as a hub for companies eager to leverage this technology to develop solutions that simplify and improve lives in sectors such as education, healthcare, finance, e-commerce and more.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account