Akamai Technologies Logo

Akamai Technologies

Senior Site Reliability Engineer

Posted 3 Days Ago
Be an Early Applicant
In-Office or Remote
Hiring Remotely in India
Senior level
In-Office or Remote
Hiring Remotely in India
Senior level
Senior Site Reliability Engineer responsible for designing and maintaining scalable infrastructure, deployment processes, monitoring platforms, and automation for distributed systems. The role improves reliability, performance, observability, and operational efficiency through Linux administration, scripting, infrastructure as code, CI/CD, Kubernetes, cloud platforms, and monitoring tools. The engineer collaborates with internal teams to resolve technical challenges, safely deploy software, and drive continuous improvement.
The summary above was generated by AI

Are you passionate about cutting edge technology?

Would you like an opportunity to effect change at a leading technology organization?

Join our Site Reliability team

The Akamai Cloud Technology Engineering team owns, develops and manages the solutions used by our engineers globally. These solutions help to run one of the largest distributed systems in the world. We work closely with internal teams and stakeholders to create innovative, powerful, scalable, highly reliable and secure systems at scale.

Partner with the best

This Senior Site Reliability Engineer (Linux) role involves improving automation and efficiency for internal teams while ensuring operational excellence. Responsibilities include enhancing system reliability, scalability, and performance by designing and maintaining infrastructure, tools, and processes. Collaborate to address technical challenges, optimize deployments, and support applications. Drive continuous improvement through workflow automation, system monitoring, and performance optimization to achieve organizational objectives.

As a Senior Site Reliability Engineer, you will be responsible for:

  • Developing processes, plans, and infrastructure to deploy new software components and updates safely and efficiently at scale
  • Improving our system monitoring and analysis platform to speed error detection and remediation, enhancing performance and reliability
  • Improving our system monitoring and analysis platform to speed error detection and remediation, enhancing performance and reliability.

Do what you love

To be successful in this role you will:

  • Have 5+ years of relevant experience and a Bachelors' degree in Computer Science or related field
  • Specialize in Linux administration, demonstrating expertise in Python and Bash scripting languages for advanced systems management and automation tasks.
  • Utilize Salt Stack, Ansible, Terraform for infrastructure automation, alongside CI/CD tools including Jenkins for streamlined deployment and management.
  • Demonstrate expertise with observability or monitoring tools like Prometheus, Grafana, ELK/Open Search, Datadog, and Splunk, Docker
  • Have hand-on mastery with Kubernetes experience with any cloud platform, such as AWS, GCP, Azure, or an equivalent alternative

About us

At Akamai, we make life better for billions of people, trillions of times a day.
Whether you're streaming live events, scrolling social media, watching your favorite series, or managing your savings, we're the engine behind the scenes. We provide the world's most distributed platform from Cloud to Edge to help the giants of the digital world work faster and stay more secure, making the internet a better experience for everyone.
Our focus is simple:
Cloud and Edge: Running apps closer to users for instant performance.
Security: Neutralizing threats before they ever reach your data.
Content Delivery: Scaling the world's biggest moments without a glitch.
AI: Enabling our customers to build, secure, and scale AI apps on the world's most distributed cloud platform.
At Akamai, we don't just support the internet; we power and protect it, because behind every great digital experience is a massive hidden challenge. And we're the ones who solve it. When millions of people hit play or pay, Akamai ensures it just works.

Benefits at Akamai: We support your health, well-being, finances, and life beyond work. See our benefits.

FlexBase adapts to your job's needs

Akamai's FlexBase program is yet another way we show our commitment to providing employees with an exceptional workplace experience. It's not about telling employees where to work; it's about supporting employees to do their best work.
We trust our incredible employees to work in ways that suit them best: at home, in an office, or a combination of both.

Connect with us on social and see what life at Akamai is like!

Similar Jobs

8 Days Ago
Easy Apply
In-Office or Remote
Easy Apply
Senior level
Senior level
Cloud • Information Technology • Security • Software
Architects and operates highly available, multi-region cloud infrastructure, Kubernetes platforms, microservices, and authentication systems. Leads disaster recovery, failover automation, observability, SLOs, incident response, FinOps, infrastructure-as-code, and reliability improvements. Develops automation in Python or Go, manages GitOps deployments, mentors engineers, authors technical documentation, and participates in on-call rotations.
Top Skills: Amazon EksArgo CdAWSAws Secrets ManagerCert-ManagerClaude CodeCursorDatadogExternal Secrets OperatorGCPGithub CopilotGkeGoHaproxyIstioKargoKubernetesLinkerdNginxPagerdutyPythonTerraformVault
One Month Ago
Remote or Hybrid
Senior level
Senior level
Digital Media • eCommerce • Gaming • Mobile • News + Entertainment
Lead reliability, scalability, observability, automation, infrastructure, disaster recovery, and security initiatives for Crunchyroll’s cloud-native data platforms. Establish SRE practices including SLIs, SLOs, error budgets, incident management, and postmortems. Operate Kubernetes and GCP environments, implement Infrastructure as Code, optimize capacity and performance, and drive vulnerability remediation, penetration-testing support, and cloud platform security.
Top Skills: Ci/CdDatadogGCPGoGrafanaIdentity And Access ManagementInfrastructure As CodeJavaKubernetesLinuxOpentelemetryOwasp Top 10PrometheusPythonShellTerraform
19 Days Ago
In-Office or Remote
Senior level
Senior level
Software
Owns Kubernetes-based development, CI, pre-production, and customer-facing production environments. Responsibilities include SRE operations, incident response, on-call support, Helm and CI/CD lifecycle management, infrastructure automation, observability, security, disaster recovery, stateful platform services, and multi-region deployments. The role also provides technical consultation, creates operational documentation, validates upgrades, and mentors engineers.
Top Skills: AnsibleApi GatewaysArgo CdAWSCluster ApiDockerFluxGithub ActionsGitopsGoGrafanaHelmKafkaKeycloakKindKubernetesKyvernoMetal3OpaOpenstackOpentelemetryPostgresPrometheusPytestPythonTemporalTerraform

What you need to know about the Pune Tech Scene

Once a far-out concept, AI is now a tangible force reshaping industries and economies worldwide. While its adoption will automate some roles, AI has created more jobs than it has displaced, with an expected 97 million new roles to be created in the coming years. This is especially true in cities like Pune, which is emerging as a hub for companies eager to leverage this technology to develop solutions that simplify and improve lives in sectors such as education, healthcare, finance, e-commerce and more.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account