NVIDIA Logo

NVIDIA

Senior DevOps Engineer

Posted Yesterday
Be an Early Applicant
In-Office
Pune, Maharashtra
Senior level
In-Office
Pune, Maharashtra
Senior level
NVIDIA seeks a Senior DevOps Engineer to architect and maintain Kubernetes-based infrastructure, manage large-scale cloud applications, and develop AI automation tools.
The summary above was generated by AI

NVIDIA is looking for an outstanding engineering Architect to join its Software Infrastructure and Operations team. The position will be part of a fast-paced crew that develops and maintains sophisticated Kubernetes based development, compute and test environments for a multitude of platforms including Windows and Linux. You will be working with a team of passionate and skilled engineers that are continuously working to provide better tools to build and manage this infrastructure. With your help we would forge the next generation of compute infrastructure multiplying the power of the CPU, GPU and DPU for the age of AI. We need a motivated, hardworking and focused individual who has a real passion for operational excellence, Infrastructure services, and automation.

What you’ll be doing:

  • Architect the scaling operation in our data centers. Deploy and Support end-to-end container management solution with Kubernetes, Docker, containerd. Design solutions with service discovery, networking, monitoring, logging, scheduling in Kubernetes.

  • Setup and Manage end to end Compute Infrastructure using PaaS & IaaS services - tools, plugins, nodes, user management, back up, restore, monitoring, etc. Design and develop AI tools needed for automating maintenance of 35000+ hosts with only 12 support engineers.

  • Design and build sophisticated automations and AI powered applications.

  • Use your depth in algorithms and system software background!

  • Work in teams to deploy new data center infrastructure.

  • Plan and implement critical metrics tracking using various data analytics mining methods and dashboards.

  • Reuse AI techniques to extract useful signals about machines and jobs from the data generated!

  • Take part in prototyping, crafting and developing cloud infrastructure for Nvidia.

What we need to see:

  • Strong Kubernetes understanding and background especially on-premises setup and extensive experience with Kubernetes components & subsystems.

  • Experience of maintaining large scale cloud/on-prim infrastructure applications using Kubernetes, Slurm and Open Stack

  • Proven programming background in python/Golang/java and/or relevant scripting languages

  • Excellent debugging and analytical skills and experience in Databases both SQL (MySQL ) and NoSQL (Elastic Search /MongoDB)

  • Proficient with configuration management tools like Ansible, Chef, Puppet and strong experience with Jenkins and/or other CI systems.

  • Hands-on experience with VMs, Dockers, Kubernetes Cluster.

  • Experience with analytics/visualization tools like Kibana, Grafana, Splunk etc. and experience with monitoring systems such as Zabbix and/or Nagios is nice to have

  • 10+ years of proven experience

  • Bachelors or Master's Degree or equivalent experience in CS, Software Engineering, or related field.

Ways to stand out from the crowd:

  • Previous experience with DevOps/SRE teams

  • Thrives in a multi-tasking environment with constantly evolving priorities and documents work well

  • Outstanding collaboration skills across organizational boundaries, experience with using and improving data centers and with computer algorithms and ability to choose the best possible algorithms to meet the scaling challenge

  • Ability to divide complex problems into simple sub problems and then reuse available solutions to implement most of those

  • Experience with designing simple systems that can work reliably without needing much support

Top Skills

Ansible
Chef
Containerd
Docker
Elastic Search
Go
Grafana
Java
Jenkins
Kibana
Kubernetes
MongoDB
MySQL
Nagios
NoSQL
Puppet
Python
Splunk
SQL
Zabbix

NVIDIA Pune, Mahārāshtra, IND Office

Survey No.144 145, Commerzone No.5, Off, Airport Rd, Yerawada, Pune, Maharashtra, India, 411006

Similar Jobs

6 Days Ago
Easy Apply
Hybrid
Pune, Maharashtra, IND
Easy Apply
Senior level
Senior level
Artificial Intelligence • Cloud • Information Technology • Machine Learning • Software
The Senior DevOps Engineer will enhance cloud infrastructure, manage multi-cloud environments, automate infrastructure with IaC, optimize costs, and strengthen security, focusing on the Edwin AI team at LogicMonitor.
Top Skills: AWSAzureBashCi/CdGCPGrafanaKubernetesPrometheusPythonTerraform
2 Days Ago
In-Office
Pune, Maharashtra, IND
Senior level
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
The role involves developing and maintaining Build systems for different OS, optimizing workflows, and collaborating with engineering teams to enhance build stability.
Top Skills: AndroidBazelCmakeGitGnu MakeJenkinsNinjaPerforcePythonUnixYocto
16 Days Ago
In-Office or Remote
Mumbai, Maharashtra, IND
Senior level
Senior level
Software • Consulting
The role involves managing DevOps practices, utilizing tools like Terraform and GitLab CI, monitoring cloud infrastructure, and ensuring high-quality software delivery.
Top Skills: AcmAlbApi GatewayAWSAws BackupAws ShieldCloudwatchDatadogDynamoDBEcs FargateElastic Container RegistryElasticache For RedisGitlabGitlab CiLambdaRds Aurora PostgresqlRoute53Secrets ManagerTerraformVisual Studio Code

What you need to know about the Pune Tech Scene

Once a far-out concept, AI is now a tangible force reshaping industries and economies worldwide. While its adoption will automate some roles, AI has created more jobs than it has displaced, with an expected 97 million new roles to be created in the coming years. This is especially true in cities like Pune, which is emerging as a hub for companies eager to leverage this technology to develop solutions that simplify and improve lives in sectors such as education, healthcare, finance, e-commerce and more.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account