Neysa Logo

Neysa

Senior Cloud Engineer

Posted 18 Days Ago
Be an Early Applicant
In-Office
Mumbai, Maharashtra
Senior level
In-Office
Mumbai, Maharashtra
Senior level
Design, deploy, automate, and operate highly available cloud infrastructure across private and public environments. Responsibilities include managing virtualization, storage, networking, Kubernetes, scaling, infrastructure automation, failure testing, security, cost optimization, and production troubleshooting across complex multi-cloud platforms.
The summary above was generated by AI

About the Role

We are building a next-generation cloud platform designed to power AI, enterprise, and high-performance computing (HPC) workloads. As a Senior Cloud Infrastructure Engineer, you will be responsible for designing, deploying, automating, and operating highly available compute, virtualization, storage, networking, and Kubernetes infrastructure across both private and public cloud environments.

In this role, you will work across the full infrastructure stack—from Linux systems, virtualization, storage, and networking to Kubernetes orchestration and infrastructure automation- ensuring our platform is scalable, resilient, secure, and operationally efficient.

You will collaborate closely with engineering, and customer-facing teams to deliver reliable cloud infrastructure, optimize performance, automate deployments, and troubleshoot complex production environments.

This is a hands-on engineering role for someone who enjoys solving complex infrastructure challenges, building automation-first platforms, and operating mission-critical cloud infrastructure at scale.

What you will be doing:

·       Work with multiple cloud providers (AWS, GCP, Azure and/or your own!) to create an infrastructure that seamlessly stitches everything together.

·       Build and maintain automatic scaling and reducing algorithms for infrastructure and software services.

·       Identify points of failure and design around them. Create and simulate failure scenarios to test and verify your design choices.

·       Work deeply with ancillary technologies, viz. networking, storage, load balancing, DNS, DHCP, logging, etc.

·       Interact with developers to optimize placement, understand cost structures to optimize expenses, and design bespoke infrastructure deployment rules.

What we need to see:

Core virtualization and Storage (10+ years)

·       Have a (very) deep understanding of virtualization, cloud computing, containers and automation across KVM, VMware, AWS, GCP and Azure.

·       Have a strong understanding of Storage Systems i.e. Block, iSCSI, NFS, FC (Fibre Channel) and Scale-Out.

·       Have intermediate or advance level understanding of Container Orchestration like Kubernetes, Docker Swarm and Nomad.

·       Be intimate with security concepts like firewalling, API security, placement policies, load balancing and availability.

·       Be an expert in automation and automation frameworks. You should be able to design, write and troubleshoot automation scripts across platforms. Which include (but not limited to) Bash (Shell), PowerShell, Python, Terraform

·       Understand how infrastructure components like DHCP, DNS, messaging queueing mechanisms, logging etc. work.

·       Have a strong, intuitive approach to troubleshooting complex, multi-domain issues.

·       Have experience on cloud networking and extranets.

·       Have an understanding of encryption, its effect on infrastructure services, and its suitability.

Ways to stand out from the rest:

·       The initiative to work independently, at your own pace, but on a schedule

·       Intuitive about what is a problem now and what could be a problem down the line

·       Kubernetes certifications (CKA, CKS) or Red Hat Certified Engineer (RHCE)

·       Good understanding of Networking

·       Linux certifications

Minimum Qualifications:

·       Bachelor's degree in Computer Science, Electrical Engineering, Electronics, Information Technology, or equivalent professional experience

·       4+ years of Linux systems administration and data center deployment

·       Soft Skills:

·       Strong problem-solving and debugging abilities across hardware, kernel, and application layers

·       Ownership mindset with accountability for deployment quality and customer success

·       Cross-functional collaboration with other teams

·       Proactive approach to continuous learning and staying current with AI infrastructure trends

·       Strong documentation and presentation skills, able to defend design decisions amongst peers.

 

Similar Jobs

7 Days Ago
Hybrid
Kharadi, Pune, Maharashtra, IND
Senior level
Senior level
Food • Information Technology • Logistics • Retail
Design, provision, deploy, and operate resilient GCP infrastructure supporting AI and automation workloads. Use Terraform to automate repeatable deployments, maintain production systems, monitor services, respond to incidents, and improve platform reliability. The role also contributes to cloud governance and compliance, collaborates with stakeholders and vendors, and may mentor less experienced team members.
Top Skills: Ai/Ml PlatformsCloud Monitoring And Observability ToolsGoogle Cloud Platform (Gcp)PythonTerraform
7 Days Ago
Hybrid
Pune, Maharashtra, IND
Senior level
Senior level
Artificial Intelligence • Cloud • Fintech • Information Technology • Analytics • Financial Services • Cybersecurity
Provides L1 production support for financial applications and enterprise data platforms. Responsibilities include monitoring ETL jobs, data pipelines, batch schedules, alerts, and incidents; troubleshooting data quality, reconciliation, ingestion, reporting, and application issues; performing SQL-based data updates, smoke testing, root-cause analysis, and impact assessment; coordinating with technical and business teams; escalating complex issues; supporting disaster recovery, audits, automation, and continuous service improvement.
Top Skills: ApvAzure Data FactoryCatchpointCloud Data ServicesControl-MDatabricksDynatraceETLItilMqPythonServicenowShell ScriptingSQLUnix/Linux
11 Days Ago
In-Office or Remote
India
Senior level
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Design and ship Java, Go, and Rust services for NVIDIA Cloud Functions, a distributed platform routing AI workloads across GPU fleets. Improve performance, reliability, scalability, and cloud-native build and release processes. Collaborate across NVIDIA technologies, contribute to an open-source project, triage community issues, review pull requests, and write documentation. The role requires expertise in systems programming, distributed architecture, Kubernetes, containerization, Linux internals, scripting, and continuous integration.
Top Skills: ArgocdBashCloud Service ProvidersGitlabGoGrpcHttp/2JavaKubernetesKubernetes Custom ResourcesKubernetes OperatorsLinuxMessage QueuesPub-Sub SystemsPythonRust

What you need to know about the Pune Tech Scene

Once a far-out concept, AI is now a tangible force reshaping industries and economies worldwide. While its adoption will automate some roles, AI has created more jobs than it has displaced, with an expected 97 million new roles to be created in the coming years. This is especially true in cities like Pune, which is emerging as a hub for companies eager to leverage this technology to develop solutions that simplify and improve lives in sectors such as education, healthcare, finance, e-commerce and more.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account