Datavail Logo

Datavail

Senior Associate Cloud SRE - Azure & AWS

Posted 3 Days Ago
Be an Early Applicant
Hybrid
Mumbai, Maharashtra
Senior level
Hybrid
Mumbai, Maharashtra
Senior level
Provides tier-two, 24x7 managed services support for AWS and Azure production environments. Responsibilities include incident response, root cause analysis, reliability engineering, infrastructure automation with Terraform, CI/CD, containerized workload operations, observability, performance and cost optimization, and hybrid-cloud support. The role manages cloud resources, networking, storage, identity, backups, monitoring, Kubernetes deployments, and escalation handling while improving runbooks, automation, service reliability, and operational efficiency.
The summary above was generated by AI

Job Title: Senior Associate Cloud SRE – Azure & AWS

Education: Any Graduate

Experience: 7+years

Location: Mumbai


Role Summary 

As a Site Reliability Engineer supporting multi-cloud infrastructure (AWS and Azure), you will manage complex operational challenges and escalations while implementing reliability best practices across production systems. You will work collaboratively with customer teams and senior engineers to ensure system stability, automate operational workflows, and maintain comprehensive observability. This is a delivery-focused role requiring both advanced technical execution and operational ownership across cloud platforms. 

Key Skills 

  • Cloud Platforms: AWS, Azure, & VMware 
  • Monitoring & Observability: Azure Monitor, New Relic, Splunk 
  • Incident & Service Management: OpsGenie, ServiceNow, Incident Triage, Incident Response 
  • Reliability Metrics: MTTR (Mean Time to Resolve), MTTA (Mean Time to Acknowledge), SLI (Service Level Indicators), SLO (Service Level Objectives), SLA (Service Level Agreement) 
  • Database: Basic Database administration/troubleshooting 
  • Containers: AKS (Azure Kubernetes Service) & Docker. 
  • Operating Systems: Windows Server, Linux Server 

Primary Responsibilities 

Tier 2 Multi-Cloud Operations & Managed Services 

Azure Operations 

  • Provide tier two support for Azure cloud environments 
  • Manage and update Azure Virtual Machine images 
  • Create and configure Azure Virtual Machines and Azure SQL databases 
  • Manage Azure Active Directory (AAD) identities, roles, and role-based access control (RBAC) 
  • Configure Azure Storage account policies and access controls 
  • Manage Virtual Networks, Network Security Groups (NSGs), and route tables 
  • Restore VM snapshots and database backups to lower environments 
  • Manage disk resizing and Azure Managed Disks optimization 
  • Implement Azure resource tagging and cost management 
  • Manage log archiving using Azure Monitor and Log Analytics 
  • PIM (Privileged Identity Management) 

AWS Operations 

  • Provide 24x7x365 tier two support and escalation handling for AWS environments 
  • Patch and manage Amazon Machine Images (AMIs) 
  • Create and configure EC2 instances and RDS databases 
  • Manage IAM roles, users, and policies 
  • Configure S3 bucket policies and Access Control Lists (ACLs) 
  • Open and manage network routes (VPC, subnets, security groups) 
  • Restore snapshots and database backups to lower environments 
  • Increase disk sizes (EBS volumes) and manage storage optimization 
  • Implement proper tagging for environment identification and cost allocation 
  • Manage log archiving using CloudWatch Logs and S3 

Cross-Cloud Responsibilities 

  • Handle escalations from tier one support with deep technical analysis across both platforms 
  • Provide root cause analysis for complex incidents in multi-cloud environments 
  • Implement consistent operational standards across AWS and Azure 
  • Support hybrid cloud connectivity and integration scenarios 

Reliability & Incident Management 

  • Implement and maintain SLIs, SLOs, and SLAs across AWS and Azure in collaboration with senior engineers and customer stakeholders 
  • Lead tier two incident response and incident triage, performing advanced troubleshooting and resolution on both cloud platforms (ServiceNow, OpsGenie) 
  • Conduct thorough post-incident analysis with actionable remediation plans; track and improve MTTR and MTTA 
  • Reduce reactive work by improving runbooks, alert configurations, and standard operating procedures for both clouds 
  • Apply reliability engineering best practices with oversight and review 
  • Mentor tier one engineers during incident response across multi-cloud scenarios 

Automation & Infrastructure as Code 

  • Build and maintain CI/CD pipelines for infrastructure and application deployments on AWS and Azure 
  • Automate complex operational tasks including patching, backups, and environment provisioning across both platforms 
  • Develop infrastructure automation using Terraform for multi-cloud environments 
  • Create scripts and tooling to eliminate manual toil and improve operational efficiency 
  • Implement Azure Resource Manager (ARM) templates or Bicep for Azure-specific automation 
  • Follow established patterns and contribute continuous improvements 
  • Document automation processes for knowledge sharing across cloud platforms 

Containerization & Deployment 

  • Deploy and operate containerized workloads using Docker on AWS services (ECS, EKS) and Azure services (AKS, Azure Container Instances) 
  • Support container reliability through health checks, autoscaling configurations, and resource management on both platforms 
  • Implement safe deployment patterns (canary, blue/green) across AWS and Azure 
  • Troubleshoot complex containerization and orchestration issues in multi-cloud Kubernetes environments 
  • Follow and enhance established containerization standards across both cloud providers 

Observability & Performance 

  • Configure and maintain comprehensive monitoring, logging, and alerting across AWS CloudWatch, Azure Monitor, New Relic, and Splunk 
  • Leverage observability data to identify issues and lead root cause analysis in multi-cloud environments 
  • Contribute to performance tuning and cost optimization initiatives across both platforms 
  • Ensure proper instrumentation and telemetry across AWS and Azure environments 
  • Identify patterns and trends to prevent future incidents 
  • Build custom dashboards and reports using CloudWatch, Azure Monitor, Datadog, and Grafana 

Collaboration & Customer Engagement 

  • Work closely with customer development and operations teams to improve system operability 
  • Participate in design reviews and reliability assessments for multi-cloud architectures 
  • Communicate technical concepts, tradeoffs, and recommendations clearly to stakeholders 
  • Provide regular operational updates and service reports covering both AWS and Azure 
  • Act as technical liaison between customers and internal engineering teams 

Required Qualifications & Experience 

  • 7+ years of hands-on experience in DevOps, SRE, or production operations roles 
  • Proven experience operating production systems in AWS OR Azure (deep expertise in one required) 
  • Working knowledge or exposure to the secondary cloud platform 
  • VMware virtualization experience 
  • Working knowledge of Windows and Linux server administration 
  • Basic database administration/troubleshooting skills 
  • Demonstrated experience managing containerized applications in production 
  • Experience delivering managed services or supporting customer-facing infrastructure 
  • Track record of handling complex technical escalations in cloud environments 

For AWS-Primary Candidates 

  • AWS Services (Expert): EC2, RDS, S3, IAM, VPC, CloudWatch, Lambda, and related services 
  • AWS Networking (Expert): VPCs, subnets, security groups, route tables, VPN/Direct Connect 
  • AWS Storage (Expert): EBS, S3, backup/restore strategies 
  • AWS Containers (Expert): ECS, EKS, or Fargate 
  • Azure (Foundational): Basic understanding, with exposure to Azure VMs, Storage, or networking a plus 

For Azure-Primary Candidates 

  • Azure Services (Expert): Azure VMs, Azure SQL, Storage Accounts, Azure AD, Virtual Networks, Azure Monitor 
  • Azure Networking (Expert): VNets, NSGs, Application Gateway, Azure Firewall, ExpressRoute 
  • Azure Storage (Expert): Managed Disks, Blob Storage, Azure Backup 
  • Azure Containers (Expert): AKS and Azure Container Instances 
  • AWS (Foundational): Basic understanding, with exposure to EC2, S3, or VPC a plus 

Cross-Platform Technical Skills (All Candidates) 

  • Infrastructure as Code: Terraform (preferred) or CloudFormation/ARM templates 
  • CI/CD: Azure DevOps, GitHub Actions, Jenkins, GitLab CI 
  • Scripting/Programming: Python, PowerShell, Bash, or similar languages 
  • Containerization: Strong Docker and Kubernetes experience 
  • Monitoring & Logging: Datadog, Splunk, ELK, Grafana, New Relic, Azure Monitor, OpsGenie, ServiceNow 
  • Version Control: Git and collaborative development workflows 
  • Troubleshooting: Advanced diagnostic and problem-solving capabilities 

Operational Capabilities 

  • Experience with 24x7 operations and tier two escalation support 
  • Strong troubleshooting, incident triage, and root cause analysis skills 
  • Understanding of networking concepts, security best practices, and compliance requirements 
  • Familiarity with backup/restore procedures and disaster recovery planning 
  • Ability to work under pressure during critical incidents 
  • Experience coordinating across distributed teams 
  • Willingness and ability to quickly learn the secondary cloud platform 

Preferred Qualifications & Certifications 

  • AWS Certifications (for AWS-primary): Solutions Architect Associate, SysOps Administrator, or DevOps Engineer Professional 
  • Azure Certifications (for Azure-primary): Azure Administrator Associate (AZ-104) or Azure Solutions Architect Expert (AZ-305) 

Additional Preferred Experience 

  • Hands-on experience with both AWS and Azure (even if limited in one) 
  • Experience with Kubernetes in production environments 
  • Prior consulting or managed services provider experience 
  • Experience with hybrid cloud or cloud migration projects 
  • Experience with configuration management tools (Ansible, Chef, Puppet) 
  • Knowledge of security and compliance frameworks (HIPAA, SOC 2, PCI-DSS) 
  • Experience in high-traffic or mission-critical industries 
  • Multi-cloud architecture or implementation experience 

About the Team
Datavail’s Team of Cloud Experts Can Save You Time and Money
Our Cloud experts are capable to overcome every obstacle in helping clients manage everything from databases, analytics, reporting, migrations, and upgrades to monitoring and overall data management.
You can free up your IT resources to focus on growing your business rather than fighting fires. Our Cloud experts can guide you through strategic initiatives or support routine database management.
Cloud Managed Services
Datavail’s business focuses on helping you use your data to drive business results through cost-saving services. The success of your business depends on how well you understand and manage your data. Our managed cloud services give you the power to unleash your organization’s potential. We provide comprehensive and technically advanced support for Cloud Operation to ensure that your infrastructure is safe, secure, and managed with the utmost level of care.
Our delivery performance in data management leads the industry. We offer highly trained Cloud administrators via a 24×7, always on, always available, global delivery model.
With the combination of a proven delivery model and top-notch experience ensures that Datavail will remain the Cloud experts on demand you desire. Datavail’s flexible and client focused services always add value to your organization.

Similar Jobs

Yesterday
Hybrid
Junior
Junior
Financial Services
Manage second-line oversight of liquidity and interest rate risks across United Kingdom legal entities. Analyze bank balance sheets, monitor risk limits and indicators, assess emerging funding risks, review controls and regulatory requirements, and provide independent challenge to Treasury and risk teams. Prepare risk committee materials, support stress testing and scenario analysis, fulfill regulatory requests, and identify scalable technology solutions for management information and risk monitoring.
Top Skills: AlteryxArtificial IntelligenceExcelPowerPointPythonTableauVisual Studio Code
Yesterday
Hybrid
Pune, Maharashtra, IND
Senior level
Senior level
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Design and develop scalable Java microservices using Spring Boot, event-driven architecture, cloud technologies, and CI/CD practices. Build real-time systems, select appropriate persistence mechanisms, conduct proofs of concept, maintain code quality, coordinate cross-functional teams, and mentor junior developers within an Agile environment.
Top Skills: AgileCi/CdCloud ComputingEvent-Driven ArchitectureJavaMicroservicesNoSQLObject-Oriented ProgrammingRdbmsScaled Agile FrameworkSolidSpring Boot
Yesterday
Hybrid
Pune, Maharashtra, IND
Expert/Leader
Expert/Leader
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Lead product experience design for Mastercard’s data-driven B2B and B2B2C products. Responsibilities include developing scalable design systems, creating high-fidelity designs and reusable components, rapidly prototyping MVPs, defining user flows and journeys, and collaborating with Product, Engineering, UX Research, and global teams. The role balances business objectives with user engagement and uses AI-assisted design and rapid design-to-code workflows.
Top Skills: Ai-Assisted Design ToolingDesign SystemsFigmaFigma MakePrototypingVibe Coding

What you need to know about the Pune Tech Scene

Once a far-out concept, AI is now a tangible force reshaping industries and economies worldwide. While its adoption will automate some roles, AI has created more jobs than it has displaced, with an expected 97 million new roles to be created in the coming years. This is especially true in cities like Pune, which is emerging as a hub for companies eager to leverage this technology to develop solutions that simplify and improve lives in sectors such as education, healthcare, finance, e-commerce and more.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account