Amtech Software Logo

Amtech Software

Platform Engineer (SRE - India)- III

Posted 5 Hours Ago
Be an Early Applicant
Remote
Hiring Remotely in India
Senior level
Remote
Hiring Remotely in India
Senior level
Senior SRE responsible for designing and operating reliable, observable, and secure AWS-hosted multi-tenant SaaS platforms. Build monitoring/alerting, SLO-driven operations, automation via Terraform and CI/CD, operate containerized/serverless workloads, lead incident response and game days, embed security and compliance, integrate AI safely into engineering workflows, and mentor junior SREs.
The summary above was generated by AI

SUMMARY

 

Amtech is scaling its Platform Engineering organization in India as our products move to a fully AWS-hosted, multi-tenant SaaS model. Our SRE practice keeps Encore, LabelTraxx, and supporting platforms reliable, observable, and secure across a multi-account AWS estate.


As a senior individual contributor on the SRE track, you will own the reliability of significant production domains, design the observability and automation frameworks the team standardizes on, and act as incident commander for complex, cross-service incidents.


You will raise the bar for the India SRE pod: setting patterns, reviewing designs, and mentoring earlier-career engineers while remaining deeply hands-on.

 

JOB DESCRIPTION

 

Reliability & Performance

  • Design and implement monitoring, alerting, and reliability frameworks on our OpenTelemetry-based stack (OpenObserve, CloudWatch) integrated with PagerDuty.
  • Define and defend SLIs, SLOs, and error budgets for the domains you own, and drive engineering priorities from them.
  • Engineer self-healing, autoscaling, and capacity management so systems recover without human intervention.
  • Lead root cause analysis for high-severity incidents and ensure permanent, verified fixes.

Automation & Operations

  • Design reusable Terraform modules and golden-path patterns adopted across Amtech's multi-account AWS organizations.
  • Build and harden CI/CD pipelines in GitHub Actions, including progressive delivery (blue/green, canary) and automated rollback.
  • Operate and optimize containerized and serverless workloads (ECS Fargate, EKS, Lambda) and RDS PostgreSQL data stores at production scale.
  • Eliminate toil systematically: identify, quantify, and automate the highest-cost operational work.

Incident Response & On-Call

  • Serve in the 24/7 on-call rotation and act as incident commander for complex, multi-service incidents.
  • Own MTTD/MTTR improvement for your domains, with measurable targets.
  • Author runbooks and drive game days or failure testing to validate them.

Security & Compliance

  • Engineer security into the platform: IAM boundaries, secrets management, network segmentation, and policy enforcement.
  • Ensure services meet SOC 2 and ISO 27001 obligations with audit evidence generated automatically where possible.

AI Competency

  • Integrate AI across the delivery workflow with reusable prompts and shared context; mentor earlier-career engineers on effective, safe usage.
  • Design validation steps for AI-assisted changes: characterization tests before AI refactors, empirical verification of AI debugging hypotheses, and rollback plans.
  • Quantify AI productivity impact against baselines rather than impressions, and apply data classification policy to all AI usage.

Technical Leadership & Collaboration

  • Mentor SRE I/II engineers through design review, paired incident response, and code review.
  • Partner with product development and Cloud-track peers on operability and migration readiness for Encore and LabelTraxx workloads.
  • Champion reliability culture and consistent global standards between India and U.S. teams.

 

QUALIFICATIONS

 

  • 4-6 years of hands-on SRE, DevOps, or cloud engineering experience, including ownership of production services at meaningful scale.
  • Deep working knowledge of AWS (EC2, ECS/EKS, RDS, S3, IAM, VPC, Lambda) in multi-account environments.
  • Strong Terraform skills, including authoring reusable modules, and fluency with GitHub Actions or equivalent CI/CD.
  • Proven software engineering ability in Python (or similar) applied to automation and tooling, not just scripts.
  • Demonstrated experience running SLO-driven operations, incident command, and blameless post-incident reviews.
  • Experience designing observability with OpenTelemetry or equivalent (metrics, logs, traces).
  • Solid understanding of networking, DNS, and cloud security architecture.
  • Demonstrated integration of AI into daily engineering workflow with measured impact and designed validation, not ad-hoc usage.
  • Bachelor's degree in Computer Science, Engineering, or a related discipline, or equivalent demonstrated skills.

PREFERRED QUALIFICATIONS

 

  • AWS Certified DevOps Engineer Professional, CKA, or Terraform Associate.
  • Experience with multi-tenant SaaS or account-per-customer architectures and ERP-class workloads.
  • Experience with PagerDuty at scale (escalation policies, service ownership models).
  • Experience operating AI/LLM workloads or building agentic automation under governance controls.
  • Prior mentorship or tech-lead experience in a distributed global team.

Similar Jobs

5 Hours Ago
Remote or Hybrid
India
Mid level
Mid level
Artificial Intelligence • Cloud • Sales • Security • Software • Cybersecurity • Data Privacy
Operate and secure SailPoint's cloud SaaS infrastructure: implement FIM, harden base images, automate audit evidence collection, secure container images, integrate SIEM/IDS/IPS/WAF, support 24x7 production, participate in on-call rotation, and collaborate with engineering and security teams to meet compliance standards.
Top Skills: AWSAzureChefContainersFile Integrity Monitoring (Fim)IdsIpsJenkinsPuppetPythonRubySIEMSirpTerraformWaf
7 Hours Ago
Remote or Hybrid
India
Entry level
Entry level
Fintech • Legal Tech • Software • Financial Services • Cybersecurity • Data Privacy
Provide onsite and remote IT support to internal and external users, handle high-volume service desk tickets, troubleshoot Microsoft OS/Office/Exchange/AD and network/infra issues, track KPIs and SLAs, participate in on-call rotational shifts, produce documentation, and ensure compliance with ITIL and security policies.
Top Skills: Active DirectoryInfrastructureItilMessagingMicrosoft ExchangeMS OfficeWindowsNetworking
7 Hours Ago
Easy Apply
Remote or Hybrid
India
Easy Apply
Mid level
Mid level
Big Data • Cloud • Software • Database
Manage full-cycle recruiting for People and Finance roles across business units. Source candidates via advanced techniques and tools, partner with hiring managers, maintain ATS data accuracy, and deliver a high-touch candidate experience from outreach through offer execution.
Top Skills: AtsBoolean SearchCRMGemGreenhouseLinkedin Recruiter

What you need to know about the Pune Tech Scene

Once a far-out concept, AI is now a tangible force reshaping industries and economies worldwide. While its adoption will automate some roles, AI has created more jobs than it has displaced, with an expected 97 million new roles to be created in the coming years. This is especially true in cities like Pune, which is emerging as a hub for companies eager to leverage this technology to develop solutions that simplify and improve lives in sectors such as education, healthcare, finance, e-commerce and more.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account