JOB TITLE: Platform Engineer (SRE II) - India
JOB STATUS: Full-Time
DEPARTMENT: Platform Engineering
REPORTS TO: Senior Manager, Platform Engineering
SUMMARY
Amtech is scaling its Platform Engineering organization in India as our products move to a fully AWS-hosted, multi-tenant SaaS model. Our SRE practice keeps Encore, LabelTraxx, and supporting platforms reliable, observable, and secure across a multi-account AWS estate.
As a Platform Engineer on the SRE track, you will build and operate the monitoring, automation, and incident response systems that keep Amtech's services running. You will own defined services end to end, automate away toil, and contribute to reliability standards used by the global team.
This is a hands-on technical role that combines operational discipline with software engineering to drive continuous reliability improvement.
JOB DESCRIPTION
Reliability & Performance
- Build and maintain monitoring, alerting, and reliability tooling on our OpenTelemetry-based stack (OpenObserve, CloudWatch) with PagerDuty for alert routing and escalation.
- Analyze production performance, capacity, and error budgets to maintain agreed SLIs and SLOs for your services.
- Implement automated health checks, scaling rules, and self-healing mechanisms to reduce manual intervention.
- Contribute to root cause analysis and post-incident reviews, driving permanent fixes.
Automation & Operations
- Build and maintain infrastructure automation in Terraform across Amtech's multi-account AWS organizations.
- Develop and maintain CI/CD pipelines in GitHub Actions.
- Operate containerized and serverless workloads on ECS Fargate, EKS, and Lambda, including RDS PostgreSQL-backed services.
- Implement safe deployment patterns: automated rollbacks, blue/green, and canary releases.
Incident Response & On-Call
- Participate in the 24/7 PagerDuty on-call rotation and lead response for incidents in your service area.
- Reduce MTTD and MTTR through proactive automation and observability improvements.
- Write and maintain runbooks used across the global SRE team.
Security & Compliance
- Embed security into automation and deployments: IAM design, secrets management, least privilege.
- Maintain systems in line with SOC 2 and ISO 27001 controls, producing audit evidence as part of normal operations.
AI Competency
- Apply AI tools across IaC, pipeline, and automation work with plan review and non-production testing before promotion; understand the blast radius of AI-generated changes.
- Validate AI output against specifications and standards, including AI-generated tests: review coverage and assertions, not just green results.
- Use approved tools only and apply Amtech's data classification policy to every AI interaction.
Collaboration & Continuous Improvement
- Partner with developers to design services for operability, scalability, and resilience.
- Coordinate with U.S. and India peers to keep reliability practices consistent globally.
QUALIFICATIONS
- 2-4 years of hands-on experience in SRE, DevOps, or cloud engineering roles.
- Demonstrated ability to operate production workloads on AWS (EC2, ECS/EKS, RDS, S3, IAM, VPC).
- Working proficiency with Terraform and Git-based CI/CD (GitHub Actions or similar).
- Solid scripting ability in Python or Bash applied to real automation problems.
- Experience with modern observability practices (metrics, logs, traces) and tools such as OpenTelemetry, CloudWatch, Prometheus, or Grafana.
- Understanding of SLO/SLI-driven operations and structured incident management.
- Sound grasp of networking, DNS, and cloud security fundamentals.
- Disciplined AI-assisted engineering practice: structured prompting, output validation, and awareness of where AI-generated code fails.
- Bachelor's degree in Computer Science, Engineering, or a related discipline, or equivalent demonstrated skills.
PREFERRED QUALIFICATIONS
- AWS Certified SysOps Administrator, DevOps Engineer, CKA/CKAD, or Terraform Associate.
- Experience supporting multi-tenant SaaS or account-per-customer AWS architectures.
- Exposure to PagerDuty or equivalent incident management platforms.
- Experience operating AI/LLM-backed services or building agentic automation under governance controls.



