Social Discovery Group Logo

Social Discovery Group

Site Reliability Engineer (SRE)

Posted One Month Ago
In-Office or Remote
Hiring Remotely in India
Mid level
In-Office or Remote
Hiring Remotely in India
Mid level
Own and improve production infrastructure reliability, deployments, Infrastructure-as-Code, Kubernetes environments, automation, CI/CD, monitoring, alerting, and observability. Investigate incidents, optimize system performance, maintain documentation and runbooks, and support DNS, WAF, CDN, and caching infrastructure. The role requires strong Linux administration, Bash scripting, networking, Git, and containerization skills, with independent ownership and collaboration across development and operations teams.
The summary above was generated by AI

Social Discovery Group (SDG) is a group of social discovery companies. SDG solves the problems of loneliness, isolation, and disconnection - transforming virtual intimacy into the new normal. SDG’s products redefine the way people interact and connect with one another.

Our portfolio includes social entertainment platforms designed to connect people online across different cultures and regions of the world.

We bring together a team of like-minded people and IT professionals who specialize in creating and developing globally impactful social discovery products. Our international team of digital nomads works remotely from all over the world.

We’re proud to be a two-time “Great Place to Work” winner (USA & Japan, 2024–2025) and a Top-5 Company for Work-From-Anywhere Jobs (FlexJobs, 2025).

We are looking for a Site Reliability Engineer (SRE) passionate about infrastructure reliability, automation, and the development of scalable production systems.

Your main tasks will be:

  • Own and improve production infrastructure reliability and stability
  • Prepare, execute, and support deployments and infrastructure changes
  • Build and maintain Infrastructure-as-Code solutions using Ansible and Terraform
  • Support and optimize Kubernetes-based and containerized environments
  • Develop automation scripts and internal operational tooling
  • Monitor system health, investigate incidents, and proactively improve observability
  • Participate in CI/CD improvements together with Development, QA, DevOps, and SRE teams
  • Work with monitoring and alerting systems to reduce downtime and improve system performance
  • Maintain technical documentation, runbooks, and operational procedures
  • Support DNS, WAF, CDN, and caching infrastructure where required

We expect from you:

  • 3+ years of experience in SRE, DevOps, System Administration, or Build/Release Engineering
  • Strong Linux administration and troubleshooting skills
  • Hands-on experience with Kubernetes and containerization technologies (Docker/Podman)
  • Experience with CI/CD pipelines, preferably GitLab CI
  • Practical experience with Infrastructure-as-Code and configuration management tools (Ansible and/or Terraform)
  • Experience with observability and monitoring tools such as Prometheus, Grafana, Zabbix, or VictoriaMetrics
  • Good understanding of networking fundamentals, DNS, HTTP/HTTPS, load balancing, and troubleshooting
  • Experience with Git and modern software delivery workflows
  • Ability to work independently, take ownership, and proactively improve infrastructure
  • Fluent Russian level for technical documentation and team communication

Nice to have:

  • AWS or GCP experience
  • RabbitMQ / AMQP experience
  • Cloudflare, Akamai, WAF, CDN experience
  • Experience with tracing and advanced observability tooling

What do we offer:

  • REMOTE OPPORTUNITY to work full-time;
  • Vacation 28 calendar days per year;
  • 7 wellness days per year (time off) that can be used to deal with household issues, to lie down and recover without taking sick leave;
  • Bonuses up to $5000 for recommending successful applicants for positions in the company;
  • 50% payment for professional training, international conferences, and meetings;
  • Corporate discount for English lessons;
  • ​Health benefits. According to the paychecks, if you are not eligible for corporate medical insurance, the company will compensate you with up to $1,000 gross per year per employee. This can be spent on self-purchase of health insurance or on doctor’s fees for yourself and close relatives (spouse, children);
  • Workplace organization. The company provides all employees with an equipped workplace and all the necessary equipment (table, armchair, wifi, etc.) in our offices or co-working locations. In the other locations, the company provides reimbursement of workplace costs up to $1000 gross once every 3 years, according to the paychecks. This money can be spent on the rent of the co-working room, on equipping the working place at home (desk, chair, Internet, etc.) during those 3 years;
  • Internal gamified gratitude system: receive bonuses from colleagues and exchange them for our merchandise, team building activities, massage certificates, etc.

Sounds good? Join us now!
The initial pay level or pay range for this role will be shared with candidates during the recruitment process and before the commencement of employment.

Similar Jobs

3 Days Ago
Remote
Senior level
Senior level
Cloud • Security • Software • Generative AI
Own and improve Elastic’s observability infrastructure across hosted cloud deployments. Responsibilities include Terraform-based infrastructure delivery, Python and Go development, production operations, incident response, on-call participation, RCA and postmortem writing, code and design reviews, mentoring, and improving operational documentation and processes. The role also involves operating Linux and containerized workloads, delivering complex projects independently, and maintaining secure, reliable platform infrastructure.
Top Skills: AnsibleArgocdBeatsElastic Cloud Enterprise (Ece)Elastic Cloud Hosted (Ech)Elastic Cloud On Kubernetes (Eck)ElasticsearchGoHelmKibanaKubernetesKyvernoLinuxLogstashPuppetPythonTeleportTerraformVault
5 Days Ago
Remote
Senior level
Senior level
Software • Automation
Owns reliability, scalability, observability, and incident response for a mission-critical SaaS platform. Responsibilities include 24x7 on-call support, root cause analysis, automation, AWS infrastructure design, EKS/Kubernetes and Docker operations, Terraform and Helm deployments, CI/CD maintenance, cloud networking, monitoring with Datadog, Grafana, and Prometheus, database operations, and migration toward Kubernetes. The role also develops self-healing systems, documentation, runbooks, and cross-functional customer-focused reliability practices.
Top Skills: AlbAmazon RdsAWSBashCi/CdDatadogDevsecopsDockerDocker SwarmEksGrafanaHelmIamInfrastructure As CodeKubernetesLinuxNlbPostgresPrometheusPythonRoute 53TerraformTerraformTransit GatewayVpcVpn
19 Days Ago
Remote
Entry level
Entry level
Cloud • Information Technology • Business Intelligence • Consulting
Designs, builds, and operates cloud infrastructure and reliability capabilities for an enterprise AI platform. Responsibilities include infrastructure as code, landing zones, Kubernetes, CI/CD, observability, incident response, SLOs, production readiness, automation, cost optimization, and support for hybrid, edge, on-premises, and customer-controlled environments. This remote, client-facing consulting role requires strong communication, production ownership, and the ability to balance reliability, delivery speed, security, and operational cost.
Top Skills: AlertingAzureAzure ArcAzure DevopsBicepCi/CdCloud InfrastructureDashboardsGithub ActionsGpu WorkloadsIdentity And Access ManagementInfrastructure As CodeKubernetesLogsMetricsNetworkingObservabilityTerraformTraces

What you need to know about the Pune Tech Scene

Once a far-out concept, AI is now a tangible force reshaping industries and economies worldwide. While its adoption will automate some roles, AI has created more jobs than it has displaced, with an expected 97 million new roles to be created in the coming years. This is especially true in cities like Pune, which is emerging as a hub for companies eager to leverage this technology to develop solutions that simplify and improve lives in sectors such as education, healthcare, finance, e-commerce and more.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account