Software Mind Logo

Software Mind

[8SN] Senior Site Reliability Engineer (SRE) – Kubernetes

Posted An Hour Ago
Be an Early Applicant
In-Office or Remote
Hiring Remotely in Montréal, QC
Senior level
In-Office or Remote
Hiring Remotely in Montréal, QC
Senior level
Own production reliability for a Kubernetes-based UI/AI stack: operate and deploy services, monitor health, respond to incidents, runbooks and postmortems, troubleshoot Node.js and JVM runtimes, support GitOps CI/CD and observability, and collaborate with engineering teams on reliability improvements.
The summary above was generated by AI
Company Description

We are Software Mind, an awesome team of engineers who are ready to ramp up any top-notch company’s projects! Our aim? To always be one step ahead. Become part of a multicultural company in constant growth with an excellent work environment certified by Great Place To Work!
 

About the Client

Our client is a leading enterprise software company building highly scalable cloud-native platforms used by organizations around the world. Their engineering teams focus on delivering reliable, secure, and high-performing services while embracing modern DevOps, Kubernetes, and cloud technologies.

You will join a team responsible for ensuring the stability, reliability, and operational excellence of a critical UI service running in production.

Contract Duration: Initial contract through the end of 2026, extending the engagement to a total 12-month term based on performance.

Job Description

About the Role

This is a Senior SRE role supporting production reliability for a Kubernetes-based UI service / AI experience framework stack.

This is not general infrastructure, and it is not a front-end developer role. The strongest candidates will have production SRE experience across Kubernetes operations, observability, Node.js runtime troubleshooting, JVM / Java service troubleshooting, Splunk, and incident ownership.

What You’ll Do

  • Support the deployment, operation, and reliability of production services running on Kubernetes.
  • Monitor service health and investigate production incidents across distributed applications.
  • Participate in on-call support, incident response, root cause analysis, postmortems, and reliability improvements.
  • Troubleshoot application runtime, networking, and service-to-service issues in collaboration with engineering teams.
  • Support CI/CD, GitOps-based deployments, observability, and production monitoring.
  • Work within a client-directed backlog and established priorities.

Qualifications

Required Qualifications

  • 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Production Engineering, or a closely related role, including strong recent hands-on experience supporting Kubernetes-based production services.
  • 3+ years of hands-on production Kubernetes experience strongly preferred. Kubernetes production operations, including deployment, scaling, rollout / rollback, resource tuning, and service-to-service troubleshooting
  • Strong production incident response experience, including on-call, runbooks, postmortems, and paging hygiene
  • Splunk experience for log aggregation, search, and production troubleshooting
  • Prometheus and Grafana experience, specifically building alert rules and dashboards, not only using existing dashboards
  • CI/CD and infrastructure-as-code for containerized deployments, including Helm and GitOps tools such as ArgoCD or Flux
  • Strong Linux and networking fundamentals, including DNS, load balancing, TCP / HTTP, HTTP/2, and Kubernetes networking
  • Node.js production troubleshooting, including heap snapshots, CPU profiles, event-loop blocking, memory growth, worker/process isolation, and V8 isolates or similar runtime models
  • JVM / Java production troubleshooting, including GC log analysis, thread dump analysis, JVM tuning, and Java service latency investigation
  • In-memory cache experience with Redis / Valkey, including key design, TTL / eviction tuning, and cache invalidation
  • Service-to-service authentication experience, including mTLS, certificate rotation, certificate format conversion, and JWT-based service authentication

Additional Information

Nice to Have

  • Web Components / Lit experience, to perform first-level debugging of UI-related issues
  • Server-side rendering or isomorphic runtime experience
  • Canary rollout / multi-version production operations
  • Distributed tracing and request-context correlation
  • KEDA or event-driven autoscaling
  • Experience with enterprise platform integration layers

What We Offer

  • Competitive salary and laptop
  • Professional development and training opportunities
  • Work with cutting-edge cloud and container technologies
  • Flexible work arrangements and collaborative team environment
  • Impact on organization-wide digital transformation initiatives

Similar Jobs

2 Hours Ago
Easy Apply
Remote
Easy Apply
Senior level
Senior level
Cloud • Security • Software • Cybersecurity • Automation
Build and maintain backend services and APIs using Ruby and GraphQL to improve source code workflows. Ensure reliability, performance, testing, observability, and on-call support. Collaborate cross-functionally and use AI agents responsibly to accelerate delivery.
Top Skills: Ai Coding AgentsBackground ProcessingCachingGitGitalyGraphQLMonorepoObservabilityRubySQL
2 Hours Ago
Easy Apply
Remote
Easy Apply
Mid level
Mid level
Cloud • Security • Software • Cybersecurity • Automation
Build and maintain backend services and APIs for source code workflows using Ruby and GraphQL. Improve performance, reliability, caching, and observability. Collaborate cross-functionally, write tests, participate in Tier 2 on-call, and use AI coding agents where appropriate.
Top Skills: GitGitalyGraphQLRubySQL
2 Hours Ago
Easy Apply
Remote
Easy Apply
Senior level
Senior level
Big Data • Fintech • Mobile • Payments • Financial Services
Lead and develop a team of Technical Account Managers supporting SMB merchants post-sale. Drive technical integrations, troubleshoot escalations, set OKRs, forecast hiring, and advocate cross-functionally to increase merchant adoption and product optimization.
Top Skills: Rest ApisWeb Development

What you need to know about the Pune Tech Scene

Once a far-out concept, AI is now a tangible force reshaping industries and economies worldwide. While its adoption will automate some roles, AI has created more jobs than it has displaced, with an expected 97 million new roles to be created in the coming years. This is especially true in cities like Pune, which is emerging as a hub for companies eager to leverage this technology to develop solutions that simplify and improve lives in sectors such as education, healthcare, finance, e-commerce and more.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account