Baker Hughes Logo

Baker Hughes

Digital Technology Senior Specialist – Observability & AI Ops

Reposted Yesterday
Be an Early Applicant
In-Office
Pearl Tower, Hadapsar, Pune, Maharashtra, IND
Senior level
In-Office
Pearl Tower, Hadapsar, Pune, Maharashtra, IND
Senior level
Lead implementation and operations of enterprise observability and AIOps platforms (Elastic stack preferred). Onboard infrastructure, cloud, and applications; build dashboards, alerts, ingestion pipelines; support incident investigation, automation, IaC, upgrades, and runbook documentation to improve reliability and operational readiness.
The summary above was generated by AI

DT Senior Specialist – Observability & AI Ops

Would you like to help shape and implement our Digital Technology teams' strategic direction?

Are you passionate about helping improve observability and digital operations?

Join our Digital Technology team!

We operate at the heart of Baker Hughes digital transformation journey. Our team delivers enterprise observability and AIOps capabilities that help technology teams detect issues earlier, troubleshoot faster, and improve service performance across cloud, infrastructure, and application environments.

Partner with the best

As an AIOps & Observability Engineer, you will support the implementation, enhancement, onboarding, and day-to-day operations of observability platforms, with a focus on Elastic Stack capabilities and practical SRE-driven operational outcomes.

As a Senior AI Ops Engineer, you will be responsible for:

  • Implement and manage the life cycle of enterprise observability platform based on elastic tech stacks [not limited to] Kibana, Logstash, Beats, Elastic Agent, Fleet, Elastic APM components, etc.
  • Onboard full suite of 20,000 plus devices under observability umbrella
  • Onboarding infrastructure, cloud services, applications, and platforms into observability and monitoring solutions.
  • Building and maintaining dashboards, visualizations, alerts, and operational reports for technology teams.
  • Configuring log, metric, trace, uptime, and APM data collection across supported environments.
  • Assisting with data ingestion, parsing, enrichment, and retention activities.
  • Supporting incident investigation, troubleshooting, and root cause analysis using observability data.
  • Collaborating with cloud, infrastructure, application, and SRE teams to improve system reliability and service visibility.
  • Contributing to automation initiatives using scripting, Infrastructure as Code, and repeatable deployment practices.
  • Participating in observability platform upgrades, patching, performance tuning, and operational support activities.
  • Creating and maintaining runbooks, knowledge articles, dashboards standards, and operational runbooks related documentation.
  • Contributing to continuous improvement of observability practices, monitoring coverage, and operational readiness.

Fuel your passion

  • Have 7+ years – SRE/DevOps experience in enterprise-scale or mission-critical environments
  • Have 5+ years – Cloud / Application / Platform operations and administration (AWS, Azure, hybrid or multi-cloud)
  • Have 5+ years – Automation, CI/CD, and scripting proficiency (Python, Bash, PowerShell, Ruby, or equivalent)
  • Have 5+ years - Exposure to containers and cloud-native platforms such as Docker, Kubernetes, Prometheus, or Grafana.
  • Have 3+ years – Proven experience administering Elastic Observability platforms across the full lifecycle, including deployment, maintenance, upgrades, patching, and capacity scaling.

Preferred qualifications

  • AWS or Azure Associate-level certification, or equivalent practical cloud operations experience. 
  • Elastic Certified Engineer or equivalent observability platform certification
  • Familiarity with infrastructure as code (GitHub Actions, CloudFormation, Terraform, Ansible) for repeatable automation
  • Exposure to cloud-native observability frameworks (Open Telemetry, service meshes)
  • Experience documenting runbooks, playbooks, and consumption guides for SMEs
  • Process knowledge such as Agile and/or ITIL

Must have Technical Skills

  • Strong background in observability platforms (Elastic.io stack preferred: Elasticsearch, Kibana, Logstash, Beats, Elastic APM, and Fleet/Elastic Agent)
  • Telemetry FundamentalsStrong, practical understanding of fundamental observability concepts, including the collection and analysis of logs, metrics, traces, and synthetic monitoring.
  • Experience with administration of leading observability platforms (Grafana, Graylog, Splunk, Sumo Logic, Tanzu, or open-source equivalents) including lifecycle management (Kubernetes, Docker, Prometheus, and Grafana installation, patching, upgrades, scaling).
  • Strong knowledge of distributed infrastructure domains (network, servers, VMs, AWS, Azure, databases) from an observability perspective.
  • Proven ability to design and tune scalable ingestion pipelines for diverse, globally distributed data sources.
  • Flexibility to adapt and evolve observability standards per domain needs while ensuring consistency across the enterprise.
  • Operational ownership mindset — accountable for uptime, reliability, and lifecycle management of the hosted observability platform.
  • Incident management and SRE practices: monitoring, alerting, troubleshooting, root cause analysis, and postmortems.
  • Proficiency in automation and scripting (Python, Bash, PowerShell, Ruby, etc.) for ingestion, upgrades, and operational tasks.
  • Familiarity with REST APIs and tools like Postman, plus DevOps constructs (GitHub, Jenkins, CI/CD pipelines, serverless technologies).
  • Configuration Proficiency: Demonstrated proficiency in managing system configurations using YAML-based configurations.
  • Knowledge of infrastructure as code (Terraform, Ansible) for repeatable automation.
  • Strong understanding of AWS/Azure services relevant to ingestion and enrichment (e.g., Kinesis, Event Hub, Lambda, Functions).
  • Data Presentation: Extensive experience in performance optimization, advanced data visualization, and sophisticated dashboarding using Kibana (or similar platforms).
  • Ability to document and enable SMEs with clear runbooks, onboarding guides, and consumption standards.
  • Collaboration skills to guide Infra domain SMEs on onboarding data sources and consuming observability outputs.

Good-to-Have

  • Exposure to cloud-native observability frameworks (Open Telemetry, service meshes).
  • Basic understanding of shared infrastructure services (DNS, DHCP, Active Directory, SSL, load balancing).
  • Experience with hybrid deployments (on-prem + cloud) and handling data sovereignty/regulatory considerations in global industries.
  • Security monitoring and compliance awareness relevant to IT domains.
  • Contribution to observability best practices (dashboards, alerting strategies, operational standards).
  • Mentoring and cross-training skills to expand in-house observability expertise.
  • Process knowledge like Agile and/or ITIL.
  • AWS Professional-level certification, or the ability to demonstrate equivalent functional knowledge and expertise.

Desired Characteristics

  • Strong analytical and troubleshooting skills.
  • Clear written and verbal communication skills.
  • Ability to collaborate with application, infrastructure, cloud, and operations teams.
  • Self-motivated with a passion for learning new technologies.
  • Customer-focused mindset and commitment towards operational excellence.
  • Ability to work independently and as part of a globally distributed team.
  • Proactive approach to maintaining documentation, fostering knowledge transfer, and supporting continuous service improvement.

Work in a way that works for you

We recognize that everyone is different and that the way in which people want to work and deliver at their best is different for everyone too. In this role, we can offer the following flexible working patterns:

  • Working flexible hours - flexing the times when you work in the day to help you fit everything in and work when you are the most productive

Working with us

Our people are at the heart of what we do at Baker Hughes. We know we are better when all of our people are developed, engaged and able to bring their whole authentic selves to work. We invest in the health and well-being of our workforce, train and reward talent and develop leaders at all levels to bring out the best in each other.

Working for you

Our inventions have revolutionized energy for over a century. But to keep going forward tomorrow, we know we have to push the boundaries today. We prioritize rewarding those who embrace change with a package that reflects how much we value their input.  Join us, and you can expect:

  • Contemporary work-life balance policies and wellbeing activities
  • Comprehensive private medical care options
  • Safety net of life insurance and disability programs
  • Tailored financial programs
  • Additional elected or voluntary benefits
The Baker Hughes internal title for this role is: Digital Technology Specialist - DevOps Engineering

Similar Jobs

Yesterday
Remote or Hybrid
India
Entry level
Entry level
Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Own the future-state business architecture for regulatory reporting within the finance controllership value stream. Define capabilities, services, processes, operating models, governance, requirements, and measurable outcomes. Represent Finance in technology, data, and AI architecture forums; ensure alignment of designs, data definitions, controls, and delivery backlogs. Support detailed solution design, resolve cross-domain impacts, influence senior stakeholders, drive simplification and standardization, and coach teams on business architecture practices.
Top Skills: Ai ArchitectureData ArchitectureTechnology Architecture
Yesterday
Remote or Hybrid
India
Mid level
Mid level
Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Manage UK STDF actuals and stress testing governance, controls, reporting, approvals, delivery timelines, and audit-ready documentation. Oversee risk and control assessments, regulatory report inventories, issue remediation, stakeholder forums, executive sign-offs, and submissions. Ensure readiness for 2027 BCST, coordinate cross-functional engagement with Finance and Treasury Controls Offices, and drive continuous improvement and automation across the stress testing control environment.
Top Skills: HeliosSharepoint
Yesterday
Remote or Hybrid
Mid level
Mid level
Artificial Intelligence • Big Data • Cloud • Information Technology • Software • Big Data Analytics • Automation
The Lead Solutions Engineer at Dynatrace will provide technical support to the sales team, demonstrate product capabilities, manage POCs, and engage with customers to promote sales and gather feedback for product improvement.
Top Skills: .NetAnsibleAWSAzureCi/CdContainersCSSGCPGoHTMLJavaJavaScriptKubernetesLinuxNode.jsOpenshiftPHPPuppetServerlessTerraformWindows

What you need to know about the Pune Tech Scene

Once a far-out concept, AI is now a tangible force reshaping industries and economies worldwide. While its adoption will automate some roles, AI has created more jobs than it has displaced, with an expected 97 million new roles to be created in the coming years. This is especially true in cities like Pune, which is emerging as a hub for companies eager to leverage this technology to develop solutions that simplify and improve lives in sectors such as education, healthcare, finance, e-commerce and more.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account