Palo Alto Labs Logo

Palo Alto Labs

Data Engineer- DataBricks

Posted One Month Ago
Be an Early Applicant
In-Office
Pune, Maharashtra, IND
Senior level
In-Office
Pune, Maharashtra, IND
Senior level
Designs and operates Databricks data platforms, pipelines, workflows, CI/CD processes, and cloud infrastructure on Azure. Responsibilities include cluster and workspace management, automated testing, Infrastructure as Code, access controls, FinOps, Airflow orchestration, monitoring, incident response, and reusable Terraform and Databricks templates. The role also improves platform reliability, security, developer productivity, and production code promotion processes while partnering with engineering, architecture, and security teams.
The summary above was generated by AI

Role: Data Engineer
Location:
Remote
Employment Type: Full-Time
Experience: 4-7 years

 

KEY RESPONSIBILITIES

·       Design and implement CI/CD pipelines for data pipelines (ingestion jobs) and transformation projects (e.g., dbt on Databricks, SQL, notebooks)

·       Orchestrate Databricks jobs and workflows end-to-end (ingestion, transformation, quality checks)

·       Integrate automated testing into CI/CD, including schema and contract checks for data models and tables

·       Implement FinOps best practices to support cost monitoring and allocation across the EDP

·       Automate platform operations for Databricks and related services, such as workspace and cluster provisioning, library and runtime mgmt. and job deployment/config

·       Implement and maintain identity and access management for Databricks and supporting cloud resources, including workspace- and cluster-level permissions, table- and view-level access controls (e.g., Unity Catalog or equivalent), service principals, groups, and roles for automated workloads, and RBAC & TBAC models in collaboration with EDP Architect

·       Provide patterns, templates, and reusable modules (Terraform modules, Airflow DAG patterns, Databricks job templates) to accelerate onboarding of new projects

·       Continuously evaluate and improve tooling, pipelines, and platform architecture to increase reliability, security, and developer productivity on Databricks

·       Define the code promotion process to minimize impacts across domains as code is promoted to production

·       Manage end-to-end orchestration using managed Airflow. Contribute to defining and tracking SLA/SLO/SLIs for the platform and participate in incident response (triage, root cause analysis)

·       Practical knowledge of IT Infrastructure technologies, cloud computing Azure), cybersecurity, and disaster recovery.

·       Working knowledge of Azure ecosystem, including hands-on experience designing, building, and optimizing scalable data pipelines within cloud-native environments.

QUALIFICATIONS

·       Hands-on experience with Databricks in production environment, including workspace and cluster management, jobs/workflows and integrations with orchestration tools

·       Strong experience with CI/CD pipelines (e.g., GitHub Actions, GitLab CI, Azure DevOps, or similar) and Git-based workflows.

·       Strong experience with Infrastructure as Code (IaC) and orchestration tools for provisioning and managing Databricks and cloud infrastructure.

·       Experience implementing automated tests and quality gates in CI/CD pipelines.

·       Ability to partner effectively with data engineers, analytics engineers, architects, and security teams.

MINIMUM EXPERIENCE & EDUCATION

·       Bachelor’s degree in Computer Science, Information Technology, or related field preferred, or equivalent work experience.

·       4–7+ years in DevOps, Cloud Engineering, Site Reliability Engineering, or Platform Engineering, with at least 2+ years supporting data/analytics platforms.

·       Experience operating production data workloads, including monitoring, logging, performance tuning, and incident response.

·       Scripting skills (e.g., Python, Bash, PowerShell) for automation and integration.

·       Experience in CPG, retail, manufacturing, or distribution environments preferred.

 

 



Palo Alto Labs Pune, Mahārāshtra, IND Office

Pune, India, 411006

Similar Jobs

7 Days Ago
Hybrid
Senior level
Senior level
Financial Services
Build and maintain Databricks-based data integration and ETL pipelines using Python, PySpark, DLT, Delta Lake, and Unity Catalog. Design governed semantic models and self-service analytics layers, support data governance and access controls, and deliver executive dashboards in Tableau or Sigma. Partner with business stakeholders to translate requirements into actionable insights while maintaining production-quality code, documentation, lineage, and compliance in a regulated environment.
Top Skills: AlteryxAWSDatabricksDatabricks GenieDatabricks SqlDelta LakeDelta Live TablesGitGithub CopilotJulesProphecyPysparkPythonServicenowSigmaSnowflakeSQLTableauUnity Catalog
12 Days Ago
Hybrid
Pune, Maharashtra, IND
Expert/Leader
Expert/Leader
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Lead the architecture and development of scalable batch, streaming, and event-driven data platforms supporting analytics, machine learning, generative AI, and agentic AI. Build governed lakehouse architectures, reusable data products, ingestion and orchestration pipelines, and AI lifecycle capabilities. Drive security, governance, observability, reliability, performance, and cost optimization while translating R&D concepts into production systems. Provide hands-on technical leadership, architecture reviews, mentorship, and cross-functional collaboration.
Top Skills: Apache AirflowSparkAWSAws Step FunctionsAzure Ai FoundryAzure Data FactoryAzure Machine LearningCi/CdDatabricksDatabricks WorkflowsDockerInfrastructure As CodeKafkaKubernetesMicrosoft FabricMicrosoft PurviewMlflowPysparkPythonSQLTerraformUnity Catalog
12 Days Ago
Hybrid
Pune, Maharashtra, IND
Expert/Leader
Expert/Leader
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Lead the architecture and hands-on development of scalable batch, streaming, and event-driven data platforms supporting analytics, machine learning, generative AI, and agentic AI. Build governed lakehouse architectures, reusable data products, ingestion and orchestration pipelines, and ML lifecycle capabilities. Establish standards for security, governance, testing, CI/CD, observability, reliability, and cost optimization while providing technical leadership, mentoring, architecture reviews, and cross-functional collaboration.
Top Skills: Apache AirflowSparkAWSAws Step FunctionsAzure Ai FoundryAzure Data FactoryAzure Machine LearningCi/CdCloud StorageData GovernanceData LakesData LineageData MeshesDatabricksDatabricks WorkflowsDockerEvent-Driven ArchitectureInfrastructure As CodeKafkaKubernetesLakehousesMetadata ManagementMicrosoft FabricMicrosoft PurviewMlflowPysparkPythonServerlessSQLStreamingTerraformUnity Catalog

What you need to know about the Pune Tech Scene

Once a far-out concept, AI is now a tangible force reshaping industries and economies worldwide. While its adoption will automate some roles, AI has created more jobs than it has displaced, with an expected 97 million new roles to be created in the coming years. This is especially true in cities like Pune, which is emerging as a hub for companies eager to leverage this technology to develop solutions that simplify and improve lives in sectors such as education, healthcare, finance, e-commerce and more.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account