Leads the design, development, deployment, and maintenance of enterprise data platforms and scalable data pipelines. Builds ETL/ELT, streaming, data quality, governance, storage, and analytics solutions using Azure, Databricks, Spark, Python, Scala, and SQL. Optimizes distributed workloads and cloud infrastructure, implements CI/CD and DevOps practices, ensures compliance and security, collaborates with technical and business stakeholders, and mentors data engineering team members.
Leads the design, development, deployment, and maintenance of data and analytics platforms. Develops reliable, scalable, and efficient data pipelines and data processing solutions that enable data to be effectively processed, stored, governed, and made available to analysts and other data consumers. Collaborates with business stakeholders, IT experts, data scientists, architects, and subject-matter experts to deliver enterprise data and analytics solutions aligned with business and technical requirements.
Key Responsibilities- Design, develop, and automate distributed data ingestion and transformation solutions using data from relational, event-based, semi-structured, and unstructured sources.
- Build reliable, scalable, and efficient ETL/ELT data pipelines using appropriate tools, technologies, and scripting languages.
- Design and implement data quality, validation, monitoring, and alerting frameworks to identify and resolve data integrity issues.
- Implement data governance practices covering metadata, data access, retention, compliance, and security.
- Design and implement physical data models, including database structures, indexing, and table relationships, to support performance and scalability.
- Develop and operate large-scale data storage and processing solutions across cloud and distributed data platforms, including data lakes, warehouses, and lakehouse environments.
- Optimize data pipelines, Spark workloads, databases, and cloud infrastructure for performance, reliability, scalability, and cost efficiency.
- Integrate data from a variety of enterprise applications and source systems and support real-time and event-driven data processing.
- Develop automation for common and repeatable data preparation, integration, deployment, and platform-management activities to minimize manual and error-prone processes.
- Implement CI/CD and DevOps practices to support automated deployment, testing, and release management.
- Participate in troubleshooting, testing, validation, and continuous improvement of data pipelines and platform solutions.
- Ensure data platforms and solutions meet applicable quality, governance, security, compliance, and regulatory requirements.
- Collaborate with data scientists, analysts, architects, IT teams, and business stakeholders to understand requirements and deliver effective data solutions.
- Document data solutions, processes, designs, and technical information to support knowledge transfer and operational effectiveness.
- Apply Agile development methodologies such as Scrum and Kanban to deliver data engineering initiatives.
- Provide technical leadership and mentor less experienced team members, promoting engineering excellence and collaboration.
- Strong technical leadership and decision-making skills, with the ability to lead large-scale data engineering initiatives from concept through production deployment.
- Strong problem-solving and analytical skills, including the ability to diagnose complex data pipeline, platform, and performance issues.
- Excellent communication and collaboration skills, with the ability to work effectively with technical and business stakeholders.
- Ability to translate complex technical concepts and stakeholder requirements into actionable data solutions.
- Strong customer focus and ability to develop solutions aligned with business objectives.
- Ability to balance strategic architecture decisions with hands-on technical execution.
- Strong project management, prioritization, and organizational skills in a fast-paced environment.
- Strong understanding of data quality, governance, security, compliance, and data management principles.
- Passion for continuous improvement, emerging technologies, and modern data engineering practices.
- Ability to mentor, guide, and develop technical talent while fostering a collaborative engineering culture.
- Advanced proficiency in Azure Databricks, Apache Spark, and distributed data processing frameworks.
- Strong expertise in enterprise-scale ETL/ELT and data ingestion pipeline design and development.
- Experience processing structured, semi-structured, streaming, and large-scale datasets.
- Strong understanding of Big Data technologies and scalable cloud-native data architectures.
- Advanced proficiency in Python and Scala for large-scale data processing and engineering solutions.
- Strong SQL skills for querying, transformation, optimization, and analysis of large datasets.
- Experience with scripting, automation, version control, testing, and build processes.
- Strong experience with Azure data services, including:
- Azure Databricks
- Azure Data Lake Storage (ADLS)
- Azure Blob Storage
- Azure Synapse Analytics
- Azure SQL Data Warehouse
- Strong understanding of cloud architecture principles, scalability, reliability, and security best practices.
- Experience with modern data storage and lakehouse technologies, including:
- Delta Lake
- Apache Iceberg
- Parquet
- ORC
- Strong understanding of data lake, data warehouse, and lakehouse architectures.
- Experience with Kafka or similar real-time streaming and event-processing technologies.
- Experience integrating data from a wide variety of enterprise applications and source systems.
- Familiarity with Qlik Replicate or similar data replication and ingestion technologies is preferred.
- Experience implementing CI/CD pipelines and automated deployment processes.
- Strong knowledge of Git, Jenkins, testing methodologies, version control, and release management.
- Experience establishing coding standards, technical documentation, and engineering governance practices.
- Experience implementing data validation frameworks, monitoring solutions, and data quality controls.
- Knowledge of cloud security, access management, data governance, compliance, and regulatory requirements.
- Experience designing reliable, auditable, and observable data platforms.
- 8+ years of hands-on experience in Data Engineering, Data Platform Engineering, Big Data, or a related discipline.
- 3+ years of experience leading technical teams, projects, or significant data engineering initiatives.
- Proven experience building and supporting large-scale cloud-based data ingestion, transformation, and analytics platforms.
- Experience developing end-to-end ETL/ELT solutions in Azure cloud environments.
- Demonstrated experience delivering high-performance and scalable distributed data processing solutions using Spark and Databricks.
- Experience optimizing data pipelines, Spark workloads, databases, and cloud infrastructure for performance, reliability, and cost efficiency.
- Experience working with modern data lake, data warehouse, and lakehouse architectures.
- Experience implementing DevOps practices, CI/CD pipelines, and automated deployment processes.
- Experience working within Agile software development methodologies.
- Experience partnering with data scientists, analysts, architects, and business stakeholders to deliver enterprise data solutions.
- Experience analyzing complex business systems, industry requirements, and data regulations.
- Experience processing and managing large datasets.
- Experience designing and developing Big Data platforms using open-source and third-party technologies.
- Experience with technologies such as Spark, Scala/Java, MapReduce, Hive, HBase, and Kafka or equivalent experience.
- Experience developing applications requiring large-scale file movement and data extraction from multiple sources.
- Experience building analytical solutions.
- Experience with Snowflake and cloud-native data warehousing solutions.
- Experience with Qlik Replicate or similar enterprise data replication technologies.
- Experience supporting real-time analytics and streaming data architectures.
- Experience implementing Infrastructure as Code (IaC) and platform automation frameworks.
- Knowledge of data observability, data lineage, and metadata management tools.
- Experience with IoT technologies.
- Experience supporting AI, machine learning, and advanced analytics workloads on enterprise data platforms.
- Exposure to multi-cloud data ecosystems and hybrid cloud architectures.
- Azure Data Engineer Associate, Azure Solutions Architect, or similar cloud certifications.
Qualifications
About UsCummins is an equal opportunity employer. Our policy is to provide equal employment opportunities to all qualified persons without regard to race, sex, color, disability, national origin, age, religion, union affiliation, sexual orientation, veteran status, citizenship, gender identity, or other status protected by law.- Bachelor’s degree or equivalent qualification in Computer Science, Information Technology, Engineering, Data Science, or another relevant technical discipline.
- Relevant equivalent professional experience may be considered.
- Relevant cloud or data engineering certifications are preferred.
- This position may require licensing or other requirements related to export controls or sanctions regulations.
Cummins Pune, Mahārāshtra, IND Office
Tower A & B, Survey No. 21, Baner - Balewadi Rd, Pune, Maharashtra, India, 411045
Similar Jobs
Fintech • Financial Services
Build and operationalize Azure-based data engineering solutions using Data Factory, Databricks, Data Lake, Azure SQL, and PySpark. Responsibilities include data migration, ingestion and transformation, Lakehouse and data warehouse implementation, pipeline performance tuning, Python APIs, CI/CD, Power BI support, security controls, SDLC participation, stakeholder collaboration, and knowledge sharing.
Top Skills:
ApigeeAzure Data FactoryAzure Data FlowsAzure Data Lake Storage Gen2Azure DatabricksAzure DevopsAzure FunctionsAzure Key VaultAzure Logic AppsAzure SqlCi/CdManaged IdentitiesAzurePower BIPysparkPythonService PrincipalsSQL
Information Technology • Professional Services • Software • App development
Lead the design, implementation, maintenance, and monitoring of scalable Big Data pipelines. Build Java/Scala-based systems using distributed processing frameworks, ETL tools, and cloud-based infrastructure. Review and optimize existing pipelines, troubleshoot production issues, lead technical teams, coordinate with US and India engineers, and deliver new modules and features using Agile practices.
Top Skills:
AkkaApache AirflowApache NifiSparkApache StormGitGradleHadoopJavaJIRAMavenSbtScala
Blockchain • Fintech • Software • Cryptocurrency • Metaverse
Build and operate production-grade AI data pipelines and financial knowledge services for equities content. Responsibilities include processing text, tables, documents, and multimedia; creating knowledge bases and search indexes; engineering RAG retrieval and citation systems; integrating LLMs and NLP models; managing orchestration, APIs, quality, cost, observability, reliability, and scalable production operations.
Top Skills:
APIsAsynchronous Task ProcessingCachingDistributed SystemsFull-Text SearchJavaKnowledge EngineeringLlmsMachine LearningMessage QueuesNlpPythonRagVector Search
What you need to know about the Pune Tech Scene
Once a far-out concept, AI is now a tangible force reshaping industries and economies worldwide. While its adoption will automate some roles, AI has created more jobs than it has displaced, with an expected 97 million new roles to be created in the coming years. This is especially true in cities like Pune, which is emerging as a hub for companies eager to leverage this technology to develop solutions that simplify and improve lives in sectors such as education, healthcare, finance, e-commerce and more.



