Design, build, and optimize PySpark applications and ETL pipelines to process large-scale datasets from relational, NoSQL, file, and streaming sources. Ensure data quality, implement error handling, and collaborate with analysts, scientists, and architects to deliver performant data solutions.
Job Title: PySpark Data Engineer
Summary:
We are seeking a skilled PySpark Data Engineer to join our team and drive the development of robust data processing and transformation solutions within our data platform. You will be responsible for designing, implementing, and maintaining PySpark-based applications to handle complex data processing tasks, ensure data quality, and integrate with diverse data sources. The ideal candidate possesses strong PySpark development skills, experience with big data technologies, and the ability to work in a fast-paced, data-driven environment.
Key Responsibilities: Data Engineering Development:
- Design, develop, and test PySpark-based applications to process, transform, and analyze large-scale datasets from various sources, including relational databases, NoSQL databases, batch files, and real-time data streams.
- Implement efficient data transformation and aggregation using PySpark and relevant big data frameworks.
- Develop robust error handling and exception management mechanisms to ensure data integrity and system resilience within Spark jobs.
- Optimize PySpark jobs for performance, including partitioning, caching, and tuning of Spark configurations.
Data Analysis and Transformation:
- Collaborate with data analysts, data scientists, and data architects to understand data processing requirements and deliver high-quality data solutions.
- Analyze and interpret data structures, formats, and relationships to implement effective data transformations using PySpark.
- Work with distributed datasets in Spark, ensuring optimal performance for large-scale data processing and analytics.
Data Integration and ETL:
- Design and implement ETL (Extract, Transform, Load) processes to ingest and integrate data from various sources, ensuring consistency, accuracy, and performance.
- Integrate PySpark applications with data sources such as SQL databases, NoSQL databases, data lakes, and streaming platforms
Qualifications and Skills:
- Bachelor's degree in Computer Science, Information Technology, or a related field.
- 5+ years of hands-on experience in big data development, preferably with exposure to data-intensive applications.
- Strong understanding of data processing principles, techniques, and best practices in a big data environment.
- Proficiency in PySpark, Apache Spark, and related big data technologies for data processing, analysis, and integration.
- Experience with ETL development and data pipeline orchestration tools (e.g., Apache Airflow, Luigi).
- Strong analytical and problem-solving skills, with the ability to translate business requirements into technical solutions.
- Excellent communication and collaboration skills to work effectively with data analysts, data architects, and other team members.
Similar Jobs
Fintech • Legal Tech • Software • Financial Services • Cybersecurity • Data Privacy
Leads the architecture, design, deployment, and operation of enterprise middleware messaging and streaming integrations. Designs multi-region, multi-datacenter platforms with geo-replication, failover, high availability, fault tolerance, and low-latency performance. Evaluates messaging technologies, establishes deployment and monitoring standards, automates infrastructure, collaborates with stakeholders, and mentors junior engineers. The role requires extensive experience with Kafka, RabbitMQ, cloud messaging services, Kubernetes, Docker, clustered environments, and resilient high-volume systems.
Top Skills:
ActivemqApache KafkaAws SnsAws SqsAzure Service BusDockerElasticGcp Pub/SubIbm MqKafka StreamsKubernetesRabbitMQRedpandaStreamnative
Artificial Intelligence • Cloud • Sales • Security • Software • Cybersecurity • Data Privacy
Leads SailPoint’s Data and AI Platform Engineering organization, setting technical strategy and managing engineering teams responsible for large-scale data processing, machine learning infrastructure, streaming, data governance, search, and distributed systems. Oversees platform roadmaps, architecture, delivery, reliability, stakeholder collaboration, talent development, and thought leadership across technologies including Flink, Spark, Snowflake, Iceberg, Airflow, and OpenSearch.
Top Skills:
Apache AirflowApache FlinkApache IcebergSparkAWSDbtGoogle Cloud PlatformJavaAzureOpensearchPythonScalaSnowflake
Artificial Intelligence • Hardware • Information Technology • Machine Learning
Leads end-to-end semiconductor tool hook-up coordination, including utility requirements, layout reviews, field installation, QA/QC, testing, commissioning, documentation, and handover. Coordinates mechanical, electrical, network, water, and gas systems with OEMs, facility teams, contractors, and quality teams. Ensures cleanroom, safety, engineering, and schedule compliance while improving installation workflows, resource planning, sequencing, productivity, and execution efficiency.
Top Skills:
Microsoft CopilotMicrosoft Power Bi
What you need to know about the Pune Tech Scene
Once a far-out concept, AI is now a tangible force reshaping industries and economies worldwide. While its adoption will automate some roles, AI has created more jobs than it has displaced, with an expected 97 million new roles to be created in the coming years. This is especially true in cities like Pune, which is emerging as a hub for companies eager to leverage this technology to develop solutions that simplify and improve lives in sectors such as education, healthcare, finance, e-commerce and more.



.jpeg)