About the Job
About Aligned Automation
At Aligned Automation, we live by our "Better Together" philosophy to build a better world. As a strategic service provider to Fortune 500 companies, we help digitize enterprise operations and drive impactful business strategies. Our purpose goes beyond projects—we strive to deliver meaningful, sustainable change that shapes a more optimistic and equitable future.
Our culture is deeply rooted in our 4Cs—Care, Courage, Curiosity, and Collaboration—ensuring that each employee is empowered to grow, innovate, and thrive in an inclusive workplace.
Location: Pune (Work from office)
Job Summary
We are seeking a highly skilled Senior PySpark Data Engineer with 7–8 years of experience in designing, developing, and optimizing large-scale data engineering solutions. The ideal candidate should have extensive experience with PySpark, Python, SQL, Databricks, cloud platforms, and modern data architectures. The role requires hands-on development, solution design, client interaction, and mentoring junior team members.
- Design, develop, and maintain scalable ETL/ELT pipelines using PySpark.
- Build high-performance data pipelines to process large volumes of structured and semi-structured data.
- Develop reusable frameworks for data ingestion, transformation, validation, and monitoring.
- Optimize Spark jobs by tuning partitions, joins, caching, and memory configurations.
- Design and implement Data Lake/Lakehouse architectures using Delta Lake.
- Write complex SQL queries, stored procedures, CTEs, and window functions.
- Collaborate with business stakeholders, architects, and data analysts to understand business requirements.
- Participate in architecture discussions and recommend scalable data solutions.
- Perform code reviews and enforce coding standards and best practices.
- Monitor production pipelines, troubleshoot issues, and implement performance improvements.
- Work with DevOps teams to implement CI/CD for data pipelines.
- Mentor junior engineers and provide technical leadership.
- Estimate effort, prepare technical documentation, and participate in Agile ceremonies.
Programming
- Python (Advanced)
- PySpark (Advanced)
- SQL (Advanced)
- Apache Spark
- Spark SQL
- Delta Lake
- Parquet
- iceberg
- IOMETE
- SQL Server
- PostgreSQL
- Azure SQL
- Azure
- AWS
- Databricks
- Apache Airflow
- Git
- Azure DevOps / GitHub
- CI/CD Pipelines


