Build and maintain scalable GCP-based ETL/ELT pipelines using Python, BigQuery, and Google Cloud Storage. Support machine learning infrastructure through feature stores, model data feeds, Vertex AI integration, and MLOps workflows. Ingest data from APIs, streaming platforms, and databases; optimize SQL and BigQuery architecture; implement CI/CD, testing, data quality checks, monitoring, and governance. Collaborate with data scientists, ML engineers, and product managers to operationalize production machine learning systems.
We are looking for someone with advanced Python programming skills who applies robust software engineering principles to data problems. You will collaborate closely with Data Scientists, ML Engineers, and Product Managers to build the scalable, automated pipelines required to train, deploy, and monitor machine learning models in production.
Key Responsibilities
- GCP Pipeline Development: Design, build, and maintain highly scalable ETL/ELT data pipelines using Python and GCP-native data processing tools (e.g., Cloud Run, Cloud Functions).
- AI/ML Infrastructure Support: Engineer feature stores, robust data feeds specifically optimized for machine learning training and inference. Work closely with ML Engineers to operationalize models using Vertex AI.
- Data Integration & Ingestion: Write clean, modular Python code to ingest data from diverse sources (APIs, streaming platforms, on-prem databases) into BigQuery and Google Cloud Storage (GCS).
- System Optimization: Optimize BigQuery architecture, partition/cluster tables, and tune complex SQL queries to ensure performance and cost-efficiency at a massive scale.
- Software Engineering Best Practices: Champion best practices in Python development, including version control (Git), CI/CD pipelines (Cloud Build / GitHub Actions), code reviews, and comprehensive unit/integration testing.
- Data Quality & Governance: Implement robust data quality checks, alerting, and monitoring to ensure the data feeding our AI models is accurate and reliable.
Qualifications
Required Qualifications
- Degree: Bachelor’s or Master’s degree in Computer Science, Engineering, Mathematics, or a related technical field (or equivalent practical experience).
- Experience: 4 to 6 years of professional experience in Data Engineering, Software Engineering, or a closely related field.
- Advanced Python: Deep expertise in Python programming. You should be highly comfortable with:
- Data processing and ML-adjacent libraries (e.g., PySpark, Pandas, NumPy).
- API development
- Writing efficient and production-grade code.
- GCP Mastery: Proven, hands-on experience designing and operating data architectures on Google Cloud Platform. Must have strong experience with:
- BigQuery (advanced SQL, architecture, and optimization).
- Google Cloud Storage (GCS).
- Compute/Serverless (Cloud Functions, Cloud Run).
- AI/ML Acumen: Experience working alongside Data Science teams. A strong understanding of the ML lifecycle, feature engineering, and the data requirements for model training and deployment.
MLOps: Understanding of MLOps principles, model registry, and continuous training pipelines.
Preferred Qualifications
- Vertex AI: Direct experience interacting with or deploying pipelines using Google Cloud's Vertex AI platform.
- Streaming Technologies: Familiarity with real-time data processing using Google Cloud Pub/Sub and streaming Dataflow jobs.
- Infrastructure as Code: Experience managing GCP resources using Terraform.
- Containerization: Proficiency with Docker.
Similar Jobs
Cloud • Information Technology • Productivity • Security • Software • App development • Automation
Lead data engineering architecture and execution for large-scale, low-latency analytics. Own technical direction, ensure operational data quality, design streaming and batch pipelines, mentor engineers, coordinate cross-functional teams, and deliver scalable solutions using Spark, Airflow, and AWS or equivalent data platforms.
Top Skills:
AirflowAthenaColumn StoresDatabricksDatabricks ApisEmrFlinkHiveKafkaKappa ArchitectureLambda ArchitectureMaster Data Management (Mdm)Microservices ArchitectureRedshiftSparkSQLStreaming Pipelines
Artificial Intelligence • Cloud • Software
Build and maintain scalable data products using Snowflake, dbt, and Airflow. Responsibilities include dimensional data modeling, SQL transformations, pipeline orchestration, data quality testing, exploratory analysis, SLA monitoring, documentation, and stakeholder requirement gathering. Partner with analytics, product, engineering, and business teams to deliver reliable, production-ready datasets while following software engineering, CI/CD, testing, and deployment practices.
Top Skills:
Apache AirflowDbtPythonSnowflakeSQL
Fintech • Financial Services
Leads data engineering initiatives involving ETL/ELT pipelines, database optimization, orchestration, big data technologies, data modeling, BI performance, and data quality monitoring. Requires strong Scala, Python, SQL, and shell scripting skills, plus payments-domain expertise covering transaction processing, settlements, disputes, reconciliation, gateways, acquiring, issuing, and PCI-DSS compliance. The role also involves cross-functional communication, planning, facilitation, negotiation, and collaboration with senior stakeholders and clients.
Top Skills:
AirflowBusiness IntelligenceData ModelingData Quality FrameworksEltETLHadoopJavaLuigiMonitoringNoSQLPci-DssPrefectPythonScalaShell ScriptingSparkSQLTest Automation
What you need to know about the Pune Tech Scene
Once a far-out concept, AI is now a tangible force reshaping industries and economies worldwide. While its adoption will automate some roles, AI has created more jobs than it has displaced, with an expected 97 million new roles to be created in the coming years. This is especially true in cities like Pune, which is emerging as a hub for companies eager to leverage this technology to develop solutions that simplify and improve lives in sectors such as education, healthcare, finance, e-commerce and more.



%20(1).png)