Build and operate scalable batch and streaming data pipelines, connectors, and RAW-to-CUBE transformations using Python, Spark, Databricks, Kafka, dbt, SQL, and Delta Lake. Ensure data quality, schema validation, observability, performance, cost efficiency, and reliable recovery. Develop governed data products and feature-serving interfaces for machine learning models, agents, and downstream applications.
Job Summary
CONVO is seeking a Senior Data Engineer to build reliable data pipelines and connectors powering its CPG-focused agentic platform. You’ll develop scalable batch and streaming pipelines using Python, Spark, Databricks, Kafka, dbt, and Delta Lake, transforming source data into governed RAW and CUBE layers while ensuring data quality, performance, observability, and reliability. You’ll also build data-product and feature-serving interfaces that enable ML models, agents, and downstream applications to consume trusted data efficiently.
Technical mission- Build reliable connectors and RAW-to-CUBE pipelines across batch and streaming paths, and expose governed data and ML features to downstream services.
- Develop Python/Spark ingestion and transformation pipelines from source systems into RAW and Enriched CUBE layers.
- Implement streaming paths alongside scheduled workloads.
- Create dbt/SQL transformations, data-quality controls and contract-validation checks.
- Optimize Delta Lake layout, SQL performance, partitioning and incremental processing.
- Build feature-serving and data-product interfaces for models, agents and applications.
- Instrument pipelines for lineage, freshness, throughput, failure recovery and cost.
- Minimum experience: 5+ years in data engineering, including 3+ years delivering production Spark or python pipelines.
- Advanced Python and production Apache Spark/Databricks development.
- Kafka or comparable event-streaming technology.
- Strong SQL performance tuning and dimensional/data-product implementation.
- dbt and Delta Lake, including incremental patterns and table optimization.
- Batch and streaming reliability patterns: idempotency, checkpointing, replay and late-arriving data.
- Automated data quality, observability and schema-contract testing.
- ML feature stores or online/offline feature consistency.
- CPG, retail, ERP, POS or syndicated-data pipelines.
- Kubernetes-based data workloads and cloud cost optimization.
- Production connectors and RAW-to-CUBE pipelines with automated tests.
- Batch/streaming operational dashboards and recovery procedures.
- Documented data products and feature-serving interfaces.
- Performance and cost baselines for the implemented workloads.
- Works under the Data Architect with source-system owners, ML/optimization teams, platform infrastructure and downstream application teams.
Similar Jobs
Artificial Intelligence • Fintech • Machine Learning • Software • App development • Conversational AI • Generative AI
Migrate SQL Server data products and analytical workloads to a modern cloud platform. Redesign legacy SQL, ETL processes, and business logic; build scalable pipelines and analytical models; validate data quality; collaborate with stakeholders; and improve platform reliability, performance, cost efficiency, and AI-assisted data engineering automation.
Top Skills:
BigQueryCi/CdCloud ComposerETLGCPGoogle Cloud StorageIamMs Sql ServerPythonSQL
Agency • Information Technology
Build and maintain Snowflake and SQL Server data pipelines for an operational data store, including ingestion, transformation, schema modeling, migrations, CI/CD, data quality testing, and legacy batch-to-API modernization. Collaborate with domain teams and architects to support customer, loan, payment, and interaction data. Ensure data accuracy, lineage, observability, and performance while applying generative AI tools to accelerate development and testing.
Top Skills:
AlationApache AirflowCollibraDagsterDbtFlywayGenerative AiGitGithub ActionsGitlab CiGreat ExpectationsInformaticaPostgresPrefectPythonRest ApisSchemachangeSnowflakeSnowflake Dynamic TablesSnowflake TasksSQLSQL ServerTerraform
Blockchain
Design, deploy, and maintain AWS and Databricks data infrastructure, data lakes, and scalable ETL pipelines. Build custom ingestion services and connectors using Airbyte, Go, and Python; optimize performance, reliability, monitoring, and data quality. Implement cloud data governance, security, access controls, encryption, and compliance practices while collaborating with platform engineers, analysts, and stakeholders.
Top Skills:
AirbyteAmazon RedshiftAmazon S3AWSAws GlueAws IamAws LambdaDatabricksGoPythonSQL
What you need to know about the Pune Tech Scene
Once a far-out concept, AI is now a tangible force reshaping industries and economies worldwide. While its adoption will automate some roles, AI has created more jobs than it has displaced, with an expected 97 million new roles to be created in the coming years. This is especially true in cities like Pune, which is emerging as a hub for companies eager to leverage this technology to develop solutions that simplify and improve lives in sectors such as education, healthcare, finance, e-commerce and more.



