We are hiring Lead Data Engineer – Analytics Platform & Technical Leadership
📍 Location: Toronto
💼 Experience: 8+ Years
About the Role
Key Responsibilities
🔹 Lead the ingestion, transformation, aggregation, and processing of large-scale datasets for analytics and downstream consumption.
🔹 Design and maintain scalable, reliable, and high-performance data pipelines across Hadoop, Databricks, and enterprise data platforms.
🔹 Drive data unification initiatives by integrating structured and semi-structured data sources.
🔹 Work with high-volume, high-velocity, and high-dimensional datasets using modern big data frameworks and cloud-native technologies.
🔹 Analyse transactional and product data to generate actionable insights and support business growth.
🔹 Partner with Product Managers, Data Scientists, Platform Strategy, and Technology teams to translate business and analytical requirements into scalable engineering solutions.
🔹 Act as a technical bridge between business, analytics, and engineering teams.
🔹 Identify innovation opportunities and deliver POCs, prototypes, and pilot solutions.
🔹 Provide technical leadership, mentorship, and guidance to data engineers and analysts.
🔹 Establish best practices around data modelling, pipeline design, performance optimisation, data quality, governance, and maintainability.
🔹 Influence architecture, engineering standards, and long-term data platform sustainability.
Required Technical Skills
✅ 8+ years of experience in Data Engineering, Big Data Analytics, or Enterprise Data Platforms.
✅ 2+ years of experience in a Lead or Technical Leadership role.
✅ Strong proficiency in Python, Pandas, NumPy, and PySpark.
✅ Hands-on experience with Impala and Hadoop-based platforms.
✅ Strong SQL skills with experience in relational and distributed data stores.
✅ Experience with ETL/ELT and data integration tools such as Apache Airflow, Apache NiFi, or Azure Data Factory.
✅ Strong experience in data modelling, data mining, querying, and reporting over large datasets.
✅ Experience with cloud-based data platforms such as Azure/AWS, Databricks, and/or Snowflake.
✅ Experience working with data lakes, distributed computing, and cloud storage services.
✅ Experience implementing CI/CD pipelines and DevOps practices for data engineering workflows.
✅ Strong understanding of data quality, governance, security, and enterprise data management.
GenAI / LLM Experience – Preferred
⭐ Experience building scalable data pipelines supporting GenAI/AI products and solutions.
⭐ Experience with batch and streaming data ingestion and transformation.
⭐ Exposure to processing unstructured and semi-structured data such as documents, logs, and text.
⭐ Understanding of PII handling, access controls, privacy, security, and auditability in AI data environments.
⭐ Familiarity with operationalising AI data workflows, including monitoring, data quality, reproducibility, and cost optimisation.
⭐ Exposure to machine learning concepts, feature engineering/calculations, and model serving is a plus.
Read Full Description