← Browse all jobs

S

Principal Data Engineer

Styli · India

FULL TIME

Job Description

Location: Bangalore (Yamlur)

Experience: 6-10 Years

Role: Principal Data Engineer

About Styli Marketplace

Launched in 2019 by Landmark Group, Styli Marketplace is the first e-commerce venture of the group, quickly becoming a leading online destination for fashion and lifestyle across the GCC, including Saudi Arabia, the UAE, Kuwait, Bahrain, and beyond. Styli connects global sellers and creators with millions of fashion-forward customers, offering the latest trends, exceptional value, and convenient services like same-day to 48-hour delivery and flexible payment options. Our mission is to make style accessible, aspirational, and exciting for all, backed by a passionate team fostering a culture of creativity and innovation.

Role Overview

We are looking for a Principal Data Engineer to take technical ownership of Styli's Data Platform and to lead its evolution from a BigQuery-centric warehouse to a modern Lakehouse on Databricks . This is a hands-on architecture and leadership role: you will set the technical direction for the platform, drive the migration end-to-end (design, staffing, partner management, cutover), and mentor a team of senior and staff data engineers. You will work closely with Platform/DevOps, Security, ML, and business stakeholders across the GCC and India to ensure the platform scales reliably, cost-efficiently, and securely as Styli grows.

What You'll Do

Data Platform Architecture & Migration Leadership

  • Own the end-to-end target-state architecture for Styli's data platform, leading the migration from BigQuery to a Databricks Lakehouse built on Delta Lake / Apache Iceberg, Unity Catalog, and GCS.
  • Lead build-vs-buy and platform evaluations (e.g., Databricks vs. Snowflake), and translate the recommendation into an executable, phased migration roadmap with clear cutover and rollback criteria.
  • Define and defend architecture decision records (ADRs) covering table format strategy, catalog ownership, and cross-engine interoperability between Databricks and existing warehouses via the Iceberg REST catalog.
  • Own the staffing and delivery model for the migration — partnering with Databricks implementation/SI partners, structuring hybrid onsite/offshore teams, and tracking cost, timeline, and risk.
  • Drive Unity Catalog adoption for centralized governance, fine-grained access control, and data lineage across workspaces and markets.

Data Pipeline Development

  • Design, build, and maintain batch and real-time pipelines migrating from Airflow/BigQuery patterns to Databricks Workflows and Delta Live Tables, alongside Apache Airflow and dbt where appropriate.
  • Handle ingestion from diverse sources: MySQL/PostgreSQL, Kafka event streams, S3/GCS data lakes, REST APIs, and SaaS platforms.
  • Ensure pipelines are fault-tolerant, recoverable, and instrumented with data quality checks at every stage, with clear migration parity/reconciliation testing against legacy BigQuery pipelines.

Data Modelling & Lakehouse Engineering

  • Design and own the medallion (bronze/silver/gold) data architecture on Delta Lake, including dimensional and analytical models (star schema, OBT, wide tables) for BI and ML consumption.
  • Own warehouse/lakehouse performance engineering — partitioning, Z-ordering/clustering, Photon acceleration, and cost-efficient query and compute (cluster policy) patterns on Databricks, alongside the legacy BigQuery estate during transition.
  • Set and enforce standards for dbt models, tests, documentation, and lineage across a well-governed transformation layer.

Streaming & Real-Time Data

  • Build and maintain real-time pipelines for live inventory updates, order event processing, and customer behavior streams, using Databricks Structured Streaming, Apache Flink, Spark Streaming, or ksqlDB.
  • Design Kafka topic schemas, partition strategies, and consumer group management for high-throughput e-commerce event flows.

Data Lake & Cloud Infrastructure

  • Build and manage a well-structured lakehouse on GCS, with Databricks running natively on GCP alongside existing cloud-native data services.
  • Partner with the Platform/DevOps team on infrastructure-as-code (Terraform) for data infrastructure provisioning and containerized data workloads on Kubernetes.

Data Quality & Observability

  • Set the data quality framework (Great Expectations, dbt tests, or Monte Carlo) to catch anomalies, nulls, duplicates, and schema drift before they reach downstream consumers — with particular rigor during migration cutover windows.
  • Own pipeline SLAs, freshness SLOs, and schema-change alerting as first-class platform metrics; own incident response for data pipeline failures with clear runbooks and escalation paths.

ML & Analytics Enablement

  • Build and maintain feature pipelines for ML models powering personalisation, recommendations, fraud detection, and demand forecasting, migrating toward Databricks Feature Store / MLflow where it improves training–serving consistency.
  • Enable self-serve analytics by maintaining clean semantic layers and well-documented data marts for BI tools (Looker, Metabase, Superset).

Team Leadership & Mentorship

  • Set technical standards and best practices for the data engineering team; review designs and code for scalability, cost, and reliability.
  • Mentor senior and mid-level data engineers, including team members based in Bengaluru, and act as the technical escalation point for the platform.
  • Partner directly with Platform Engineering, Security, and business stakeholders to align the migration roadmap with compliance and cost objectives.

What We're Looking For

Required

  • 6-10 years of hands-on data engineering experience, including demonstrated ownership of platform-level architecture, ideally in a high-transaction e-commerce, fintech, or consumer tech environment.
  • Proven, hands-on experience leading a large-scale migration onto Databricks — including Delta Lake, Unity Catalog, Databricks Workflows/Delta Live Tables, and cluster/cost optimization.
  • Strong proficiency in SQL — complex analytical queries, window functions, CTEs, query optimisation, and warehouse-specific dialects (BigQuery / Databricks SQL / Redshift).
  • Deep experience architecting lakehouse platforms on GCP, including evaluating and integrating open table formats (Apache Iceberg or Delta Lake) and cross-engine interoperability patterns.
  • Hands-on experience with Apache Airflow for orchestration and dbt for transformation layer management, at scale — handling millions of events per day with reliability and observability.
  • Working knowledge of Apache Kafka or equivalent event streaming platforms for real-time data ingestion.
  • Proficiency in Python for pipeline development, data transformation, and automation, including PySpark for distributed processing on Databricks.
  • Track record of technical mentorship and cross-functional stakeholder management, including partner/vendor (SI) management for large migration programs.
  • Understanding of data modelling principles: normalisation, dimensional modelling, medallion architecture, and analytical patterns.

Good to Have

  • Experience running Databricks and BigQuery (or Snowflake) in parallel during a phased migration, including reconciliation and cutover strategy design.
  • Familiarity with data cataloguing and governance tooling (DataHub, Amundsen, Collibra, or Alation) beyond Unity Catalog.
  • Knowledge of data mesh or data product principles and federated ownership models.
  • Experience with real-time analytics databases: ClickHouse, Druid, or Pinot for high-throughput OLAP workloads.
  • Familiarity with infrastructure-as-code (Terraform) and deploying data workloads on Kubernetes.
  • Relevant certifications: Databricks Certified Data Engineer Professional, Databricks Lakehouse Architect, Google Professional Data Engineer, or dbt Analytics Engineering.

The Data Problems You'll Solve at Styli

  • Leading the BigQuery-to-Databricks lakehouse migration — architecting and executing the move for a live, high-traffic e-commerce platform with zero-downtime tolerance.
  • Personalisation at scale — building event pipelines that capture every click, view, and purchase to power real-time recommendations for millions of customers.
  • Inventory & demand forecasting — reliable pipelines feeding ML models that optimise stock levels across hundreds of SKUs and multiple markets.
  • Campaign & marketing analytics — ingesting data from paid channels, CRM, and app attribution to give growth teams a single source of truth.
  • Order & payments intelligence — real-time event streams for order lifecycle tracking, fraud signals, and settlement reconciliation.
  • Cross-market reporting — unified data models serving multiple country markets (GCC and India) with different currencies, catalogues, and fulfilment providers.

Details

CompanyStyli
LocationIndia
TypeFULL TIME
Nichetech

Similar Jobs

C

TOSCA Automation Engineer (SAP)

CloudLabs Inc

S

Senior UPS Service Engineer

STARLITE COMPUTERS SERVICES

C

FullStack Web and AI Developer

CarTrade Tech Ltd.

S

Principal Data Engineer

Serko

U

Printed Circuit Board Design Engineer

Univision Technology Solutions Pvt Ltd (UTS) )