Skip to content
ReferMeAJob
LiveRemoteFull-timeApply by 1 Nov 2026

Data Engineer

Remote

Experience
4–6 years
Employment
Full-time
Work mode
Remote
Salary
Not disclosed
Deadline
Apply by 1 Nov 2026
Posted
2026-08-12

Required skills

SkillExperienceLevel
Python1+ yearsIntermediate
SQL2+ yearsIntermediate
PostgreSQL2+ yearsIntermediate
JavaScript1+ yearsIntermediate
AWS3+ yearsIntermediate
Azure3+ yearsIntermediate
Typescript1+ yearsIntermediate
CI/CD3+ yearsIntermediate
ETL(Extract, Transform, Load)4+ yearsIntermediate
Snowflake2+ yearsIntermediate
Data Engineer4+ yearsIntermediate
Databricks4+ yearsIntermediate

About the role

About the role

The Data Engineer in our AI & Data team will be responsible for designing and building the data

structures and pipelines our AI Engineers rely on across Azure, Snowflake, Databricks, and

Lakebase (our managed Postgres / OLTP layer). The primary mission of this role is to enable the AI

Engineering team- translating the needs of machine-learning and computer-vision workflows into

reliable, well-modelled, and cost-effective data foundations.

Tasks include setting up new data pipelines and transformations, ingesting structured and

unstructured data into the data lake and warehouse, monitoring the performance and cost

effectiveness of existing data jobs, and docking machine-learning processes into the existing data

landscape. The Data Engineer works hand in hand with AI Engineers and is the go-to person for

making trusted data available for models, products, and analytics.

Main Responsibilities

● As part of the AI & Data team, design and build the data structures, schemas, and models

that AI Engineers depend on for training, feature engineering, and inference.

● Develop and orchestrate scalable data pipelines on Databricks (Spark, Delta Lake) and load

curated, analytics-ready data into Snowflake.

● Own data ingestion, transformation (ELT/ETL), and storage across the Azure cloud (e.g.

ADLS, Data Factory, Event Hubs/Synapse), including structured, semi-structured, and

unstructured data such as text, images, and video.

● Dock machine-learning and computer-vision models into the data pipelines, and design the

data flow that feeds and consumes those AI services.

● Sync curated lakehouse data into Lakebase (managed Postgres) for low-latency serving,

manage change-data-capture back into Delta tables, and support online feature stores and

agent state for AI Engineers.

● Build and maintain API integrations and automated data ingestion from internal systems

and external third-party sources.

● Monitor pipeline performance, reliability, and cost; troubleshoot failed jobs and optimize

Snowflake and Databricks workloads.

● Implement data quality, validation, and lineage, and document the data dictionary and ETL

processes.

● Partner with AI Engineers and stakeholders to translate model and business requirements

into extensions of the data platform.

Skills, Qualifications & Education

● Bachelor’s degree in Computer Science, Data Engineering, or a related field.

● At least 4 years of work experience in data engineering or a similar data-focused role.

● Hands-on production experience with Databricks (Apache Spark, Delta Lake, notebooks,

workflows).

● Hands-on production experience with Snowflake (data modeling, performance tuning,

access control, cost management).

● Solid experience with Microsoft Azure data services (e.g. ADLS, Data Factory, Event Hubs /

Synapse).

● Experience with PostgreSQL and OLTP databases; familiarity with Lakebase (Databricks

managed Postgres) is a strong plus.

● Working knowledge of JavaScript / TypeScript, used for data APIs, microservices, or app

facing integrations.

● Strong expertise in SQL (will be tested during the recruitment process).

● Robust Python literacy, especially for data handling and pipeline development.

● Comfortable working with both structured and unstructured data; does not shy away from

troubleshooting failed ETL processes or API integrations.

● Outstanding data-structure and data-modeling design skills.

● Working knowledge of machine-learning, NLP, or computer-vision workflows is a plus.

● Experience with dbt, Airflow, or Databricks Workflows, and with CI/CD and infrastructure

as-code, is a plus.

● Strong ability to translate ideas between technical and non-technical audiences.

● Curious, collaborative, self-motivated, and organized; able to run multiple projects against

tight deadlines.