We work on enterprise data platforms and scalable pipelines on Google Cloud. We build reliable systems for ingesting, transforming, and analyzing large volumes of data.
Your role:
Design and develop batch and streaming data processing pipelines using Apache Beam and Apache Spark on Google Cloud Dataproc .
Model, optimize, and query large datasets in BigQuery , with a focus on cost and performance.
Work with BigQuery Studio to explore data, develop analytical notebooks, and collaborate with data science and analytics teams.
Monitor data quality and implement tests and alerts on production pipelines.
Work with the Data Science team to ensure that data is accessible, reliable, and well-documented.
Help define the team's data engineering standards: naming conventions, data catalog, and data lineage.
Your profile:
BigQuery — Advanced SQL, Optimization — Intermediate/Advanced Level
Python — Pipelines and Scripting — Intermediate Level
Advanced SQL — Window Functions, CTE, Optimization — Intermediate Level
Git/CI-CD — Collaborative Workflow — Intermediate Level
Nice to have:
Experience with Google Cloud Dataflow (fully managed Beam pipelines).
Knowledge of pipeline orchestrators: Apache Airflow / Cloud Composer .
Introduction to data modeling : star schema, snowflake schema, Data Vault.
Google Cloud Professional Data Engineer certification (or currently pursuing it).
What we offer:
Permanent contract (National Collective Bargaining Agreement for the Retail Sector).
Company canteen
Structured professional development plan .
J-18808-Ljbffr
📌 Data Engineer (Italia)
🏢 Altro
📍 Italia