Junior data engineer - Italia (Milano)

Junior data engineer - Italia (Milano)

12 set
|
Cato
|
Milano

12 set

Cato

Milano

Your mission

Build and run the pipelines that carry a tender from the source portal to the customer's screen: ingestion, merge, enrichment, delivery. You own concrete pieces of the data pipeline end to end — not tickets handed to you, but the sources, jobs and tables behind them.

What you'll actually do

- Own scrapers and ingestion for a set of national portals across Spain and Italy, where reading the source in its original language is part of the job.
- Write and maintain orchestrator flows: retries, backfills, alerting, and a clear answer to "Did today's run actually land?"
- Work on merge, dedup and reconciliation — the same tender arrives three times, in three shapes, and only one version can reach the customer.
- Ship AI enrichment steps: batch LLM extraction of requirements, embeddings, OCR on attachments.
- Write SQL that survives production: query plans, indexes, JSONB, partitioning, CONCURRENTLY migrations.
- Guard data quality with tests and checks that fail loudly before a customer finds the gap.

Ideal profile

- Real SQL: you can read an EXPLAIN and say why the plan is bad, not just that it is slow.
- Python you'd put in production: typed, tested, and readable six months later.




- Pipelines you've actually operated: with an orchestrator (Prefect, Airflow, Dagster, ArgoWorkflows) and the 3 a.m. failures that come with them.
- Builder by default: you see a manual process and your first instinct is to automate it.
- Comfortable with messy sources: broken HTML, inconsistent XML, PDFs that were scans of scans.

Experience

- At least 1 year building data pipelines in production.
- Hands-on with PostgreSQL beyond writing queries — you've had to make one fast.
- Exposure to LLM-based extraction is welcome; curiosity about it is mandatory.

What you won't find here

- No micromanagement: we trust you to own your part of the stack.
- No "standard" 9-to-5 mentality: we care about outcomes and we are looking for people who are willing to go the extra mile.
- No "we've always done it this way" excuses: we're here to disrupt, not to follow old patterns.

Our Tech Stack

- Data & Infra: Python, PostgreSQL, Prefect, K3s/ArgoCD, AWS
- AI: batch LLM extraction, embeddings, OCR

Compensation

RAL €35,000 – €45,000 + equity, depending on seniority and profile.

📌 Junior data engineer - Italia (Milano)
🏢 Cato
📍 Milano

Candidati a questo annuncio

Mostra le tue capacità professionali all'azienda, compila il form e lascia un tocco personale nella lettera di presentazione, aiuterà il recruiter nella scelta del candidato.

Iscriviti a questa job alert:

Ricevi via email le nuove offerte di lavoro per: junior data engineer - italia (milano) / milano

Iscriviti a questa job alert:

Ricevi via email le nuove offerte di lavoro per: junior data engineer - italia (milano) / milano