pstrong Your mission /strongp Bring a tender from the source portal into Cato: scraping, parsing, merging, enrichment. You'll start by owning a handful of sources end to end - the scraper, the job behind it, and the data that comes out - and take on more as you go. Not tickets handed to you: sources you're responsible for. /pp La preghiamo di verificare di possedere il livello di esperienza e le qualifiche adeguate leggendo la panoramica completa di questa occasione qui sotto. brpstrong What you'll actually do /strong /pulli Build and maintain scrapers for national tender portals, where reading the source in its original language is part of the job. /lili Keep them alive: portals change their HTML, move endpoints, break pagination, throttle you. You find out before the customer does. /lili Turn messy sources into clean records: broken HTML, inconsistent XML, APIs that lie about their own schema. /lili Write and maintain orchestrator flows: retries, backfills, alerting, and a clear answer to "Did today's run actually land?" /lili Work on merge and dedup - the same tender arrives three times, in three shapes, and only one version can reach the customer. /lili Ship AI enrichment steps: batch LLM extraction of requirements, embeddings, OCR on attachments. /lili Guard data quality with tests and checks that fail loudly before a customer finds the gap. /li /ul pstrong Ideal profile /strong /pullistrong Python that holds up: /strong typed, tested, and readable six months later. /lilistrong You've scraped something real: /strong HTML and XML parsing, pagination, sessions,
rate limits - and you know why a scraper that worked yesterday is broken this morning. /lilistrong SQL you're comfortable in: /strong joins, aggregations, window functions. You'll read from the database every day; tuning and running it isn't your job. /lilistrong Builder by default: /strong you see a manual process and your first instinct is to automate it. /lilistrong Comfortable with messy sources: /strong broken HTML, inconsistent XML, PDFs that were scans of scans. /lilistrong You close your own loop: /strong you check that what you shipped actually ran, before someone else has to ask. /li /ul pstrong Experience /strong /pulli1-2 years writing Python in production: scrapers, ETL scripts, automation - anything that had to run unattended and be fixed when it didn't. /lili Exposure to an orchestrator (Prefect, Airflow, Dagster) is a plus, not a requirement: you'll learn ours properly. /lili Exposure to LLM-based extraction is welcome; curiosity about it is mandatory. /li /ul pstrong What you won't find here /strong /pulli No micromanagement: we trust you to own your part of the stack. /lili No "standard" 9-to-5 mentality: we care about outcomes and we are looking for people who are willing to go the extra mile. /lili No "we've always done it this way" excuses: we're here to disrupt, not to follow old patterns. xrdbqlu /li /ul pstrong Our Tech Stack /strong /pulli Data Infra: Python, PostgreSQL, Prefect, AWS /lili AI: batch LLM extraction, embeddings, OCR /li /ul pstrong Compensation /strong /pp RAL €35,000 - €45,000 + equity, depending on profile. /p pstrong Hiring Manager /strong /pp Lorenzo Rossetto /p /p /p
📌 Junior data engineer - Italia (Milano)
🏢 Cato
📍 Milano