11 ago
|
Bitrock
|
Italia
ppBitrock is a high-end consulting and system integration company, bstrongly committed to offering cutting-edge and innovative solutions /b. Our tailored consulting services enable our clients to preserve the value of legacy investments while migrating to more efficient systems and infrastructure. We take a holistic approach to technology: we consider each system in its totality, as a set of interconnected elements that work together to meet business needs. /p pWe thrive on overcoming challenges to help our clients reach their goals, by supporting them in the following areas: Data, AI ML Engineering; Back-end Engineering, Platform Engineering, Front-end Engineering, Product Design UX Engineering, Mobile App Development, Quality Assurance, FinOps, GovernanceThe effectiveness our solutions also stems fromb partnerships /b with key technology vendors, like HashiCorp, Confluent, Lightbend, Databricks, and Meterian. /p h3Who we are looking for /h3 pWe're looking for a bDatadog Observability Specialist Freelance /b to own the end-to-end rollout of a primary observability platform for a multi-region, multi-tenant environment on Azure Kubernetes. /p pYou will lead the migration from a mature but fragmented self-hosted stack (Prometheus/Thanos, Grafana, ELK, AppDynamics) to Datadog. The project is past the proof-of-concept phase: the agent is deployed via Helmfile, a Terraform repository for Datadog resources is bootstrapped, and a tagging taxonomy is agreed upon. We need a genuine Datadog expert who knows the platform's intricacies and sharp edges to drive this rollout to completion and leave behind an autonomous, self-service platform for our client's engineering teams. /p h3What You'll Do /h3 ul libRollout and agent configuration /b — Take the agent from a handful of QA environments to every cluster and tier, as code in Helm/Helmfile,
including host-based instrumentation on the legacy Tomcat and httpd VMs. /li libTagging and service ownership /b — Enforce Unified Service Tagging and our internal taxonomy across agent, chart and application CD sources, and populate the Software Catalog with real team ownership. /li libEverything as code /b — Own the Terraform estate (DataDog/datadog provider) for monitors, dashboards, SLOs, log indexes and RBAC, wired into Azure DevOps pipelines with remote state. /li libMonitoring, alerting, incident routing /b — Design actionable P1–P4 monitor tiers and SLOs, route them through On-Call and Microsoft Teams, and run the dual-run and cutover from Prometheus. /li libIntegrations /b — Connect Azure, Kubernetes, Kafka/Confluent with Data Streams Monitoring, the databases, the JVM/Node/.NET runtimes, RUM, and the OpenTelemetry and Splunk HEC pipelines. /li libCost governance /b — Control spend through log index and retention strategy, exclusion filters, metric cardinality limits and per-team usage attribution. /li libSecurity and access /b — Design RBAC and Restricted Datasets, and keep log scrubbing, PII redaction and SQL obfuscation a day-one constraint on regulated data. /li libEnablement /b — Leave behind runbooks, naming conventions and onboarding docs, and run workshops so application teams and partner vendors can operate the platform themselves. /li /ul h3Requirements /h3 ul libDeep, demonstrable Datadog expertise /b — at least one large multi-team,
multi-environment rollout personally driven, not just product usage. You know the sharp edges (Terraform provider gaps, API limits around restriction policies, index ordering, cardinality billing traps) and can talk about them concretely. /li liStrong command across bAPM, Logs, Infrastructure, RUM, Database Monitoring, Data Streams Monitoring, SLOs, Software Catalog, On-Call, Watchdog /b. /li libTerraform /b at professional level, specifically the Datadog provider — module design, remote state, CI/CD integration. /li libKubernetes /b in production, ideally bAKS /b, with Helm and Helmfile. /li libOpenTelemetry /b — collector configuration, processors, filters, exporters, and a clear view of where OTel ends and vendor instrumentation begins. /li liInstrumentation across a mixed estate: JVM (Spring Boot and legacy Tomcat/Java EE), Node.js, .NET. /li libAzure /b fundamentals and bAzure DevOps /b pipelines. /li liMigration experience from Prometheus/Grafana, ELK/Elastic APM, AppDynamics, Dynatrace or New Relic — including the political part, not just the technical part. /li liFluent English, and the ability to write a design document other engineers actually read. /li /ul h3Nice to have /h3 ul liKafka / Confluent Cloud operations. /li liSplunk, particularly HEC ingestion patterns. /li liRegulated-data environments (GDPR and equivalents). /li liWorking alongside multiple system integrators on a shared platform. /li liFinOps or observability cost optimization with real numbers attached. /li /ul pbDuration: /b 1 year /p pbStart: /b September /p h3Recruitment process /h3 pOur recruitment process has 3 stages: /p ul liFirst discovery short interview with our HR team /li liTechnical interview with our Team Leaders /li liFinal interview with our Client /li /ul /p #J-18808-Ljbffr
📌 Senior Datadog Consultant - Freelance (Italia)
🏢 Bitrock
📍 Italia