04 ago
|
Ntt Data Europe & Latam
|
Lazio
04 ago
Ntt Data Europe & Latam
Lazio
ppNTT DATA is looking for an Enterprise Search Engineer (Elastic), to perform on remote or onsite from Rome for a UN Agency.
You will work closely with ICT teams to deliver a well-engineered functioning enterprise search with Artificial Intelligence (AI) augmentation based on Enterprise Elastic search.
/ppIn this role you will be responsible for the following tasks: /pulliDesign and implement scalable data ingestion pipelines and connectors to ingest structured, semi-structured, and unstructured content from enterprise sources (SharePoint, Liferay, Web crawls, Data Lake, Corporate Systems, etc.) into Elasticsearch or an equivalent search index, supporting batch, incremental, and near-real-time indexing processes.
/liliDesign mechanisms to track document versions, source provenance, access permissions, timestamp updates, and deletion events to keep the search index accurate and current; develop content extraction pipelines for PDF, Word, Excel, PowerPoint, HTML, emails, scanned documents, and other enterprise formats, converting them into standard markdown and/or vector embeddings to improve AI readability.
/liliDesign and implement semantic chunking strategies and hybrid search logic optimized for retrieval quality (chunk size, overlap, section-aware splitting, heading preservation, table handling, context retention), and implement metadata extraction, enrichment, and deduplication of content during ingestion.
/li /ulh3Build Retrieval Capabilities: /h3ulliDevelop hybrid search capabilities combining keyword-based search, semantic vector search, metadata filtering, and contextual retrieval, complemented by re-ranking pipelines using specialized embedding models, ranking logic, or other suitable re-ranking techniques to improve relevance of retrieved results.
/liliImplement advanced retrieval techniques such as query rewriting, query expansion, function or tool calling, multi-query retrieval, metadata-aware retrieval, parent-child retrieval, contextual document embeddings, and contextual compression, along with security controls to ensure users can only retrieve and access relevant information they are authorized to view.
/li /ulh3Build RAG Pipelines: /h3ulliDesign and build the Retrieval-Augmented Generation (RAG)
pipeline that retrieves relevant enterprise content and uses large language models and/or specialized AI models to generate grounded answers, including agentic workflows where the AI application can invoke tools, perform multi-step reasoning, call enterprise APIs, refine searches, and retrieve additional context to answer user queries.
/liliImplement prompt engineering and orchestration patterns for reliable and relevant response generation, including system prompts, retrieval prompts, guardrails, context assembly, and response formatting, along with fallback strategies for insufficient context, ambiguous questions, or low-confidence retrieval results.
/li /ulh3Your profile: /h3ulliFirst level University degree IN Computer Science, Computer Engineering, Information Systems or related with at least 5 years of professional work experience, of which minimum 3 years hands?on experience building enterprise search, AI-powered search, semantic search, Retrieval-Augmented Generation, or LLM-based applications is required.
/liliElasticsearch engineering.
Proven and deep hands?on experience with query DSL, BM25 tuning, function_score, boosting/decay functions, and multi-field matching strategies amongst other Elasticsearch features.
/liliIndex data modelling architecture.
Ability to design mappings, choose the right field types, and configure custom analyzers/tokenizers per content type (e.g., code vs. prose vs. structured records vs. multimedia content).
/liliConnector / ingestion pipeline.
Real experience building or configuring pipelines for SharePoint, Liferay, databases, and Azure Data Lake - including incremental sync/CDC, handling deletes/updates, and dealing with rate limits and API intricacies of each source.
/liliHybrid semantic search (lexical, vector, ELSER or similar).
br/Proven experience of implementing hybrid search including a deep understanding of semantic search and keyword search.
/liliPerformance, scaling cluster operations.
Shard strategy,
index sizing, reindexing strategy, query latency tuning, and general cluster health management /liliSearch evaluation relevance testing methodology.
Experience of building a ground truth and gold-standard benchmark of relevant samples and test queries with expected results, measure precision/recall and/or NDCG and similar evaluation metrics, and iterate against it.
/liliProficiency in Python and experience with data processing frameworks and libraries.
/liliStrong hands?on experience with Elasticsearch, OpenSearch, Azure AI Search, or similar enterprise search platforms.
/liliExperience implementing semantic chunking strategies that split content into contextually coherent sections while preserving headings, structure, metadata, and parent?document relationships to improve retrieval accuracy in RAG applications.
/liliExperience designing ingestion pipelines that convert enterprise documents into structured Markdown using tools such as Marker, Docling, or equivalent document?conversion frameworks to preserve layout, tables, headings, and metadata for downstream RAG indexing.
/liliExperience with embedding models, re?ranking models, cross?encoders, prompt engineering, context window management, and response grounding techniques.
/liliExperience with LLM orchestration frameworks such as LangChain, LlamaIndex, Haystack, or equivalent frameworks.
/liliExperience with tool calling, agentic workflows, function calling, multi?step retrieval, and AI application orchestration.
/liliExperience working with commercial or open?source LLMs, such as Azure OpenAI, OpenAI, Anthropic, Google Gemini, Meta Llama, Mistral, Jina, or similar models.
/li /ulpKnowledge of React and/or similar front-end technologies and frameworks for web application development /pulliFluent in English-language (both written and verbal) is required.
/li /ulpNTT DATA is a multinational consultancy company, employing over ******* professionals world-wide.
Within the International Institutions we have framework contracts with European Institutions like: European Commission; European Parliament; European Court of Auditors; Europol; NATO; Court of Justice; EPO; European Council, United Nations, etc.
/p /p #J-*****-Ljbffr
📌 Enterprise Search Engineer (Elastic)- Elasticsearch / Data Modelling- Un Agency (Lazio)
🏢 Ntt Data Europe & Latam
📍 Lazio