07 ago
|
Mindrift
|
Italia
pMindrift is building a dataset to evaluate AI coding agents by creating challenging tasks and robust evaluation criteria within realistic environments. You will craft prompts, define success criteria, and assemble tests that accept multiple valid solutions while rejecting incorrect ones. /ppThis project-based role seeks experienced developers to design tasks, implement tests, and iterate based on QA feedback, contributing to fair, robust evaluation methodologies for AI agents. /p #J-18808-Ljbffr
📌 AI Evaluation Engineer — Shape Real-World Coding Benchmarks (Italia)
🏢 Mindrift
📍 Italia