11 ago
|
Mindrift
|
Italia
Mindrift is building a dataset to evaluate AI coding agents by creating challenging tasks and robust evaluation criteria within realistic environments. You will craft prompts, define success criteria, and assemble tests that accept multiple valid solutions while rejecting incorrect ones.
This project-based role seeks experienced developers to design tasks, implement tests, and iterate based on QA feedback, contributing to fair, robust evaluation methodologies for AI agents.
#J-18808-Ljbffr
📌 AI Evaluation Engineer — Shape Real-World Coding Benchmarks (Italia)
🏢 Mindrift
📍 Italia