ph3About the Role /h3 ul liMercor is partnering with a leading AI research lab to support a Frontier Code Agents project. /li liContributors help evaluate and improve frontier AI coding models through structured technical assessments. /li liThe work focuses on realistic infrastructure engineering workflows and model evaluation. /li liSpots are limited and filling quickly on a first come, first serve basis. /li /ul h3What You'll Do /h3 ul liUse frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks. /li liReview model-generated implementations involving cloud platforms, Kubernetes, CI/CD systems, observability, and infrastructure automation. /li liIdentify bugs, edge cases, reliability issues, and failure modes. /li liCompare outputs from multiple frontier models and assess their strengths and weaknesses. /li liApply professional engineering judgment to realistic infrastructure engineering scenarios.
/li /ul h3Time Commitment /h3 ul liSprint based project that runs in 12-24 hour stretches based on client requirement. /li /ul h3Compensation /h3 ul li$400 per accepted task. /li liTypical tasks take approximately 2–3 hours after ramp-up. /li liCompensation is tied to accepted work. /li /ul h3Who Should Apply /h3 ul li2+ years of professional DevOps, SRE, or Cloud Engineering experience. /li liExperience with AWS, Azure, GCP, Kubernetes, Terraform, CI/CD pipelines, or observability tooling. /li liRegular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools. /li liAbility to evaluate model-generated infrastructure and reliability engineering solutions. /li liExperience supporting production-scale systems is preferred. /li /ul /p #J-18808-Ljbffr
📌 DevOps Engineer - AI Model Evaluator (Roma)
🏢 Obsidian
📍 Roma