03 ott
|
BRAINTRUST
|
Italia
Help a top AI lab evaluate and improve large language models through security-focused coding tasks. Bring your software-engineering judgment and hands-on security experience to work involving vulnerabilities, exploit verification and security patches.
This is a contracting engagement, with potential for a longer-term engagement. Remote candidates in the selected countries are elegible.
What you will do
• Evaluate coding tasks involving software vulnerabilities, exploit verification and security patches.
• Create high-quality coding prompts and reference answers for benchmark-style problems.
• Evaluate model outputs for code generation, refactoring, debugging and implementation.
• Identify and document model failures, edge cases and reasoning gaps.
• Compare private language models with leading external models.
• Build or configure coding environments for evaluation and reinforcement learning.
• Follow detailed annotation and evaluation guidelines consistently.
What you bring
• At least five years of professional software-development experience and strong Python skills.
• Hands-on experience with vulnerability research, exploit reproduction or verification, or implementing, backporting or validating security patches.
• The ability to apply structured evaluation criteria and write clear technical feedback.
• Fluency in written and spoken English.
Helpful, not required
• Professional code review, coding annotation, LLM/code evaluation or benchmark design.
• Knowledge of another programming language.
• Team leadership or mentoring experience.
📌 Senior Application Security Engineer (Italia)
🏢 BRAINTRUST
📍 Italia