01 ott
|
Cti Clinical Trial And Consulting Services
|
Lazio
01 ott
Cti Clinical Trial And Consulting Services
Lazio
pbThe work /b /p pIcaro Foundation is an independent non-profit AI safety lab based in Rome.
We study advanced AI systems: what they can do, how they fail, and how those findings can support developers and institutions responsible for their governance.
/p pbWe see AI safety as one of the defining scientific and societal challenges of our time.
/b As AI systems become more capable, autonomous, and widely deployed, understanding and reducing their risks is increasingly urgent.
We are looking for people who are deeply interested in these questions and motivated to contribute through rigorous research.
/p pYou will bhelp produce new research and develop the lab's shared codebase and knowledge base /b, working closely with our researchers across the research process: reviewing literature, refining questions, implementing experiments, analysing results, and contributing to papers and technical reports.
/p pOur research focuses particularly on bagentic, multi-agent, and compositional safety /b: how risks emerge across extended interactions, tool use, and systems involving multiple AI agents.
We also study btesting awareness and evaluation validity /b, including whether models behave differently when they recognise that they are being evaluated.
/p pAlongside our research, we evaluate frontier models for international model providers as independent third-party evaluators, using public and proprietary benchmarks and red-teaming environments.
/p pOur public work includes: /p ul liBoiling the Frog, on multi-turn agentic safety; /li liAdversarial Humanities Benchmark, on the robustness of safety behaviour under stylistic reformulations; /li liresearch on LLM-to-LLM risks, multi-agent collusion, and interaction-level safety.
/li /ul pYou can explore our research programme and papers to learn more.
/p pbWhat you would do /b /p pYour work will combine three closely connected areas.
/p pbContribute to research /b /p ul liReview relevant literature, compare methods, and identify questions worth investigating.
/li liHelp turn research questions into experimental protocols, including baselines, controls, and clear evaluation criteria.
/li liImplement and run experiments with frontier and open-weight models, including agentic and multi-agent environments.
/li liAnalyse results and model traces, investigate unexpected behaviour, and assess confounders and alternative explanations.
/li liContribute to research papers, benchmarks, technical reports, and presentations.
/li /ul pbDevelop the research codebase /b /p ul liWrite and improve Python code for experiments, evaluations, data processing, and analysis.
/li liExtend existing tools and environments, fix bugs, and participate in code review.
/li liAdd tests,
documentation, and reproducible configurations so other researchers can inspect, rerun, and build on your work.
/li /ul pbBuild the lab's knowledge base /b /p ul liProduce concise, source-grounded notes on papers, methods, benchmarks, and research questions.
/li liDocument experimental setups, findings, limitations, and negative results.
/li liOrganise and connect references, datasets, code, and research notes so the team can find relevant evidence and reuse previous work.
/li /ul pYou may bring stronger skills in research or engineering.
The role involves both writing code and reasoning carefully about evidence.
/p pbWho should apply /b /p pWe welcome applications from bmaster's students, PhD students, recent graduates, and researchers at the beginning of their careers /b, including those who have recently completed a PhD.
/p pRelevant experience may come from a thesis, academic research, independent experiments, open-source contributions, internships, or previous employment.
We also welcome applicants from non-traditional backgrounds who can demonstrate strong research or engineering ability.
/p pA completed PhD, previous AI safety employment, and published papers are not required.
We care about the quality of your work, your contribution to it, and your ability to learn.
/p pIf you are currently studying, please tell us about your availability and how you would combine the role with your academic commitments.
/p pbWhat we are looking for /b /p ul liExperience with LLM evaluations, red-teaming, or benchmark development; /li liA solid technical or quantitative background, developed through university study, independent projects, or relevant work.
/li liGood Python skills and familiarity with Git, debugging, and working with an existing codebase.
/li liPractical experience with machine learning or LLMs through at least one substantive project or research contribution.
/li liAn understanding of basic experimental reasoning and statistics: comparing conditions, interpreting results, and recognising uncertainty and possible confounders.
/li liThe ability to read technical papers critically and explain methods, findings, and limitations clearly in English.
/li libA strong interest in AI safety and a sense of urgency about understanding and reducing the risks posed by increasingly capable AI systems.
/b /li liIntellectual curiosity, openness to criticism, and a willingness to revise your views in response to evidence.
/li /ul pbUseful, not required /b /p pExperience with: /p ul liagentic or multi-agent systems; /li listatistical analysis or experimental replication; /li lisoftware testing, containers, or reproducible research workflows; /li liliterature reviews, research documentation, or open-source contributions; /li liInspect AI, the open-source evaluation framework developed by the UK AI Security Institute and Meridian Labs, or comparable tools.
/li /ul pFor an example of our research software, see the Adversarial Humanities Benchmark codebase, also listed in Inspect Evals as an externally maintained evaluation.
/p pYou do not need experience in all of these areas.
/p pbHow we work /b /p pWe are a small research team.
You will work closely with experienced researchers and receive feedback on experimental design, code, analysis, and writing.
/p pYou will begin with clearly scoped contributions to ongoing projects and take on greater responsibility as your skills and familiarity with the work develop.
We encourage everyone to ask questions, challenge assumptions, and propose ideas.
/p pExisting evaluation infrastructure, technical support, and API budget are available.
Contributions may become public papers, benchmarks, datasets, or tools where compatible with confidentiality obligations.
Authorship and acknowledgement will reflect contributions.
/p pWe value work that others can understand and build on: clear reasoning, reliable code, well-documented experiments, and honest reporting of uncertainty.
/p ul libLocation: /b Flexible, with a preference for working in person with the team in bRome, Italy /b.
/li libIn-person collaboration: /b We particularly welcome applicants who are based in Rome or would be interested in relocating.
We value regular in-person discussion, collaborative experimentation, and learning from one another.
/li libRemote arrangements: /b May be considered for candidates based in Europe or China, with substantial overlap with European working hours.
/li libEngagement: /b Contractor role.
/li libCompensation: /b The specific compensation range will be shared during the first interview.
/li /ul pReferrals increase your chances of interviewing at Icaro Foundation by 2x /p pFind curated posts and insights for relevant topics all in one place.
/p h3Seniority level /h3 ul liMid-Senior level /li /ul h3Employment type /h3 ul liContract /li /ul h3Job function /h3 ul liEngineering and Information Technology /li liIT System Testing and Evaluation /li /ul #J-*****-Ljbffr
📌 Ai Safety Researcher (Lazio)
🏢 Cti Clinical Trial And Consulting Services
📍 Lazio