AI Verification and Evaluation Research Institute (AVERI) Logo

AI Verification and Evaluation Research Institute (AVERI)

Research Scientist

Posted An Hour Ago
Be an Early Applicant
In-Office or Remote
Hiring Remotely in United Kingdom
Mid level
In-Office or Remote
Hiring Remotely in United Kingdom
Mid level
Conduct independent research on frontier AI evaluation and auditing. Design rigorous evaluations, audit methods, and reproducible experimental pipelines; analyze model capabilities, risks, robustness, and alignment-relevant behavior. Produce audit reports, recommendations, publications, tools, and presentations while collaborating with technical, policy, engineering, and standards teams. The role requires careful handling of confidential information and may involve pilot audits with frontier AI companies.
The summary above was generated by AI
About AVERI:

The AI Verification and Evaluation Research Institute (AVERI) is a US-based nonprofit with the mission to make third-party auditing of frontier AI effective and universal. AVERI aims to bring about a world in which the most powerful AI systems – and the companies that build them – are rigorously audited for safety and security by third parties. We believe this independent auditing layer is essential for enabling confident deployment of AI and managing critical risks from this increasingly powerful technology.

About the role:

As a Research Scientist at AVERI, you will conduct research on frontier AI systems and the safety and security practices surrounding them, often as part of pilot audits with frontier companies, sometimes working externally to those companies and sometimes embedded inside them. The work sits at the practical and scientific foundations of assessing model capabilities, risks, robustness, and alignment-relevant behavior in ways that are technically deep, decision-relevant, and useful to the broader AI governance ecosystem. You will work closely with our policy, engineering, and standards teams, and you will help shape a new field as it develops. We are hiring at two levels: Research Scientist and Research Director (see below). Given AVERI’s early stage (launched in January 2026), early staff take on a wide range of responsibilities.

What You’ll Do:
  • Own a line of research on frontier AI evaluation and auditing, from scoping through delivery, within AVERI's research agenda.

  • Design and run evaluations of closed and open-weight AI systems, including as part of pilot audits with frontier AI companies.

  • Develop audit methods and evaluation protocols for model behavior, capabilities, safety properties, robustness, deployment risks, or compliance with stated policies.

  • Build reproducible experimental pipelines for testing models, analyzing outputs, and interpreting results.

  • Produce rigorous audit reports and translate technical findings into actionable recommendations for AI developers, policymakers, auditors, and the broader public.

  • Publish and present completed research as conference papers, white papers, open-source tools, talks, and public writing.

  • Collaborate with researchers, engineers, policy experts, and standards bodies inside and outside AVERI, and contribute to research planning and methodological standards for frontier AI audits.

  • Handle confidential access to company systems and practices and sensitive findings with appropriate care.

About you:
  • A track record of independent research in AI/ML, preferably in AI safety, evaluations, robustness, or security (roughly three or more years beyond a PhD or equivalent, or a comparable record of delivered work).

  • Strong technical judgment: you can design rigorous empirical studies and reason carefully about measurement, uncertainty, and evidence.

  • Experience working with large language models or other modern AI systems, especially in safety-relevant contexts.

  • You can turn open-ended questions into tractable research plans and communicate clearly with collaborators while working independently in a remote environment.

  • You are motivated by AI safety, AI governance, and third-party auditing, and want to contribute to it full-time.

Strong candidates may also have:
  • Prior work on AI auditing, model evaluations, red teaming, interpretability, robustness, privacy, cybersecurity, or responsible AI.

  • Published research or high-quality public technical writing, and experience building reproducible evaluation pipelines, benchmarks, or research infrastructure.

  • Familiarity with AI governance, standards, regulatory processes, or institutional audit practices, and experience communicating technical findings to non-technical audiences or working in interdisciplinary teams.

  • Domain knowledge in areas relevant to AI risk, such as cybersecurity, biosecurity, misinformation, privacy, legal compliance, or model misuse.

Research Director

We are also open to hiring this role at the Research Director level. The Research Director owns and sets AVERI’s technical research agenda, priorities, and roadmap. They would also represent AVERI’s research externally to frontier labs, standards bodies, policymakers, and the research community, and would help to recruit a larger research team (which they may manage directly or just provide technical leadership to). We expect a Research Director to bring roughly seven or more years in the field, an exceptional profile, a publication record peers recognize, and a demonstrated record of setting research direction that others followed. If you think you might be a strong fit at this level, please say so in your application.

Someone with several years of independent research experience who will lead their own line of work will typically come in as a Research Scientist. Level and title are set during the interview process.

What we offer:
  • Salary: Your salary depends on the scope, autonomy, and impact we expect you to have at AVERI. The salary ranges for this role are:

    • Research Scientist: $220,000 to $450,000

    • Research Director: $330,000 to $600,000

  • Comprehensive medical benefits, including generous reimbursement of eligible HRA expenses

  • 6% unrestricted retirement contribution

  • Unlimited PTO

  • A work enablement and professional development package of up to $22,200 per year

  • A mission-driven team committed to integrity and purpose, providing a supportive environment with meaningful opportunities for career growth

  • Location and travel: Fully remote, with a preference for candidates in the San Francisco Bay Area. Travel at least twice per quarter and possibly significantly more, especially to Washington, DC and San Francisco.

If you think you are perfect for this role, but you do not meet all of the requirements, we strongly encourage you to apply and tell us why you’d be a great fit. We are committed to fostering a diverse and inclusive environment where all employees have the opportunity to succeed. We do not discriminate based on race, religion, color, national origin, gender, sexual orientation, or any other legally protected status. Please let us know what accommodations you might need when applying or if asked to be interviewed by emailing us at [email protected]

Similar Jobs

4 Days Ago
Remote
Junior
Junior
Professional Services • Consulting
Conduct original research on reliable agentic AI, developing methods for simulation, evaluation, optimization, regression control, and continuous learning. Analyze agent failures, traces, and human feedback; create experiments and prototypes; and translate research into production-facing systems. Collaborate with product and engineering teams, contribute to technical strategy, and provide rigorous experimental evidence. The role requires strong research depth, publication or open-source contributions, hands-on Python development, and expertise in LLM agents or agentic systems.
Top Skills: Agent FrameworksAi AgentsContinuous LearningLarge Language Models (Llms)PythonReinforcement LearningSimulation Systems
14 Days Ago
Remote
Mid level
Mid level
Artificial Intelligence • Fintech • Machine Learning • Software • Financial Services
Conduct applied research on foundation models for real-time fraud detection using large-scale behavioral and sequential data. Design experiments, benchmarks, holdouts, and monitoring systems; develop models from data preparation through deployment; and optimize training and inference efficiency. Collaborate with engineering, client-facing teams, legal, compliance, and customers to productionize models and establish explainability, governance, and risk documentation.
Top Skills: Deep LearningDistillationEmbedding StoresFeature StoresFine-TuningFoundation ModelsGpu ComputingModel MonitoringModel ServingModel VersioningPythonQuantizationReal-Time InferenceSelf-Supervised LearningSQLTokenization
One Month Ago
In-Office or Remote
Mid level
Mid level
Angel or VC Firm
Founding research hire to develop clinical reasoning and computer-use models using proprietary hospital and clinic data. Own end-to-end training and evaluation pipelines, design benchmarks exposing LLM weaknesses on clinical reasoning, run experiments, and publish technical reports and papers. Collaborate closely with founders and deploy research into products that improve physician workflows.
Top Skills: LlmsPython

What you need to know about the Edinburgh Tech Scene

From traditional pubs and centuries-old universities to sleek shopping malls and glass-paneled office buildings, Edinburgh's architecture reflects its unique blend of history and modernity. But the fusion of past and future isn't just visible in its buildings; it's also shaping the city's economy. Named the United Kingdom's leading technology ecosystem outside of London, Edinburgh plays host to major global companies like Apple and Adobe, as well as a growing number of innovative startups in fields like cybersecurity, finance and healthcare.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account