Control-LM (101165402)

  https://cordis.europa.eu/project/id/101165402

  Horizon Europe (2021-2027)

  Controlling Large Language Models

  ERC STARTING GRANTS (ERC-2024-STG)

  artificial intelligence

  2024-11-01 Start Date (YY-MM-DD)

  2029-10-31 End Date (YY-MM-DD)

  € 1,500,000 Total Cost


  Description

Large language models (LMs) are quickly becoming the backbone of many artificial intelligence (AI) systems, achieving state-of-the-art results in many tasks and application domains. Despite the rapid progress in the field, AI systems suffer from multiple flaws inherited from the underlying LMs: biased behavior, out-of-date information, confabulations, flawed reasoning, and more. If we wish to control these systems, we must first understand how they work, and develop mechanisms to intervene, update, and repair them. However, the black-box nature of LMs makes them largely inaccessible to such interventions. In this proposal, our overarching goal is to: *Develop a framework for elucidating the internal mechanisms in LMs and for controlling their behavior in an efficient, interpretable, and safe manner.* To achieve this goal, we will work through four objectives. First, we will dissect the internal mechanisms of information storage and recall in LMs, and develop ways to update and repair such information. Second, we will illuminate the mechanisms of higher-level capabilities of LMS to perform reasoning and simulations. We will also repair problems stemming from alignment steps. Third, we will investigate how training processes of LMs affect their emergent mechanisms and develop methods for fine-grained control over the training process. Finally, we will establish a standard benchmark for mechanistic interpretability of LMs to consolidate disparate efforts in the community. Taken as a whole, we expect the proposed research to empower different stakeholders and ensure a safe, beneficial, and responsible adoption of LMs in AI technologies by our society.


  Complicit Organisations

1 Israeli organisation participates in Control-LM.

Country Organisation (ID) VAT Number Role Activity Type Total Cost EC Contribution Net EC Contribution
Israel TECHNION - ISRAEL INSTITUTE OF TECHNOLOGY (999907720) IL557585585 coordinator HES € 1,500,000 € 1,500,000 € 1,500,000