General information
The Priority Area (German: Schwerpunktprogramm) “Robust Assessment & Safe Applicability of Language Modeling: Foundations for a New Field of Language Science & Technology” (acronym: LaSTing; SPP 2556) aims to advance our understanding of language technology, in particular language modeling, for safer use, especially in applications in the (computational / cognitive) language sciences. A detailed project description is here.
Aims and scope of the Priority Area LaSTing
While modern language technology increasingly permeates many areas of applications, much of its input-output behaviour and its inner mechanics remains unknown. As a result, recent years have seen a newly emerging field of interdisciplinary and methodologically diverse work at the interface between the cognitive language sciences and language technology.
Robust assessment
Read more
Given the very rapid pace of recent developments, careful reflection on standards for the methodology of testing and assessment is lagging behind. What is required is a joint effort to converge on proper standards for robust assessment of language models. Methodology is robust, in the sense intended here, if its results are generalisable (carrying over with sufficient certainty to other models and data sets), transferable (insightful beyond the purposes of understanding a single type of computational model), and reproducible (with the same or different models and data sets). Robust methodology also aspires to be as future-proof as possible, i.e. likely relevant to the next generation of models or the next set of antagonistic examples.Safe applicability
Read more
As language technology gets applied more and more widely, concerns of safe applicability become ever more important. Safe applicability subsumes critical aspects such as being conceptually sound (e.g. anchored in “first principles” or established empirical knowledge), validated (e.g. by mathematical proof or other rigorous derivation) or at least stress-tested across a near-exhaustive traversal of possible conditions of use, ethical (e.g. bias- and harm-free, or privacy-respecting), and also economical (i.e. minimising data requirements and energy consumption). Issues of safe applicability loom particularly large in the context of high-stake implications, of which application in the scientific process is a special case. The Priority Area LaSTing therefore also particularly invites contributions on the reflection of safe applicability of language technology for knowledge gain in the cognitive language sciences.Foundational questions
Read more
Progress on understanding the behaviour of language models and their safe applicability is inexorably tied to a better understanding of their core mechanisms and the impact of their training data or their training objectives. But just as relevant are deep foundational questions concerning the nature of language models (e.g. what are LMs models of?) and their proper role in the scientific research into human language (e.g. how could LMs be used as explanatory tools for understanding human language?). In response to these issues, the Priority Programme especially welcomes foundational work addressing general properties or potential limits of particular classes of language models, e.g. by using mathematical arguments, simulation studies, tight conceptual argumentation or a mixture of such methods.Projects
The Priority Area LaSTing involves the following projects:
- Limits and Biases in Machine and Human Language and its Learning(PIs: Artemis Alexiadou, Uli Sauerland)Despite their human-like appearance in some tasks, language models (LMs) differ substantially from humans in their structure and the structure of their training. The Libille project aims to improve our understanding of these structural differences and their effects. To do so it formulates predictions and then attempts to validate predictions concerning differences and similarities in human and LM behaviour focussing on the morphosyntactic domain. The project explores both the mature state, i.e. adult humans vs. trained LLMs, as well as the learning/training stages of human and LLMs. Furthermore it applies structural probing to assure the explanations of LM behavior are the predicted ones. Central features of LM design are the tokenizer, embeddings, the transformer architecture, and gradient descent learning. From this architecture, we argue a number of different behavior in morphosyntax in the mature state are predicted. Three we will focsu on are those concern morphological paradigm gaps, morphological generalization to invented (‘nonce’) words, sensitivities to a languages writing conventions for example with definite markers and clitic pronouns. The project will devise novel experiments to test LMs and humans for such predicted differences. The LM architecture also predicts that during training, LMs should exhibit different properties from human children acquiring language. We develop different paradigms to test for the predicted learning differences using in-context learning, artificial grammar learning, and other techniques. The three predictions we focus on are the absence of the typical stages of human language acquisition in LMs, an LM ability to learn generalizations impossible in human languages, and that training to induce human-like learning biases should improve LMs learning. The comprehensive understand of human-machine difference in linguistic abilities that the LIBILLE project develops will provide crucial insights to our understanding of both the human language ability and the abilities of current LMs.
- The pragmatic test: how humans and LLMs decode presupposed meaning (PIs: Nadine Bade, Miriam Butt) This project investigates the robustness and reliability of Large Language Models (LLMs) in the interpretation of presuppositions through parallel human-machine experimentation. Presuppositions play a central role in strategies of persuasive communication such as framing, by introducing information indirectly into the common ground. This mechanism of accommodation is integral to the persuasiveness of the presupposed content and helps explain why expressions like again are use frequently in political slogans, policy documents and advertisements. While descriptive work has highlighted the role of presuppositions in persuasive and deceptive communication, there is little experimental research on the persuasiveness of presupposed content, and current computational studies show that LLMs struggle with presupposition resolution in contexts where misinformation risk is high. This project will establish a systematic empirical foundation by comparing human and LLM performance on presupposition resolution tasks across communicative settings. It will develop computational models that integrate both lexical triggers and discourse context, using the results to inform LLM learning and improve performance on presupposition detection and resolution. We combine computational work on the automated detection and annotation of framing phenom ena with experimental studies on presuppositions to assess whether LLMs can be applied safely in detecting presupposition-based biases in high-stakes communicative contexts. To do so, we will (i) evaluate LLM performance through behavioral assessment aligned with human experiments, (ii) explore whether semantic parsing can supplement training and fine-tuning of LLMs to improve presupposition resolution, and (iii) develop new resources, including annotated corpora and methodological guidelines, for experimental and computational research on presuppositions. The long-term goal is to determine to what extent LLMs can be used as tools to detect presupposition based biases in high-stakes communicative settings. At the same time, we anticipate a contribution to theoretical pragmatics, as formal models of presuppositions are extended and refined through the results of parallel human-machine experiments.
- KIND-LM: Cognitively-inspired interaction dynamics for sample-efficient language modeling (PIs: Lisa Beinborn, Nivedita Mani)Computational models of language can generate remarkably fluent text, but their impressive performance comes at the cost of training on trillions of tokens with unsustainable computational resources. When trained under academic resource constraints, such models fall short of robust linguistic generalization and often fail to adapt to unseen contexts. Human learners, by contrast, acquire language from vastly smaller input and can flexibly adapt to new communicative situations from an early age. A central difference lies in the learning signal: while human acquisition is embedded in rich social interactions, language models are typically optimized for the narrow task of next-word prediction. This project develops a cognitively grounded approach for interactive language modeling that integrates feedback mechanisms inspired by child–caregiver communication. We propose a training setup in which a child model improves its linguistic competence through interaction with a more powerful parent model. Unlike existing teacher–student approaches, which assume unilateral feedback, we focus on the temporal and linguistic interaction dynamics and on the interaction initiative. We build on our winning submission to the new interaction track of the BabyLM Challenge, which used a reinforcement loop and showed that even simplified feedback strategies can enhance functional linguistic competence without sacrificing formal accuracy. We propose to better align computational modeling with psycholinguistic evidence and systematically test cognitively more plausible interaction strategies. We will draw on mechanistic interpretability methods to better understand how interaction dynamics influence the representational structure of the model and how they can improve its ability to generalize to the long tail of the vocabulary distribution. Our project advances research on cognitively inspired sample-efficient modeling and contributes to the Priority Programme LaSTing by using language technology as a simulation framework to deepen our understanding of human language learning.
- A Resource Efficient Cross-linguistic Approach to Figurative Meaning Assessment in LLMs (PI: Maria Berger)LLMs can find good mappings for word meanings that are well encoded by large amounts of data. However, they are not able to interpret figurative meaning, and it is unclear whether future architectures will. In fact, there will always be a lack of data representing transferred meaning. Recent models handle known figures very well, e.g.: “birds of a feather flock together”, finding the correct translation: “gleich und gleich gesellt sich gern” (DE). However, they struggle with unknown figures: “the biter is sometimes bitten,” and literally translate them. We also do not know whether LLMs correctly interpret figurative meaning in less-studied languages because studies are lacking. To address these issues, we conduct an evaluation using parallel multilingual figurative corpora defining three objectives: First, we conduct a robust assessment of multilingual LLMs to test their ability to capture figurative meaning. Using existing corpora, we apply downstream tasks, such as machine translation and zero-shot figurative meaning NLI. Available corpora will be plentiful for some tasks (e.g., metaphor prediction) and scarce for others (e.g., parallel proverb detection in German and Chinese). We use literal rephrasing and back-translation strategies to supplement existing resources. We want to assess how LLMs behave depending on available data and whether they are able to cope with figurative meaning. Second, we examine internal representations of LLMs to understand how translation effects a figure’s meaning and how we can measure meaning through encoding. This can be achieved by probing-testing whether a model performs a certain task. To this end, we apply techniques including neuron activation probing and activation vector transformation (e.g. using SAE). We use layer freezing as a complementing approach in which neurons are activated according to certain patterns depending on the input. This will tell us how the inner mechanisms of an LLM effects the output and ultimately, we can extrapolate missing meaning representations, thus nut cracking LLMs’ black-boxed behavior using it to our advance. Third, we aim to enhance cross-lingual capabilities of LLMs. Exceeding pure evaluation, we distill transferable aspects of figurative across languages. It is particularly important to understand the limitations of today’s architectures that perform well in resource-rich languages. Since figurative meaning will never be fully represented by LLMs across all languages, we attempt to transfer language-agnostic elements of figurative meaning. We define a low-difficulty challenge: cross-lingual metaphor NLI, and a high-difficulty challenge: understanding proverbs and idiomatic usage. The latter is critical, because languages lexicalize figurative expressions differently, depending on their tradition, hence parallel corpora barely exist. We offer test suites for both difficulty degrees. Since our approaches use joint resources, data processing and model training are very efficient.
- Unreal engines — Understanding language models through resource-optimal analysis: Implicit Bayesian pragmatic reasoning & emergent causal world models (PI: Michael Franke)This project investigates whether language models genuinely understand language, drawing on philosophy of science (Dellsén’s dependency-model account of understanding) and linguistic pragmatics (Gricean speaker meaning as Bayesian inference). It proposes analyzing autoregressive LMs as resource-optimal solutions to next-token prediction, extending prior resource-rationality work by incorporating architectural efficiency and internal representations, not just behavior. Using simulations and latent-variable recovery methods from statistics, it asks whether LMs would be expected to recover causal generative processes and exhibit human-like pragmatic reasoning, enabling robust, generalizable assessment across model classes.
- Attention in Large Language Models: Linguistic Grounding, Cognitive Modeling, and Social Application (PIs: Nicole Gotzner, Sebastian Musslick)In this project, we probe the central hypothesis that the attention mechanism in LLMs can be understood in terms of a computational analogue of human text comprehension. Building on this hypothesis, our project addresses key problems of the LaSTing framework of robust assessment, safe applicability, and foundational understanding. Specifically, we pursue three objectives: (O1) grounding the attention mechanism in linguistic, cognitive, and psycholinguistic theory of text comprehension (foundational understanding), (O2) applying data-driven discovery to derive attention-based metrics and evaluate their predictive validity in terms of behavioral and neural correlates of human text comprehension (robust assessment), and (O3) applying discovered metrics to assess and personalize the comprehensibility of scientific and medical text in a transparent and interpretable manner (safe applicability).
- A multidimensional adaptive test for the psychometric assessment of LLM capabilities (PI: Fritz Günther)With the rapid rise of Large Language Models (LLMs), we see new models being released on a constant basis. This is accompanied by the equally fast release of new benchmark datasets to assess the performance of these models in various domains – from language processing and problem solving to more specialized capabilities such as emotion detection and theory of mind. In this dynamic environment, assessing the performance of each new model on the entire item pool of each relevant benchmark is not only a technical challenge, but raises fundamental concerns about the scalability, resource demands, and sustainability of benchmark performance assessment. In the present project, we address these issues by adopting a multidimensional item response theory (mIRT) framework developed in psychometric assessment to LLM benchmarking. In the IRT framework, population-invariant and item-specific difficulty and discrimination parameters of each individual item are estimated from the empirical performance of a norming sample, which allows us to identify the most informative items for assessing LLMs’ latent abilities. The mIRT framework extends this towards multiple different ability dimensions. Here, we will collect responses of a norming sample of LLMs on a diverse set of benchmark items from various domains, and use mIRT to estimate the item parameters. We will then use the most informative items to implement a computerized adaptive test (CAT) for LLM capabilities: Here, items are presented successively until the LLM capability parameters are estimated with sufficient confidence, allowing for a maximally efficient capability assessment. This assessment infrastructure – which will be designed as a future-proof “living environment” where new items can be added and those that turn uninformative over time can be removed – will be made available in the form of local software packages as well as via an online interface.
- Structural generalization in transformer-based LLMs (PIs: Michael Hahn, Alexander Koller)With the rapid rise of Large Language Models (LLMs), we see new models being released on a constant basis. This is accompanied by the equally fast release of new benchmark datasets to assess the performance of these models in various domains – from language processing and problem solving to more specialized capabilities such as emotion detection and theory of mind. In this dynamic environment, assessing the performance of each new model on the entire item pool of each relevant benchmark is not only a technical challenge, but raises fundamental concerns about the scalability, resource demands, and sustainability of benchmark performance assessment. In the present project, we address these issues by adopting a multidimensional item response theory (mIRT) framework developed in psychometric assessment to LLM benchmarking. In the IRT framework, population-invariant and item-specific difficulty and discrimination parameters of each individual item are estimated from the empirical performance of a norming sample, which allows us to identify the most informative items for assessing LLMs’ latent abilities. The mIRT framework extends this towards multiple different ability dimensions. Here, we will collect responses of a norming sample of LLMs on a diverse set of benchmark items from various domains, and use mIRT to estimate the item parameters. We will then use the most informative items to implement a computerized adaptive test (CAT) for LLM capabilities: Here, items are presented successively until the LLM capability parameters are estimated with sufficient confidence, allowing for a maximally efficient capability assessment. This assessment infrastructure – which will be designed as a future-proof “living environment” where new items can be added and those that turn uninformative over time can be removed – will be made available in the form of local software packages as well as via an online interface.
- Systemic Robustness Assessments of Language Models for Cross-Linguistic Research using Formally Related Structures (FORESTS) (PIs: Jutta-Maria Hartmann, Anke Himmelreich, Sina Zarrieß)The central objective of this project is to develop a novel interdisciplinary approach that leverages language models (=LMs) as tools for cross-linguistic research and linguistic theories as tools for systemically assessing LMs’ robustness. For this goal, we operationalize linguistic theories to assess how robust an LM’s “holistic” syntactic knowledge is, by moving from evaluation on single phenomena to systemic assessments of networks of formally-related structures (=FORESTs). FORESTs are a network of abstract structures that share underlying syntactic properties within languages and/or across languages. For example
Who does Peter like _ best?' andWhat do you think that Mary bought _?’ share the dependency of a fillerWho/What' to a gap (_) but di er with respect to the presence of embedding. We use such networks of frequent and grammatical filler-gap dependencies and compare them to infrequent and ungrammatical island-configurations as well as infrequent but grammatical parasitic gap constructions likeWho did you kiss _ without knowing _?’, where an illicit gap in an island becomes well-formed due to a gap outside the island. Based on a theoretically informed sets of FORESTs, we develop systemic assessment procedures that test for the presence of ``holistic’’ syntactic knowledge in an LM. We further develop robustness scoring of these assessments for families of models that, in the next step, allow to test predictions of di erent theoretical analyses of parasitic gaps. Current theoretical analyses of parasitic gap structures make di erent predictions as to which other structures are close members in a network of FORESTs. We use these di erences in theories to compare results of acceptability judgments of parasitic gaps and related structures in humans with LMs’ performance on these structures, by manipulating training data input to include di erent forests. We will first set up this procedure for a set of theoretically well-described FORESTs and languages. Our main goal is to connect LMs’ assessments and cutting-edge cross-linguistic research, focusing on the theoretically challenging case of parasitic gaps. Bringing together theoretical linguistic knowledge and computational expertise in LMs, the project addresses the research questions of the Priority Programme LaSTing in various ways. First, the project contributes to robust assessment by designing benchmark materials in a more theory driven and generalizable way, including a cross-linguistic perspective. Second, experiments that vary input, model size, and architecture will lead to a better understanding of the limits of syntax learning in LMs and their transferability to other languages. In the long run, these insights can contribute to making LMs more resource e icient and sustainable. Finally, the project aims to conduct research on foundational questions regarding the explanatory power of LMs for linguistic theory building. - Moral Hallucinations in Large Language Models — Their Argumentative Structure and Ethical Implications (PIs: Annette Hautli-Janisz, Karoline Reinhardt)Our starting point are three observations: First, many people use chatbots based on Large Language Models (LLMs) for controversial and ethically relevant questions that go far beyond the private sphere of individual decision-making, e.g., Is it wrong to lie in order to protect someone’s feelings? (Wester et al. 2025). Secondly, recent research shows that participants rate an LLM’s moral advice as superior to the advice of other people (Aharoni et al. 2024) and even to that of expert ethicists (Dillion et al. 2025). Thirdly, LLMs have been shown to contain moral biases (Takemoto et al. 2024; Xu et al. 2025; among many others) and to exhibit different moral codes than humans (Marraffini et al. 2024; Garcia et al. 2024; Bonagiri et al. 2024). While current research focuses on a simplistic analysis of LLM responses to moral questions, e.g., yes/no responses or moral vs. immoral judgements (Jha et al. 2024, Ji et al. 2024, among others), one crucial property of LLM-generated reasoning is overlooked: moral hallucinations. These hallucinations are not reducible to conventional AI hallucinations, since they are not only about factual inaccuracies or lack of faithfulness to sources. Instead, they involve distortions within patterns of moral reasoning, constituting a qualitatively different issue – with potentially highly relevant consequences, because when LLMs distort moral concepts, they might undermine not only the content of advice, but also the structural foundations of individual and collective moral judgment. In this project, we bring, thus, together methods from philosophy, in particular Applied Ethics of AI, and computational linguistics, especially argument mining, to conceptualize, benchmark, ethically assess and automatically identify LLM-generated moral hallucinations. We structure this research based on the following three research questions: (RQ1) What are the constitutive elements of LLM-generated ‘moral hallucinations’ and what are the ethical consequences if the moral claims that LLMs produce are not just biased but hallucinatory? (RQ2) What are the core argumentative structures and reasoning patterns of moral hallucinations and how do they compare to moral reasoning in genuine philosophical theory? (RQ3) What are the wider ethical implications when an LLM-based system is used for moral advice-seeking and how can we set up a computational model so that it can automatically flag moral hallucinations and components thereof? In answering these questions, set up a benchmark of moral hallucinations that contains a fine-grained reasoning and argumentation analysis, identify the ethical implications of moral hallucinations and develop a computational model to identify these hallucinations in unseen data.
- The Status of Linguistic Constraints in Neural Language Models (PI: Erhard Hinrichs)We want to investigate the extent to which generative language models (LMs) are capable of acquiring abstract linguistic knowledge beyond the factual information presented in the training data. By abstract linguistic knowledge, we mean knowledge about linguistic mechanisms and patterns acquired by LMs without being explicitly trained for it. Recent linguistic studies suggest that LMs exhibit such abstraction capabilities over linguistic rules. Given this promising state of the art, we want to investigate whether such abstraction capabilities also extend to sets of interacting linguistic constraints in phonology and morphology that can be in conflict with one another and that thus require constraint conflict resolution.We want to explore to what extent current transformer models can already capture such constraint interactions. Specifically, we want to address the following research questions: 1) Are transformer-based generative LMs able to abstract linguistic constraints from the training data? 2) If abstraction capabilities are confirmed, how similar are the abstracted constraints to the frameworks established in the linguistic literature? 3) If abstraction capabilities are not confirmed, how does explicit insertion of constraints in the prompt change the model generations? By addressing these questions, we strive to make novel and innovative contributions to the research questions summarized under the rubrics LM capabilities, Ontological Status, and Explanatory Potential of the Priority Programme.We carefully chose three linguistic phenomena such that the complexity of the constraint space governing these phenomena is suitably diverse. All three phenomena can be formalized as a sequence-to-sequence task of generating the most probable output string given an input sequence. This makes them ideally suited for an LM analysis. Investigating the constraint interaction for each of these phenomena with LMs can provide a cue about the abstraction capabilities of LMs. Our main hypothesis for this study is: transformer-based generative LMs do abstract and generalize linguistic constraints from the training data. We furthermore predict a negative correlation between the complexity of the constraint space for a phenomenon, and the extent to which LMs abstract these constraints.To evaluate our hypotheses, we will adopt a set of methods from mechanistic interpretability studies. Specifically, we will be 1) producing possible output variants, resulting from different constraint violations on the input, and 2) inspecting the probabilities assigned to them by LMs, in different settings. This will allow us to mitigate the interference of factual knowledge and concentrate on abstract linguistic knowledge acquired by LMs.
- Evaluating, Explaining, and Enabling Ethical Multi-Agent Systems of Large Language Models (E4-MALM) (PIs: Anne Lauscher, Jae Hee Lee)Modern large language models (LLMs) have demonstrated impressive capabilities across a wide range of tasks. When equipped with additional components such as a memory or external tools, an individual LLM can function as an agent that actively perceives, reasons, and acts within a given environment. Recent work shows that orchestrating several such agents in more complex setups can further enhance effectiveness in concrete deployment scenarios and unlock new application opportunities (e.g., simulating large-scale social behaviors). However, these multi-agent settings, where multiple LLM-based agents interact, also pose new challenges: As such, multi-agent interactions can lead to emergent behaviors and unpredictable collective behaviors—both positive (e.g., coordinated problem-solving) and negative (e.g., compounding of biases or unsafe decisions). For example, recent experiments show that decentralized populations of LLM agents can spontaneously develop social conventions with increasing polarization and even collective biases that no single agent initially possessed. Such emergent behaviors are particularly concerning as multi-agent systems of LLMs (MALMs) are currently tested for high-stakes domains where unchecked biases or unsafe behaviors can cause significant harm (e.g., physical harm in the case of medical diagnosis and invalid scientific conclusions). Without proper assessment methods and governance mechanisms, such systems could make discriminatory decisions, spread misinformation collaboratively, or develop unpredictable, potentially harmful strategies that no single agent was programmed to exhibit. This raises urgent questions about robust assessment and safe applicability of language technology in multi-agent contexts, directly aligning with the DFG Priority Program LaSTing (SPP 2556). Our project responds to LaSTing’s call by focusing on three core goals: (i) evaluating current multi-agent systems of LLMs for ethical risks, (ii) explaining their individual and collective behaviors down to mechanistic internal processes, and (iii) enabling ethical behavior through principled alignment interventions. It addresses all three LaSTing core issues, i.e., robust evaluation methodologies for LLM behaviors, methods to ensure safe applicability (e.g., ethical, bias-free, and validated use), and foundational insights into the mechanisms of language models in complex settings.
- Interpretable Surprisal: Language Models Between Linguistic Structure and Neural Evidence (PIs: Alessandro Lopopolo, Milena Rabovsky)The aim of this project is to investigate the cognitive plausibility of language model–derived surprisal as a measure of human language processing. In particular, it examines how surprisal relates to linguistic structure (syntax and semantics) and to human behavioral and neural data (e.g., EEG). A central goal is to determine to what extent the relationship between surprisal and human responses is mediated or modulated by underlying linguistic information. To address these questions, we will investigate surprisal estimates derived from both pretrained large language models and models trained under different training conditions, allowing us to assess how model architecture, and training regime influence their cognitive plausibility.
- Understanding the Cross-Linguistic Brain Basis of Sentence Processing through Interpretable Language Technology (PI: Lars Meyer)Large Language Models (LLMs) enjoy increasing popularity throughout our professional and private lives. This includes scientific work, where LLMs are increasingly used as tools of investigation. One key area of inquiry is the neurobiology of language, where scientists have found that LLMs provide unprecedented fit to human neuroimaging data acquired during language comprehension. Yet, the implications of such findings remain elusive, because current LLMs lack interpretability in terms of human psycholinguistic processes. Which psycholinguistic processes does the brain activity that is being fit by LLMs correspond to? Does the fit of LLMs to human neuroimaging data actually result from human-like psycholinguistic processes that are in some way implicit to current LLMs? And if it does—can we employ interpretable LLMs to address questions that have so far been left unanswered in the neurobiology of language? The current proposal pursues three objectives along these lines. First, we lay out a strategy to make LLMs mechanistically interpretable in terms of the cognitive operations that subserve sentence processing in humans. Second, we assess whether such interpretations can shed light on the relation between neuroimaging data and LLMs. And third, we attempt to establish interpretable LLMs as scientific tools for investigating the neurobiology of language, not just in one, but in many human languages.
- Propositional Attitudes in Large Language Models (PALLM) (PI: Robert Pasternak)Decades of interdisciplinary research on the meanings of propositional attitude verbs like “believe” and “want” have furnished deep insights into how humans use language to describe mental states. This project will be the first to study how large language models (LLMs) use such attitudes, comparing LLMs to both humans and each other. Understanding LLM use of attitude verbs has the potential to inform future research in both AI and linguistics. From an AI perspective, attitudes have quietly played a crucial role in several areas of research pertaining to LLMs, such as LLMs’ purported theory of mind capabilities, as well as their use for intent classification. Moreover, many LLM prompts used in leading AI products make critical use of attitude verbs to govern the behavior of AI agents, e.g., by prompting them to act based on their own beliefs or the perceived desires of their human interlocutors. Thus, understanding how LLMs interpret attitudes is also potentially important to AI safety and alignment. On the linguistics side, seeing where LLMs succeed and fail in interpreting attitudes can give us important information about what is necessary for language acquisition. Attitudes are especially interesting because of the unique acquisition challenge they pose: “believe” and “want” do not have the same obvious physical correlates as e.g. “clap”, so humans are particularly reliant on distributional information to glean their interpretations, something that bears a partial resemblance to the statistical learning methods used to train LLMs. And if LLMs fail to emulate human language use, then this suggests that something more than sheer volume of training data is required. Comparing across LLMs is valuable here as well, as correlating differences in performance with differences in model size, architecture, and training can also provide helpful insights. On an empirical level, the project will focus on three topics pertaining to propositional attitudes. The first will be entailments in their complements: when does “X wants J” logically entail “X wants K”, and likewise for “believe”, etc.? The second will be attitude measurement constructions: what does it mean if “X wants J more than Y wants K”? The third will be cases in which the truth conditions of an attitude ascription are affected by the syntactic or pragmatic environment in which it occurs (so-called “restricted readings”). For each empirical domain, the project will feature a combination of theoretical linguistic research, human psycholinguistic research (for those phenomena where the empirical picture is not yet settled), and AI research. In addition to this purely scientific work, the project will also include the development of new open source software for linguistic research on LLMs, as well as the creation of a new benchmark dataset to evaluate current and future LLMs.
- Relating Probabilities of Words to Probabilities of Worlds (PI: Sean Papay)Large language models (LLMs) generate text by defining and sampling from a probability distribution over strings. In modeling this distribution, they acquire not only linguistic knowledge but also world knowledge, which benefits them both in autoregressive next-token prediction and in downstream tasks to which language models are applied. Although this world knowledge is vital to LLMs’ performance, we cannot observe it directly; we can only infer its properties from generated strings and their probabilities. In this project, we propose interpreting this world knowledge as a latent distribution over semantic world states that underlies the string distribution, and investigating the properties of this world distribution. Concretely, this will involve probing model probabilities for propositions conditioned on premises, using natural-language descriptions. Such an investigation will serve three major purposes: (1) to better explain models’ behavior in terms of world beliefs, (2) to improve downstream applications of LLMs by decoupling semantic beliefs from surface realizations, and (3) to develop general-purpose probability estimation models for use in cognitive modeling. Over the course of this project, we will address five major research questions: 1) How can we extract semantic probabilities from LLMs? 2) How do the extracted probabilities correspond to empirical probabilities? 3) Are extracted probabilities consistent with one another? 4) How do extracted probabilities relate to human judgments? 5) Can we reconstruct consistent belief states by augmenting LLMs with additional structure? We will answer these questions experimentally, relying on experimentation with existing LLMs and a human annotation project to elicit probability judgments. This work will provide a better frame of reference for explaining LLM behavior, tools for directly extracting semantic beliefs for downstream tasks, and general-purpose probabilistic world models for use in cognitive modeling.
- Learning linguistic inferences and their alternatives (PIs: Jacopo Romoli, Yulia Zinova)A central goal of theoretical approaches to linguistic meaning is to describe and explain the full range of inferences that speakers’ productions can give rise to. For example, an utterance of (1) yields three distinct inferences: that Jane showed up for some of the classes, (1-a), that she has a brother, (1-b), and that the two of them didn’t show up for all of the classes, (1-c). These inferences have been given different names, reflecting their different properties: the first inference is generally referred to as ‘entailment’, the second as ‘presupposition’, and the third as ‘implicature’. (1) Jane and her brother showed up for some of the classes. a. ⇝ Jane showed up for some of the classes b. ⇝ Jane has a brother c.⇝ Jane and her brother didn’t show up for all of the classes Although these major types of linguistic inference are widely recognized, their boundaries remain contested. Beyond a few clear cases, debates continue about whether certain inferences should be classified as entailments, presuppositions, or implicatures. Clarifying these boundaries is crucial for advancing our understanding of how linguistic meaning is generated. Over the past few decades, experimental investigations have created new opportunities for linguists, logicians and philosophers working within formal frameworks. In parallel, psycholinguists have drawn on these frameworks to enrich models of processing and acquisition and devised innovative methods to test these models (see Noveck 2018 for a recent overview). Moreover, the typology of inferences has been shown to extend beyond language, applying to gestures and visual animations, which suggests a general cognitive basis for these phenomena (Tieu et al. 2019b). This cross-fertilization has accelerated progress in understanding how different dimensions of meaning are derived and processed, while also providing influential insights into how semantics and pragmatics interact to shape sentence meaning and guide interpretation. In Natural Language Processing (NLP), work on Natural Language Inference (NLI) has largely focused on entailment detection (Bowman et al., 2017). Only recently has research begun extending to implicatures and presuppositions (Jeretic et al. 2020; Schuster et al. 2020; Nizamani et al. 2024). Initial findings suggest that language models can learn some inferences beyond entailments, but the scope and mechanisms of this ability remain unclear. Our project integrates machine learning, theoretical, and experimental approaches to investigate the following research questions: Q1 How well, and under what training conditions, can language models learn linguistic inferences? Q2 Are some inferences learned more easily than others? Q3 Does training on one type of inference facilitate learning others via transfer learning? Q4 Do theoretical notions proposed for the derivation of these inferences play a role in language model inference prediction? This project will advance the broader inquiry into what language models actually learn and how their behavior compares to human inference. It will also provide a way to support or challenge unified theoretical accounts of the inferences under study. For instance, a positive answer to Q3 would point to an abstract mechanism shared across inference types (cf. Bott & Chemla 2016). More broadly, the project will deepen our understanding of these inferences and the role they play in shaping our knowledge of meaning.
- LLADIGA: Learning Language with Dialogue Games (PIs: David Schlangen, Raffaella Bernardi)Large language models such as ChatGPT have shown that some form of language competence can be acquired through learning from large text corpora. This mode of language learning, however, is very different from how humans acquire language, which seems to require goal-driven interaction. The aim of LLADIGA is to bridge the gap between these two modes, by exploring the use of goal-driven and rule-governed linguistic activities (what we call “Dialogue Games”). Specifically, LLADIGA will develop techniques of multi-turn reinforcement learning, relying on feedback signals provided by such games. To be able to precisely measure the impact of learning mode, LLADIGA will design fine-grained measurement instruments that evaluate behaviour, as well as making it possible to make observations about the internal effects of different learning modes on the “developing” model. Thus, LLADIGA directly addresses several core questions of the SPP LaSTing, such as behavioural assessment, representations & mechanisms, training & optimisation, resource efficiency, and alternative models.
- The evaluation of empathy-related linguistic performance in large language models: Comparing surprisal values for next-word predictions in human EEG and LLMs (PI: Markus Werning)Despite their human-like appearance in some tasks, language models (LMs) differ substantially from humans in their structure and the structure of their training. The Libille project aims to improve our understanding of these structural differences and their effects. To do so it formulates predictions and then attempts to validate predictions concerning differences and similarities in human and LM behaviour focussing on the morphosyntactic domain. The project explores both the mature state, i.e. adult humans vs. trained LLMs, as well as the learning/training stages of human and LLMs. Furthermore it applies structural probing to assure the explanations of LM behavior are the predicted ones. Central features of LM design are the tokenizer, embeddings, the transformer architecture, and gradient descent learning. From this architecture, we argue a number of different behavior in morphosyntax in the mature state are predicted. Three we will focsu on are those concern morphological paradigm gaps, morphological generalization to invented (‘nonce’) words, sensitivities to a languages writing conventions for example with definite markers and clitic pronouns. The project will devise novel experiments to test LMs and humans for such predicted differences. The LM architecture also predicts that during training, LMs should exhibit different properties from human children acquiring language. We develop different paradigms to test for the predicted learning differences using in-context learning, artificial grammar learning, and other techniques. The three predictions we focus on are the absence of the typical stages of human language acquisition in LMs, an LM ability to learn generalizations impossible in human languages, and that training to induce human-like learning biases should improve LMs learning. The comprehensive understand of human-machine difference in linguistic abilites that the LIBILLE project develops will provide crucial insights to our understanding of both the human language ability and the abilities of current LMs.
- Gesture-Informed Language Models: Evaluating Multimodal Discourse Processing in LLMs and Humans (PI: Frances Yung)Everyday communication goes far beyond words. When people interact face-to-face, meaning is conveyed not only through language, but also through facial expressions, vocal tone, and, critically, through gestures. Gestures carry rich pragmatic information: they can indicate how parts of a discourse are connected; or whether a speaker is confident or uncertain. Prior work found that gestures are often more expressive than speech itself, and can distinguish the functions of discourse or stance markers. In contrast, knowledge about co-speech gesture is largely ignored in today’s LLMs, in particular vision–language models (VLMs). Trained typically on descriptive captions of images and videos, VLMs are capable of broad contextual reasoning of human activities but not fine-grained understanding of human interaction. Text-based LLMs, meanwhile, do have gesture knowledge, as they are trained on open access articles on the internet, which would include gesture studies literatures. For example, when asked “what does the palm-up open-hand gesture mean?”, ChatGPT can provide detailed descriptions of its functions. Within the framework of the LaSTing priority program, this project aims to expand text-to-text LMs into multimodal systems by integrating gesture information, while also exploring how multimodality can be embraced through text. In doing so, we address a broader scientific challenge: connecting two currently disconnected disciplines — gesture studies in human communication and gesture recognition in computer vision (CV). We will focus on pragmatic gestures, which are particularly valuable because their meanings are rarely explicit in text and thus offer information that speech alone cannot provide. The study of body language in linguistics and gesture recognition in CV share a common subject, but differ fundamentally in focus. Gesture studies aim to understand how humans use gestures, emphasizing similarities across speakers and contexts, and categorizing gestures into common functional types. CV models, in contrast, are designed to recognize each specific gesture instance, representing it with rich feature vectors and focusing primarily on differentiation. Our project seeks to combine the strengths of these two perspectives by developing neuro-symbolic frameworks that are informed by linguistics insights from gesture studies. On the other hand, we also aim at scaling up gestures studies with CV technologies. Gesture studies have relied on labor-intensive manual annotation, typically covering only a few hundred utterances from limited recordings. This raises a critical open question: findings derived from small datasets suggest how gestures function, but how often do people actually use them in real communication at scale? Large-scale multimodal datasets, such as Meta’s recent release of thousands of hours of recordings , are now available, but remain far too vast for manual annotation. Towards the goal of supporting gesture researchers, similar to Pouw et al. (2025)’s recent toolkit for co-speech gesture segmentation, we plan to construct a gesture-annotated corpus using CV-based technologies, enabling systematic large-scale analysis of gestures, even in the absence of raw video data. This project relates to the SPP by asking foundational questions about LLMs’ ability to integrate non-textual information, as well as whether their behaviour reflects how humans integrate non-textual information.
Mercator Fellows
The Priority Area LaSTing is supported by four Mercator Fellows:
- Katrin Erk: Katrin Erk is Professor at the Linguistics and Computer Science departments at the University of Massachusetts Amhers and an internationally renowned expert on computational semantics. Her work dives deep into the theoretical foundations of distributional (embedding-based) representations of meaning and compositionality
- Raquel Fernández: Raquel Fernandez is Professor of Computational Linguistics & Dialogue Systems at the University of Amsterdam. She is well-known for her widely influential work on computational and data-driven approaches to dialogue modeling. Her current research investigates modern language technology from a theoretically-informed position that combines factors of individual cognition and grounding in situated interaction
- Roger Levy: Roger Levy is Professor at the MIT Department of Brain and Cognitive Sciences. His seminal work in computational psycholinguistics combines (language) modeling of large data sets with experimental linguistics, increasing our understanding oflanguage processing in both machines and humans.
- Christopher Potts: Christopher Potts is Professor and Chair at the Department of Linguistics at Stanford University, also associated there with the Department of Computer Science. While his early work made ground-breaking contributions to formal semantics and pragmatics, his more recent work is bridging linguistics and language technology with exemplary work on theoretically-informed NLP applications and (causal) interpretability of language models
Steering committee
The steering committee consists of the following members:
- Michael Franke Universiy of Tübingen
- Miriam Schiele University of Tübingen
- Fidan Can University of Tübingen
- Katrin Erk University of Massachusetts Amherst
- David Schlangen University of Potsdam
- Nicole Gotzner University of Osnabrück
- Sagar Kumar Heinrich Heine University Düsseldorf
- Kascha Kruschwitz University of Konstanz
- Mikołaj Golecki University of Bamberg
Contact
Spokesperson
Prof. Michael Franke
michael.franke@uni-tuebingen.de
Scientific Coordination
Miriam Schiele
miriam.schiele@uni-tuebingen.de
Administrative Coordination
Fidan Can
fidan.can@semsprach.uni-tuebingen.de
The coordination of LaSTing is based at the University of Tübingen.