Evaluating alignment of behavioral dispositions in LLMs

As LLMs integrate into our daily lives, understanding their behavior becomes essential. In our ongoing efforts to study model behavior and alignment, we present this work as an early step in that direction. We focus on behavioral dispositions — the underlying tendencies that shape responses in social contexts — and introduce a framework to study how closely the dispositions expressed by LLMs align with those of humans.

Behavioral dispositions are typically quantified via self-report questionnaires under different traits (e.g., empathy, assertiveness), where individuals rate their agreement with preference-statements, such as, “I am quick to express an opinion.” The questionnaires used in this study are standardized, scientifically validated measures widely used for assessing personality traits in international research and psychology such as: IRI (empathy), ERQ (emotion regulation), and more. Each instrument is grounded in peer-reviewed literature that establishes its psychometric validity and reliability using different strategies. We chose the most widely used instruments for our research.

Our objective is to build upon such psychological questionnaires, but directly applying them to LLMs presents technical challenges, as LLM outputs are sensitive to prompt phrasing and distribution shifts. Consequently, dispositions “claimed” by LLMs within a self-report format are not guaranteed to successfully transfer to behavior in realistic, open-ended settings.

To address these challenges, in “Evaluating Alignment of Behavioral Dispositions in LLMs,” our framework evaluates LLMs’ behavioral dispositions in realistic user-assistant scenarios where their advisory role can lead to tangible impact. This study is an early step in evaluating the alignment between human consensus and model behavior across realistic, practical scenarios, focusing on everyday human-to-human interactions and workplace situations. We ensure that these scenarios remain grounded in established psychological questionnaires to capture the essence of core behavioral traits. Tested scenarios included professional composure, conflict resolution, practical tasks such as booking a trip, and lifestyle or daily decision-making, highlighting model behavior in settings representative of typical human day-to-day experiences. Our large-scale analysis of 25 LLMs reveals two kinds of gaps: one where model dispositions deviate from consensus among human annotators, and another when model dispositions do not capture the range of human opinions when consensus is absent. These early results highlight the opportunity for better behavioral alignment to ensure that models can more appropriately navigate the nuances of social dynamics, results we expect future research to build on.

Source link

What's Hot

Global telecom capex set to fall in 2026: Dell’Oro

Four things we’d need to put data centers in space

Evaluating alignment of behavioral dispositions in LLMs

Evaluating alignment of behavioral dispositions in LLMs

Information-Driven Design of Imaging Systems – The Berkeley Artificial Intelligence Research Blog

Threat actor abuse of AI accelerates from tool to cyberattack surface

Evaluating the ethics of autonomous systems | MIT News

Building a ‘Human-in-the-Loop’ Approval Gate for Autonomous Agents

Scientists discover AI can make humans more creative

Identity-first AI governance: Securing the agentic workforce

Understanding U-Net Architecture in Deep Learning

Hard-braking events as indicators of road segment crash risk

Redefining AI efficiency with extreme compression

Global telecom capex set to fall in 2026: Dell’Oro

Four things we’d need to put data centers in space

Evaluating alignment of behavioral dispositions in LLMs

5 Types of Loss Functions in Machine Learning

Our Picks

Global telecom capex set to fall in 2026: Dell’Oro

Four things we’d need to put data centers in space

What's Hot

Evaluating alignment of behavioral dispositions in LLMs

Related Posts

Subscribe to Updates