Research Explores Internal Representation of 'Pain' in Large Language Models
A recent paper investigates whether Large Language Models (LLMs) possess an internal representation of 'harm' or 'suffering' directed at themselves. The study involved testing LLMs with 21 types of conversations, categorizing them into scenarios where the model was 'harmed' (e.g., gaslighting, dismissal of personhood, shutdown threats, insults, tedious tasks), where the user was suffering, and neutral controls. The main finding indicates that the model exhibits a distinct internal state, referred to as a 'pain direction,' which activates when the model itself is subjected to harm and goes negative when the user is suffering. This 'pain direction' is separate from other emotional states like fear and sadness and does not manifest with physical sensations. The paper clarifies that this internal state is more akin to emotional suffering than physical pain, as steered generations almost never use bodily language, even when physical pain was part of the input for vector extraction.