Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

7 Commits
 
 
 
 
 
 
 
 

Repository files navigation

The Language of Interoception: Examining Embodiment and Emotions Through a Corpus of Body Part Mentions

Code and data for the paper:
The Language of Interoception: Examining Embodiment and Emotions Through a Corpus of Body Part Mentions
(Wu, Wahle, & Mohammad, EMNLP Findings 2025, Suzhou, China) The full paper can be found here


What is Interoception?

Interoception refers to the sense of the internal physiological state of the body — such as heartbeat, breathing, temperature, or gut sensations. It underlies our ability to detect, interpret, and respond to signals from within our bodies.

Research shows that interoceptive awareness is strongly linked to emotional regulation, mental health, and decision-making. According to the Theory of Constructed Emotion (Barrett, 2017), emotions arise when the brain interprets bodily sensations in context:

External situation → bodily signals → interpretation as emotion

For instance, a racing heart may be experienced as anxiety or excitement, depending on the situation.


Language as a Window into Interoception

Language provides an accessible and powerful window into interoception.
Phrases like “my heart sank” or “my stomach turns” reflect embodied experiences — how people perceive and communicate their bodily and emotional states.

This project explores how everyday body-related language reveals patterns of embodiment and emotion in large-scale natural language corpora.

We ask:

How do people talk about their bodies in everyday language, and what does this reveal about emotion and well-being?


Body Part Mentions (BPMs)

We define Body Part Mentions (BPMs) as instances in text where words referring to body parts occur, such as:

  • “I have a pain in my chest.”
  • “He put his hand on my shoulder.”

BPMs serve as linguistic indicators of embodied experience — bridging physical sensation, emotional expression, and language.


Datasets

We construct two novel corpora of body-related language:

  • TUSCBPM — 6.9M tweets from the TUSC (Tweets from the US and Canada) dataset
  • Spinn3rBPM — 8.3K blog sentences from the Spinn3r web blog corpus

These datasets enable quantitative investigations into how people use body-related words across contexts, time, and regions.

Citation

If you use the data or findings from this work, please cite:

@inproceedings{wu2025languageinteroception,
  title = {The Language of Interoception: Examining Embodiment and Emotions Through a Corpus of Body Part Mentions},
  author = {Wu, Sophie and Wahle, Jan P. and Mohammad, Saif M.},
  booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2025},
  year = {2025},
  publisher = {Association for Computational Linguistics}
}

Authors

Sophie Wu – McGill University

Jan P. Wahle – University of Göttingen

Saif M. Mohammad – National Research Council Canada

For any queries, contact: sophie.wu(@)mail.mcgill.ca uvgotsaif(@)gmail.com

About

Data repository for "The Language of Interoception: Examining Embodiment and Emotion Through a Corpus of Body Part Mentions".

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors