Start Over

Using Natural Language Processing to Explore 'Dry January' Posts on Twitter: Longitudinal Infodemiology Study

Authors :: Alex M Russell
Danny Valdez
Shawn C Chiang
Ben N Montemayor
Adam E Barry
Hsien-Chang Lin
Philip M Massey
Source :: Journal of Medical Internet Research, Vol 24, Iss 11, p e40160 (2022)
Publication Year :: 2022
Publisher :: JMIR Publications, 2022.
Abstract: BackgroundDry January, a temporary alcohol abstinence campaign, encourages individuals to reflect on their relationship with alcohol by temporarily abstaining from consumption during the month of January. Though Dry January has become a global phenomenon, there has been limited investigation into Dry January participants’ experiences. One means through which to gain insights into individuals’ Dry January-related experiences is by leveraging large-scale social media data (eg, Twitter chatter) to explore and characterize public discourse concerning Dry January. ObjectiveWe sought to answer the following questions: (1) What themes are present within a corpus of tweets about Dry January, and is there consistency in the language used to discuss Dry January across multiple years of tweets (2020-2022)? (2) Do unique themes or patterns emerge in Dry January 2021 tweets after the onset of the COVID-19 pandemic? and (3) What is the association with tweet composition (ie, sentiment and human-authored vs bot-authored) and engagement with Dry January tweets? MethodsWe applied natural language processing techniques to a large sample of tweets (n=222,917) containing the term “dry january” or “dryjanuary” posted from December 15 to February 15 across three separate years of participation (2020-2022). Term frequency inverse document frequency, k-means clustering, and principal component analysis were used for data visualization to identify the optimal number of clusters per year. Once data were visualized, we ran interpretation models to afford within-year (or within-cluster) comparisons. Latent Dirichlet allocation topic modeling was used to examine content within each cluster per given year. Valence Aware Dictionary and Sentiment Reasoner sentiment analysis was used to examine affect per cluster per year. The Botometer automated account check was used to determine average bot score per cluster per year. Last, to assess user engagement with Dry January content, we took the average number of likes and retweets per cluster and ran correlations with other outcome variables of interest. ResultsWe observed several similar topics per year (eg, Dry January resources, Dry January health benefits, updates related to Dry January progress), suggesting relative consistency in Dry January content over time. Although there was overlap in themes across multiple years of tweets, unique themes related to individuals’ experiences with alcohol during the midst of the COVID-19 global pandemic were detected in the corpus of tweets from 2021. Also, tweet composition was associated with engagement, including number of likes, retweets, and quote-tweets per post. Bot-dominant clusters had fewer likes, retweets, or quote tweets compared with human-authored clusters. ConclusionsThe findings underscore the utility for using large-scale social media, such as discussions on Twitter, to study drinking reduction attempts and to monitor the ongoing dynamic needs of persons contemplating, preparing for, or actively pursuing attempts to quit or cut down on their drinking.

Subjects :: Computer applications to medicine. Medical informatics
R858-859.7
Public aspects of medicine
RA1-1270

Details

Language :: English
ISSN :: 14388871
Volume :: 24
Issue :: 11
Database :: Directory of Open Access Journals
Journal :: Journal of Medical Internet Research
Publication Type :: Academic Journal
Accession number :: edsdoj.1700e2e7fecd440aa2ba38c64fdcdaaf
Document Type :: article
Full Text :: https://doi.org/10.2196/40160

Full Text Access

View/download PDF

Tools

Email
Cite

Printer

Authors Abstract Subjects Details

Searchworks

Select search scope, currently: Articles

Catalog

books, media & more in Jio Institute collections

Articles

journal articles & other e-resources

Using Natural Language Processing to Explore 'Dry January' Posts on Twitter: Longitudinal Infodemiology Study

Abstract

Subjects

Details

Tools

Searchworks

Select search scope, currently: Articles Catalog books, media & more in Jio Institute collections Articles journal articles & other e-resources

Using Natural Language Processing to Explore 'Dry January' Posts on Twitter: Longitudinal Infodemiology Study

Abstract

Subjects

Details

Tools

Select search scope, currently: Articles

Catalog

books, media & more in Jio Institute collections

Articles

journal articles & other e-resources