JMIR Infodemiology · Published 2026-07-15 · DOI 10.2196/86593
Abstract BackgroundThe variation of language concerning tobacco products and tobacco use is known to impact the understanding of related risks and influence behaviors including use uptake and product cessation. Transnational tobacco companies can use such complexities to change the acceptability of tobacco use and influence public understanding of related risks. These changes, in turn, impact tobacco use behaviors. Looking at variations in the language used by different groups can therefore offer helpful insights into tobacco use cultures and make imbalances of information between groups plain. This paper examines the language of tobacco use, specifically smoking, across a sample of health organizations (the National Health Service, the World Health Organization, the National Institute for Health and Care Excellence, and the Centers for Disease Control and Prevention), the tobacco industry (British American Tobacco and Philip Morris International), and in “general English.” ObjectiveThrough corpus-assisted analysis, the study aims to illustrate differences in tobacco use and tobacco user characterizations by 2 transnational tobacco companies, health organizations, and users of general English. The study assesses the possible implications of these differences for public health. MethodsWe queried 4 bodies of text (corpora) from 3 different groups; 2 of these corpora were preexisting corpora of general spoken and written English in the United Kingdom and the United States. The remaining 2 sampled tobacco-related documents from 2 transnational tobacco companies and from the United Kingdom and international organizations with a focus on health between 2003 and 2023. The 2 sampled corpora contained 1355 documents and 10,023,538 words. We used the corpus analysis software LancsBox (Lancaster University) to identify variations in characterizations of tobacco users and tobacco use behaviors between these groups, using the stems “smoker*” and “smok*.” ResultsFrequency and collocation analysis showed clear differences in how the 3 groups described smokers and smoking, with only limited overlap in the terms they used. Only 7 of 23 unique categorizations of “smoker*” were shared. There was a significant association (P ConclusionsWhile there was some overlap in terminology used between corpora, the most common categorizations in each corpus were highly varied, showing very little shared language between groups in their descriptions of either tobacco users or use behaviors. This variance indicates that these groups may not share the same sense-making resources related to tobacco use, which may render information flows around tobacco vulnerable to distortion.
Abstract from DOAJ. Public domain (CC0 1.0).
Read the article at the publisher →
Fitzpatrick, I., Sun, X. (2026). Investigation of Differences in Tobacco Use Language Between Groups: Corpus-Assisted Analysis. JMIR Infodemiology. https://doi.org/10.2196/86593