Better cannabis tracking starts with more transparent health data

Natural Language Processing to Identify Substance Use in Electronic Health Records: A Scoping Review.

Substance use & addiction journal β€’ β€’ Review β€’ Moderately Relevant
πŸ€–

AI Summary

This scoping review examined how natural language processing (NLP) is used to identify substance use in electronic health records (EHRs). Across 86 studies, researchers used rule-based systems, conventional machine learning, deep learning, and large language or transformer-based models to detect tobacco, alcohol, opioids, cannabinoids, stimulants, and polysubstance use. Most studies reported performance metrics above 0.80, although the abstract does not provide a single overall accuracy estimate.

Cannabinoid use appeared in only 9 studies, compared with much greater attention to tobacco, opioids, and alcohol. This means NLP tools may be useful for recognizing cannabis-related information in clinical records, but the evidence base is comparatively limited. The review also found limited transparency and reproducibility: only 22 studies published their code, and few shared detailed model specifications or datasets. Better reporting and routine sharing of research materials could improve cannabis surveillance and future clinical research.

πŸ’‘ Key Findings

1
Across 86 studies, NLP methods generally reported performance metrics above 0.80 for identifying substance use in EHRs.
High
85%
2
Cannabinoid use was identified in only 9 studies, showing that cannabis-related NLP research is markedly underrepresented compared with tobacco, alcohol, and opioids.
High
85%
3
Only 22 studies published their code, and few provided detailed model specifications, highlighting limited transparency and reproducibility.
High
90%
4
More consistent sharing of code, datasets, and model specifications could strengthen cannabis-related health surveillance and research.
High
80%

πŸ“„ Original Abstract

This scoping review aimed to characterize natural language processing (NLP) techniques deployed for identifying substance use in electronic health records (EHRs) and to compare the performance of these techniques by substance type. We conducted a systematic search of PubMed, the Cochrane Library, Embase, Web of Science, ACM Digital Library, IEEE Xplore, and Scopus for peer-reviewed original research published in English. Studies were eligible if they applied NLP to identify non-prescription or problematic substance use in EHRs, provided full-text access, and reported quantitative performance metrics. A total of 86 studies met the inclusion criteria. NLP-identified substance use included tobacco (n = 42), alcohol (n = 26), non-prescription opioids (n = 30), cannabinoids (n = 9), stimulants (n = 6), and polysubstance use and/or other drugs (n = 16). NLP techniques included rule-based (n = 46), conventional machine learning (n = 39), deep learning (n = 17), and large language/transformer-based models (n = 22), with some studies applying multiple techniques. Annotation guidelines were available for 30 studies, and only 22 published their codes. Most studies reported performance metrics exceeding 0.80. Tobacco, opioids, and alcohol were the most frequently identified substances, whereas stimulants and cannabinoids were markedly underrepresented. Across the reviewed literature, transparency and reproducibility were limited, with few studies publishing code or detailed model specifications. Nevertheless, reported performance metrics were generally high. NLP techniques showed high performance in identifying tobacco, opioids, and alcohol use in EHRs, but stimulants and cannabinoids remain underrepresented. Transparency and reproducibility remain limited, underscoring the need for routine sharing of code, datasets, and model specifications.

Explore More Research

Stay informed about the latest cannabis science.

Your stash, decoded.