Professor Ryu Joo-seok, Senior Researcher Kim Jeong-min.
A study has found that the risk of aspiration can be screened using just a 2-second “ah~” sound made immediately after swallowing food. The technology is expected to have potential as a simple method to check for swallowing disorders, which are common among older adults and patients with neurological diseases, in everyday settings.
The research team led by Professor Ryu Joo-seok of the Department of Rehabilitation Medicine at Seoul National University Bundang Hospital announced that it has developed a model that distinguishes an aspiration risk group from a normal group by training artificial intelligence (AI) on voice recordings made immediately after swallowing food. The study results were published in the international journal “Scientific Reports.”
Aspiration is a phenomenon in which food or saliva enters the airway instead of the esophagus. It is common in patients whose swallowing function has declined due to aging or neurological diseases such as stroke. When repeated, bacteria in food or saliva can reach the lungs and lead to aspiration pneumonia.
The representative test currently used to determine whether aspiration has occurred is the “videofluoroscopic swallowing study” (VFSS). Using X-rays, it observes in real time how food moves from the mouth to the esophagus. Although it enables accurate evaluation, it requires equipment and specialized personnel, and involves radiation exposure, which limits its use as a test that can be performed repeatedly in daily life.
The research team focused on the fact that the voice changes after swallowing food. When swallowing occurs normally, the larynx blocks the airway so that food passes into the esophagus. However, when swallowing function is impaired, food can remain around the vocal cords or enter the airway, and this can alter the vibration of the vocal cords and the characteristics of the voice.
From October 2021 to February 2023, the research team collected voice recordings from patients who underwent swallowing disorder tests and from individuals without symptoms. Among 198 participants aged 40 or older who met the study criteria, including recording quality, 70 were in the aspiration risk group and 128 were in the normal group. In particular, taking into account the actual phonation duration achievable by patients, all voice data were standardized into 2-second segments. The average phonation time for individuals who underwent swallowing disorder tests was 2.22–2.48 seconds, shorter than the 6.19 seconds observed in the general population.
The AI model trained on combined male and female data achieved an AUC of 0.8090 in distinguishing the aspiration risk group. An AUC closer to 1 indicates superior classification performance. The highest sensitivity of the integrated model was 82.77%. Overall performance was also better than that of models trained separately for men and women. The research team analyzed that when aspiration occurs, voice characteristics commonly associated with aspiration become more prominent than the usual differences between male and female voices.
However, this study was conducted on 198 individuals at a single medical institution. To use it as a diagnostic test for aspiration in actual clinical practice, additional validation involving more patients and multiple medical institutions is required.
Professor Ryu stated, “Because aspiration occurs in the process of swallowing food or saliva, it must be possible to continuously monitor the risk in daily life,” adding, “We will secure multicenter data to verify the performance of the AI model and advance it to a level where it can be applied in clinical settings.”
ⓒ dongA.com. All rights reserved. Reproduction, redistribution, or use for AI training prohibited.
Popular News