Home / News / Details
ZJUI Team Led by Associate Professor Liu Zuozhu Publishes in Nature Communications
Date:14/09/2026 Article:Wang Chuxi Photo:From interviewee
Recently, a research team led by Associate Professor Liu Zuozhu at the Zhejiang University–University of Illinois Urbana-Champaign Institute (ZJUI) reported significant advances in computer-enabled dental diagnosis and clinical practice and introduced DentVLM, the first multimodal vision-language model designed for clinical dentistry. The study, A Multimodal Vision-Language Model for Comprehensive Dental Diagnosis and Enhanced Clinical Practice, has been published in Nature Communications.
 
Meng Zijie, a 2024 PhD student in Electronic Information at Zhejiang University, and Dai Xiwei, a 2025 master's student in Electronic Information at Zhejiang University, are among the co-first authors. The co-corresponding authors are Dr. Hao Jin of Shanghai Ninth People's Hospital, Shanghai Jiao Tong University School of Medicine, Professor Jimeng Sun of the University of Illinois Urbana-Champaign, Professor Wu Jian of the College of Computer Science and Technology at Zhejiang University, and Associate Professor Liu Zuozhu of ZJUI.
 
Oral diseases are a major global public health challenge, affecting around 3.5 billion people worldwide. The gap between growing demand for dental care and available healthcare resources is particularly pronounced in low- and middle-income countries. Dental diagnosis often requires clinicians to integrate information from multiple imaging sources, while conventional image interpretation relies heavily on specialist expertise, making it difficult to maintain both efficiency and accuracy in primary care and high-volume clinical settings. Existing intelligent diagnostic systems are also largely designed for a single imaging modality or a specific disease, limiting their ability to meet the multimodal, multitask, and interactive requirements of real-world dental care.

 

“”

 

To address these challenges, the team developed DentVLM, a multimodal vision-language model for comprehensive dental diagnosis. The model can jointly process seven types of two-dimensional dental images, including panoramic radiographs, lateral cephalograms, and five types of intraoral photographs, and supports 36 clinical diagnostic tasks covering dental diseases, prior treatment recognition, and malocclusion assessment. In addition to diagnostic results, DentVLM can provide supporting clinical rationales in natural language together with relevant anatomical localization, extending conventional point-solution recognition tools into a more interactive and clinically reviewable decision-support system.
 
To strengthen the model's understanding of dental knowledge and clinical reasoning, the team constructed approximately 2.46 million bilingual Chinese-English visual question-answering pairs using more than 100,000 dental images from over 20,000 patients. A two-stage training strategy was then adopted to progressively improve DentVLM's accuracy, interpretability, and clinical usability in complex dental diagnostic tasks.
 
“”

 

In comparisons with 18 general-purpose, open-source, medical, and dental-specific multimodal models, DentVLM outperformed leading baselines by an average of 23.8%. The model demonstrated strong overall performance in general dental disease and treatment recognition, malocclusion assessment, and lesion localization, highlighting its ability to understand and integrate multimodal dental information.
 
The team further conducted a multicenter clinical reader study involving 32 participants with different levels of experience. DentVLM demonstrated overall diagnostic performance comparable to that of intermediate-level dentists. With DentVLM support, junior participants improved their overall diagnostic accuracy by 10.5 percentage points, while diagnostic time was reduced by approximately 15.0% to 37.0% across participants with different levels of experience. These findings suggest that DentVLM can not only perform a wide range of dental imaging tasks independently, but also serve as an intelligent clinical support tool to improve diagnostic efficiency and assist less experienced clinicians in decision-making.
 
The study advances intelligent dental care beyond the conventional single-image, single-task paradigm toward multimodal, multitask, and explainable clinical decision support. DentVLM provides a new technological pathway for applications including at-home oral health screening, intelligent diagnosis in hospitals, and dental care in primary healthcare settings.The technology has already been deployed at several medical institutions across China. The team is now accelerating the development of a second-generation multimodal foundation model toward general-purpose intelligence in dentistry, with the aim of supporting a wider range of data modalities, more complex clinical tasks, and more natural interaction.
 
The paper was published in Nature Communications, vol. 17, Art. no. 8933, July 2026, doi: 10.1038/s41467-026-75718-x.
Paper: https://www.nature.com/articles/s41467-026-75718-x

回到顶部