RiceNet: A Multimodal CNN–Vision Language Framework for Explainable Rice Leaf Disease Classification
Published in International Conference on Computing Advancements (ICCA), Dhaka, Bangladesh, 2026
Rice leaf diseases greatly reduce crop yields and threaten food security, especially in South Asia where rice is a key staple. Man- ual diagnosis is slow, subjective, and often unreliable in field set- tings. Although deep learning has shown promise in classifying plant diseases, most methods depend solely on convolutional neural networks (CNNs), which lack semantic understanding and often face issues with class imbalance and limited generalization. To overcome these challenges, this study introduces RiceNet, a multi- modal CNN–Vision Language Model (VLM) for explainable rice leaf disease classification. The architecture combines an EfficientNet- B4 backbone with a BLIP vision encoder via bidirectional cross- attention to jointly learn spatial and semantic features. A gated fusion mechanism adaptively merges multimodal data for accu- rate disease recognition. To boost minority-class performance, a hybrid loss function combining Focal Loss and PolyLoss is used. Ad- ditionally, techniques like Sharpness-Aware Minimization (SAM), Stochastic Weight Averaging (SWA), progressive resizing, MixUp, and CutMix are incorporated to improve optimization stability and model generalization. Evaluated on an 8-class rice leaf dis- ease dataset with 3,646 training, 770 validation, and 772 test im- ages, RiceNet achieved 96.1% accuracy, with macro-average and weighted-average F1-scores of 96.1% and 95.5%. Grad-CAM visu- alizations and BLIP-generated captions enhance interpretability visually and semantically. Overall, the combination of multimodal cross-attention learning and flat-minima optimization significantly enhances classification performance and robustness for real-world precision agriculture applications.
Recommended citation:
Download Paper
