TY - GEN
T1 - Food Image Recognition
T2 - 13th International Conference on E-Health and Bioengineering, EHB 2025
AU - Constantin, Onisim
AU - Tapu, Ruxandra
AU - Mocanu, Bogdan
AU - Grosu, Mirela
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2026.
PY - 2026/1/1
Y1 - 2026/1/1
N2 - Food recognition is a challenging fine-grained classification task with practical applications in health monitoring, dietary assessment, and intelligent food services. Although convolutional neural networks, transformer-based vision models, and multimodal approaches have advanced rapidly, their comparative strengths for food recognition remain insufficiently explored. To bridge this gap, in this paper we present a unified experimental framework that enables robust evaluation and yields new insights into the design and deployment of effective food recognition systems. Results reveal that transformer architectures consistently outperform convolutional baselines, achieving over 94% accuracy, while multimodal frameworks demonstrate competitive zero-shot performance without task-specific fine-tuning. Beyond raw performance, our analysis uncovers systematic error patterns linked to high intra-class variability and visual similarity across food categories. We further introduce a lightweight web application for real time inference and explainable predictions, highlighting the practical implications of our findings.
AB - Food recognition is a challenging fine-grained classification task with practical applications in health monitoring, dietary assessment, and intelligent food services. Although convolutional neural networks, transformer-based vision models, and multimodal approaches have advanced rapidly, their comparative strengths for food recognition remain insufficiently explored. To bridge this gap, in this paper we present a unified experimental framework that enables robust evaluation and yields new insights into the design and deployment of effective food recognition systems. Results reveal that transformer architectures consistently outperform convolutional baselines, achieving over 94% accuracy, while multimodal frameworks demonstrate competitive zero-shot performance without task-specific fine-tuning. Beyond raw performance, our analysis uncovers systematic error patterns linked to high intra-class variability and visual similarity across food categories. We further introduce a lightweight web application for real time inference and explainable predictions, highlighting the practical implications of our findings.
KW - Convolutional Neural Networks
KW - Fine-grained image classification
KW - Food recognition
KW - Multimodal learning
KW - Vision Transformers
UR - https://www.scopus.com/pages/publications/105039331738
U2 - 10.1007/978-3-032-23952-5_32
DO - 10.1007/978-3-032-23952-5_32
M3 - Conference contribution
AN - SCOPUS:105039331738
SN - 9783032239518
T3 - IFMBE Proceedings
SP - 290
EP - 297
BT - Advances in Digital Health and Medical Bioengineering II - Telemedicine, Biomaterials, Environmental Protection, Medical Imaging, and Biomechanics
A2 - Costin, Hariton-Nicolae
A2 - Magjarevic, Ratko
A2 - Petroiu, Gabriela-Gladiola
PB - Springer Science and Business Media Deutschland GmbH
Y2 - 13 November 2025 through 14 November 2025
ER -