Passer à la navigation principale Passer à la recherche Passer au contenu principal

DeViL: Decoding Vision features into Language

Résultats de recherche: Le chapitre dans un livre, un rapport, une anthologie ou une collectionContribution à une conférenceRevue par des pairs

Résumé

Post-hoc explanation methods have often been criticised for abstracting away the decision-making process of deep neural networks. In this work, we would like to provide natural language descriptions for what different layers of a vision backbone have learned. Our DeViL method generates textual descriptions of visual features at different layers of the network as well as highlights the attribution locations of learned concepts. We train a transformer network to translate individual image features of any vision layer into a prompt that a separate off-the-shelf language model decodes into natural language. By employing dropout both per-layer and per-spatial-location, our model can generalize training on image-text pairs to generate localized explanations. As it uses a pre-trained language model, our approach is fast to train and can be applied to any vision backbone. Moreover, DeViL can create open-vocabulary attribution maps corresponding to words or phrases even outside the training scope of the vision model. We demonstrate that DeViL generates textual descriptions relevant to the image content on CC3M, surpassing previous lightweight captioning models and attribution maps, uncovering the learned concepts of the vision backbone. Further, we analyze fine-grained descriptions of layers as well as specific spatial locations and show that DeViL outperforms the current state-of-the-art on the neuron-wise descriptions of the MILANNOTATIONS dataset.

langue originaleAnglais
titrePattern Recognition - 45th DAGM German Conference, DAGM GCPR 2023, Proceedings
rédacteurs en chefUllrich Köthe, Carsten Rother
EditeurSpringer Science and Business Media Deutschland GmbH
Pages363-377
Nombre de pages15
ISBN (imprimé)9783031546044
Les DOIs
étatPublié - 1 janv. 2024
Modification externeOui
Evénement45th Annual Conference of the German Association for Pattern Recognition, DAGM-GCPR 2023 - Heidelberg, Allemagne
Durée: 19 sept. 202322 sept. 2023

Série de publications

NomLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Volume14264 LNCS
ISSN (imprimé)0302-9743
ISSN (Electronique)1611-3349

Une conférence

Une conférence45th Annual Conference of the German Association for Pattern Recognition, DAGM-GCPR 2023
Pays/TerritoireAllemagne
La villeHeidelberg
période19/09/2322/09/23

Empreinte digitale

Examiner les sujets de recherche de « DeViL: Decoding Vision features into Language ». Ensemble, ils forment une empreinte digitale unique.

Contient cette citation