Katalog Plus
Bibliothek der Frankfurt UAS
Bald neuer Katalog: sichern Sie sich schon vorab Ihre persönlichen Merklisten im Nutzerkonto: Anleitung.
Dieses Ergebnis aus BASE kann Gästen nicht angezeigt werden.  Login für vollen Zugriff.

Leveraging Concept Annotations for Trustworthy Multimodal Video Interpretation through Modality Specialization

Title: Leveraging Concept Annotations for Trustworthy Multimodal Video Interpretation through Modality Specialization
Authors: Ancarani, Elisa; Tores, Julie; Sun, Rémy; Sassatelli, Lucile; Wu, Hui-Yin; Precioso, Frederic
Contributors: Université Côte d'Azur (UniCA); Laboratoire d'Informatique, Signaux, et Systèmes de Sophia-Antipolis (I3S) / Equipe SIGNET; COMmunications, Réseaux, systèmes Embarqués et Distribués (Laboratoire I3S - COMRED); Laboratoire d'Informatique, Signaux, et Systèmes de Sophia Antipolis (I3S); Université Nice Sophia Antipolis (1965 - 2019) (UNS)-Centre National de la Recherche Scientifique (CNRS)-Université Côte d'Azur (UniCA)-Université Nice Sophia Antipolis (1965 - 2019) (UNS)-Centre National de la Recherche Scientifique (CNRS)-Université Côte d'Azur (UniCA)-Laboratoire d'Informatique, Signaux, et Systèmes de Sophia Antipolis (I3S); Université Nice Sophia Antipolis (1965 - 2019) (UNS)-Centre National de la Recherche Scientifique (CNRS)-Université Côte d'Azur (UniCA)-Université Nice Sophia Antipolis (1965 - 2019) (UNS)-Centre National de la Recherche Scientifique (CNRS)-Université Côte d'Azur (UniCA); Modèles et algorithmes pour l’intelligence artificielle (MAASAI); Centre Inria d'Université Côte d'Azur; Institut National de Recherche en Informatique et en Automatique (Inria)-Institut National de Recherche en Informatique et en Automatique (Inria)-Université Nice Sophia Antipolis (1965 - 2019) (UNS)-Laboratoire Jean Alexandre Dieudonné (LJAD); Université Nice Sophia Antipolis (1965 - 2019) (UNS)-Centre National de la Recherche Scientifique (CNRS)-Université Côte d'Azur (UniCA)-Centre National de la Recherche Scientifique (CNRS)-Université Côte d'Azur (UniCA)-Scalable and Pervasive softwARe and Knowledge Systems (Laboratoire I3S - SPARKS); Université Nice Sophia Antipolis (1965 - 2019) (UNS)-Centre National de la Recherche Scientifique (CNRS)-Université Côte d'Azur (UniCA)-Centre National de la Recherche Scientifique (CNRS); Institut universitaire de France (IUF); Ministère de l'Education nationale, de l’Enseignement supérieur et de la Recherche (M.E.N.E.S.R.); Signal, Images et Systèmes (Laboratoire I3S - SIS); Biologically plausible Integrative mOdels of the Visual system : towards synergIstic Solutions for visually-Impaired people and artificial visiON (BIOVISION); Institut National de Recherche en Informatique et en Automatique (Inria)-Institut National de Recherche en Informatique et en Automatique (Inria); Scalable and Pervasive softwARe and Knowledge Systems (Laboratoire I3S - SPARKS); ANR-21-CE38-0012,TRACTIVE,Vers une analyse multimodale automatique de l'esthétique discursive filmique(2021); European Project: 951911,H2020-ICT-2018-20,H2020-ICT-2019-3,AI4Media(2020)
Source: MM '25:The 33rd ACM International Conference on Multimedia ; https://hal.science/hal-05393708 ; MM '25:The 33rd ACM International Conference on Multimedia, Oct 2025, Dublin Ireland, France. pp.11-19, ⟨10.1145/3746277.3760415⟩
Publisher Information: CCSD; ACM
Publication Year: 2025
Collection: HAL Université Côte d'Azur
Subject Terms: Explanatory concept annotation; Trustworthiness; Multimodality; Video interpretation; [INFO.INFO-AI]Computer Science [cs]/Artificial Intelligence [cs.AI]; [INFO.INFO-LG]Computer Science [cs]/Machine Learning [cs.LG]
Subject Geographic: Dublin Ireland; France
Description: International audience ; Multimodal datasets usually come as multimodal data annotated for a certain construct (such as depression). However, for such tasks of video interpretation, models must not only make accurate predictions, but make them for the right reasons. Ensuring model trustworthiness is however hampered by the lack of per-modality information. We consider the case of a recently introduced dataset for the video interpretation task of detecting objectification, annotated for this end task along with multimodal explanatory concepts that provide per-modality labels. With such additional knowledge, we study how to design models with both high task accuracy and modality trustworthiness. We first introduce the MSpecF framework articulating and fusing a spectrum of variably specialized models, and two trustworthiness metrics. We show that modality-specialized models generally maximize trustworthiness, and maximize task accuracy for confident modalities. For less certain modalities, task accuracy is maximized by non-specialized models. We show that the full fusion of specialized models MSpecF(All*) achieves advantageous trade-offs between task accuracy and trustworthiness compared to other fusion choices. This work shows that rich per-modality annotations of moderate-size datasets allow to make more trustworthy models, essential for applications such as supporting social scientists in analyzing complex social constructs.
Document Type: conference object
Language: English
ISBN: 979-84-00-72059-8
Relation: info:eu-repo/semantics/altIdentifier/arxiv/2504.11232v1; info:eu-repo/grantAgreement//951911/EU/A European Excellence Centre for Media, Society and Democracy/AI4Media; ARXIV: 2504.11232v1
DOI: 10.1145/3746277.3760415
Availability: https://hal.science/hal-05393708; https://hal.science/hal-05393708v1/document; https://hal.science/hal-05393708v1/file/3746277.3760415.pdf; https://doi.org/10.1145/3746277.3760415
Rights: https://creativecommons.org/licenses/by/4.0/ ; info:eu-repo/semantics/OpenAccess
Accession Number: edsbas.89151689
Database: BASE