| Title: |
Leveraging Concept Annotations for Trustworthy Multimodal Video Interpretation through Modality Specialization |
| Authors: |
Ancarani, Elisa; Tores, Julie; Sun, Rémy; Sassatelli, Lucile; Wu, Hui-Yin; Precioso, Frederic |
| Contributors: |
Université Côte d'Azur (UniCA); Laboratoire d'Informatique, Signaux, et Systèmes de Sophia-Antipolis (I3S) / Equipe SIGNET; COMmunications, Réseaux, systèmes Embarqués et Distribués (Laboratoire I3S - COMRED); Laboratoire d'Informatique, Signaux, et Systèmes de Sophia Antipolis (I3S); Université Nice Sophia Antipolis (1965 - 2019) (UNS)-Centre National de la Recherche Scientifique (CNRS)-Université Côte d'Azur (UniCA)-Université Nice Sophia Antipolis (1965 - 2019) (UNS)-Centre National de la Recherche Scientifique (CNRS)-Université Côte d'Azur (UniCA)-Laboratoire d'Informatique, Signaux, et Systèmes de Sophia Antipolis (I3S); Université Nice Sophia Antipolis (1965 - 2019) (UNS)-Centre National de la Recherche Scientifique (CNRS)-Université Côte d'Azur (UniCA)-Université Nice Sophia Antipolis (1965 - 2019) (UNS)-Centre National de la Recherche Scientifique (CNRS)-Université Côte d'Azur (UniCA); Modèles et algorithmes pour l’intelligence artificielle (MAASAI); Centre Inria d'Université Côte d'Azur; Institut National de Recherche en Informatique et en Automatique (Inria)-Institut National de Recherche en Informatique et en Automatique (Inria)-Université Nice Sophia Antipolis (1965 - 2019) (UNS)-Laboratoire Jean Alexandre Dieudonné (LJAD); Université Nice Sophia Antipolis (1965 - 2019) (UNS)-Centre National de la Recherche Scientifique (CNRS)-Université Côte d'Azur (UniCA)-Centre National de la Recherche Scientifique (CNRS)-Université Côte d'Azur (UniCA)-Scalable and Pervasive softwARe and Knowledge Systems (Laboratoire I3S - SPARKS); Université Nice Sophia Antipolis (1965 - 2019) (UNS)-Centre National de la Recherche Scientifique (CNRS)-Université Côte d'Azur (UniCA)-Centre National de la Recherche Scientifique (CNRS); Institut universitaire de France (IUF); Ministère de l'Education nationale, de l’Enseignement supérieur et de la Recherche (M.E.N.E.S.R.); Signal, Images et Systèmes (Laboratoire I3S - SIS); Biologically plausible Integrative mOdels of the Visual system : towards synergIstic Solutions for visually-Impaired people and artificial visiON (BIOVISION); Institut National de Recherche en Informatique et en Automatique (Inria)-Institut National de Recherche en Informatique et en Automatique (Inria); Scalable and Pervasive softwARe and Knowledge Systems (Laboratoire I3S - SPARKS); ANR-21-CE38-0012,TRACTIVE,Vers une analyse multimodale automatique de l'esthétique discursive filmique(2021); European Project: 951911,H2020-ICT-2018-20,H2020-ICT-2019-3,AI4Media(2020) |
| Source: |
MM '25:The 33rd ACM International Conference on Multimedia ; https://hal.science/hal-05393708 ; MM '25:The 33rd ACM International Conference on Multimedia, Oct 2025, Dublin Ireland, France. pp.11-19, ⟨10.1145/3746277.3760415⟩ |
| Publisher Information: |
CCSD; ACM |
| Publication Year: |
2025 |
| Collection: |
HAL Université Côte d'Azur |
| Subject Terms: |
Explanatory concept annotation; Trustworthiness; Multimodality; Video interpretation; [INFO.INFO-AI]Computer Science [cs]/Artificial Intelligence [cs.AI]; [INFO.INFO-LG]Computer Science [cs]/Machine Learning [cs.LG] |
| Subject Geographic: |
Dublin Ireland; France |
| Description: |
International audience ; Multimodal datasets usually come as multimodal data annotated for a certain construct (such as depression). However, for such tasks of video interpretation, models must not only make accurate predictions, but make them for the right reasons. Ensuring model trustworthiness is however hampered by the lack of per-modality information. We consider the case of a recently introduced dataset for the video interpretation task of detecting objectification, annotated for this end task along with multimodal explanatory concepts that provide per-modality labels. With such additional knowledge, we study how to design models with both high task accuracy and modality trustworthiness. We first introduce the MSpecF framework articulating and fusing a spectrum of variably specialized models, and two trustworthiness metrics. We show that modality-specialized models generally maximize trustworthiness, and maximize task accuracy for confident modalities. For less certain modalities, task accuracy is maximized by non-specialized models. We show that the full fusion of specialized models MSpecF(All*) achieves advantageous trade-offs between task accuracy and trustworthiness compared to other fusion choices. This work shows that rich per-modality annotations of moderate-size datasets allow to make more trustworthy models, essential for applications such as supporting social scientists in analyzing complex social constructs. |
| Document Type: |
conference object |
| Language: |
English |
| ISBN: |
979-84-00-72059-8 |
| Relation: |
info:eu-repo/semantics/altIdentifier/arxiv/2504.11232v1; info:eu-repo/grantAgreement//951911/EU/A European Excellence Centre for Media, Society and Democracy/AI4Media; ARXIV: 2504.11232v1 |
| DOI: |
10.1145/3746277.3760415 |
| Availability: |
https://hal.science/hal-05393708; https://hal.science/hal-05393708v1/document; https://hal.science/hal-05393708v1/file/3746277.3760415.pdf; https://doi.org/10.1145/3746277.3760415 |
| Rights: |
https://creativecommons.org/licenses/by/4.0/ ; info:eu-repo/semantics/OpenAccess |
| Accession Number: |
edsbas.89151689 |
| Database: |
BASE |