基于联合嵌入方法的视频动作细粒度识别
    点此下载全文
引用本文:金扬,顾亦然.基于联合嵌入方法的视频动作细粒度识别[J].计算技术与自动化,2026,(2):75-79
摘要点击次数: 57
全文下载次数: 0
作者单位
金扬,顾亦然 (南京邮电大学自动化学院、人工智能学院,江苏 南京 210023) 
中文摘要:联合嵌入方法的核心思想是学习一个多模态共享空间,使不同模态数据在该空间中的表征既能保留各自模态特性,又能实现跨模态的语义对齐。针对视频动作细粒度识别任务,提出了一种新的联合嵌入模型JEDAM。首先,使用I3D网络提取视频的视觉特征,然后通过带有动态查询向量的多头注意力机制将视觉特征嵌入到多模态向量空间中。其次,将用于描述细粒度动作的文本特征通过残差门控机制嵌入到同一向量空间中。最后,引入自适应学习策略,模型能够根据实际表现动态调整参数更新策略。实验结果显示,JEDAM模型在视频动作细粒度识别任务上取得了更优的性能表现。
中文关键词:视频理解  视频动作识别  联合嵌入  多头注意力机制  残差门控机制  自适应学习
 
Fine-grained Video Action Recognition Based on Joint Embedding Methods
Abstract:The core idea of joint embedding methods is to learn a multimodal shared space where representations of different modalities can retain their respective characteristics while achieving cross-modal semantic alignment. This paper proposes a new joint embedding model, JEDAM, for the task of fine-grained video action recognition. First, the I3D network is used to extract visual features from the video, which are then embedded into the multimodal vector space through a multi-head attention mechanism with dynamic query vectors. Second, the textual features describing fine-grained actions are embedded into the same vector space using a residual gating mechanism. Finally, an adaptive learning strategy is introduced, enabling the model to dynamically adjust the parameter update strategy based on actual performance. Experimental results show that the JEDAM model achieves superior performance in the task of fine-grained video action recognition.
keywords:video understanding  video action recognition  joint embedding  multi-head attention mechanism  residual gating mechanism  adaptive learning
查看全文   查看/发表评论   下载pdf阅读器