Back to Search Start Over

MMED: A Multi-domain and Multi-modality Event Dataset

Authors :
Zehang Lin
Wenyin Liu
Zhenguo Yang
Lingni Guo
Qing Li
Publication Year :
2019
Publisher :
arXiv, 2019.

Abstract

In this work, we release a multi-domain and multi-modality event dataset (MMED), containing 25,052 textual news articles collected from hundreds of news media sites (e.g., Yahoo News, BBC News, etc.) and 75,884 image posts shared on Flickr by thousands of social media users. The articles contributed by professional journalists and the images shared by amateur users are annotated according to 410 real-world events, covering emergencies, natural disasters, sports, ceremonies, elections, protests, military intervention, economic crises, etc. The MMED dataset is collected by the following the principles of high relevance in supporting the application needs, a wide range of event types, non-ambiguity of the event labels, imbalanced event clusters, and difficulty discriminating the event labels. The dataset can stimulate innovative research on related challenging problems, such as (weakly aligned) cross-modal retrieval and cross-domain event discovery, inspire visual relation mining and reasoning, etc. For comparisons, 15 baselines for two scenarios have been quantitatively and qualitatively evaluated using the dataset.

Details

Database :
OpenAIRE
Accession number :
edsair.doi.dedup.....88f318bf865898e7d8ced4e679e226bd
Full Text :
https://doi.org/10.48550/arxiv.1904.02354