datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BHM-Bengali-Hateful-Memes
Dataset Description
BHM is a novel multimodal dataset for Bengali Hateful Memes detection. The dataset consists of 7,148 memes with Bengali as well as code-mixed captions,
tailored for two tasks: (i) detecting hateful memes and (ii) detecting the social entities they target (i.e., Individual, Organization, Community, and Society).
Paper Information
Paper: https://aclanthology.org/2024.acl-long.454/
Code:… See the full description on the dataset page: https://huggingface.co/datasets/Eftekhar/BHM-Bengali-Hateful-Memes.memefact-templates
MemeFact Templates Dataset
This dataset contains 663 meme templates enriched with contextual knowledge for fact-checking meme generation. Each template includes comprehensive information about its origin, cultural significance, visual characteristics, and typical caption patterns to support Retrieval Augmented Generation (RAG) systems.
Dataset Description
Overview
The "MemeFact Templates" dataset is the result of extensive data engineering applied to the… See the full description on the dataset page: https://huggingface.co/datasets/sergiogpinto/memefact-templates.hateful_memesfactcheck-memes-x
Fact-checking Memes - X Dataset
This dataset contains 119 meme correction posts and their associated engagement metrics from a real-world deployment of fact-checking memes on X (formerly Twitter). The memes were specifically designed to counter misinformation by providing visually engaging explanations of fact-checking verdicts.
Dataset Description
Overview
The "Fact-checking Memes - X" dataset documents a social media experiment conducted between October 25… See the full description on the dataset page: https://huggingface.co/datasets/sergiogpinto/factcheck-memes-x.Hateful_Memes_in_VLMThis dataset contains the response of VLMs (InstructBlip, ShareGPT4V, LLaVA and CogVLM) to hateful memes and the annotation to these responses. For more information, please refer to paper "From Meme to Threat: On the Hateful Meme Understanding and Induced Hateful Content Generation in Open-Source Vision Language Models."
VidMsg
VidMsg
VidMsg is a manifest-only benchmark for evaluating implicit message understanding in short, internet-native videos. It contains 400 public YouTube video references paired with 52 fine-grained narrative messages across 9 thematic topics.
VidMsg evaluates whether models can infer what a video communicates, rather than only what it visibly depicts. The repository does not redistribute raw video files. It provides YouTube IDs/links, topic and message labels, metadata, and a… See the full description on the dataset page: https://huggingface.co/datasets/mememe321/VidMsg.nepali-meme-captions
NeMeme-CAP: Nepali Meme Captions
Dataset Summary
English-language captions generated by Google Gemini for the CHiPSAL 2026 SubtaskA Nepali Meme Datset.
The context-aware captions was generated accross the training, validation, and test splits.
Supported Tasks
Hateful Meme Classification: Predict whether the meme is non-hateful (label=0) and hateful (label=1).
Multimodal Meme Understanding: Useful as auxiliary text features or as ground-truth explanations for… See the full description on the dataset page: https://huggingface.co/datasets/Anish/nepali-meme-captions.memefact-llm-evaluations
MemeFact LLM Evaluations Dataset
This dataset contains 7,680 evaluation records where state-of-the-art Large Language Models (LLMs) assessed fact-checking memes according to specific quality criteria. The dataset provides comprehensive insights into how different AI models evaluate visual-textual content and how these evaluations compare to human judgments.
Dataset Description
Overview
The "MemeFact LLM Evaluations" dataset documents a systematic… See the full description on the dataset page: https://huggingface.co/datasets/sergiogpinto/memefact-llm-evaluations.Genshin-Impact-Character-Memes
market-meme-post-velocity-gamma-exposure-coherence-risk-v0.1What this repo is for
Detect squeeze risk by tracking coherence between
retail narrative acceleration
and options market structure.
Use case
watchlists for meme-style tickers
early warning for reflexive moves
separating hype from structural ignition
How to read it
coherent
post velocity rises
GEX rises or flips in a way that implies hedging demand
near-dated calls concentrate
price reaction starts to confirm
incoherent
social ramps but GEX and near-dated structure do not support it
structure… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/market-meme-post-velocity-gamma-exposure-coherence-risk-v0.1.semantic-memessemantic-memesahv_iv_mementokol-decision-datasetMemecoins-DatasetConversationSamples,question,answer
0,"hi, how are you doing?",i'm fine. how about yourself?
1,i'm fine. how about yourself?,i'm pretty good. thanks for asking.
2,i'm pretty good. thanks for asking.,no problem. so how have you been?
3,no problem. so how have you been?,i've been great. what about you?
4,i've been great. what about you?,i've been good. i'm in school right now.
5,i've been good. i'm in school right now.,what school do you go to?
6,what school do you go to?,i go to pcc.
7,i go to pcc.,do you like it… See the full description on the dataset page: https://huggingface.co/datasets/Memeathon/ConversationSamples.corpus-text-meme-indonesiameme-viralityMemecoinsmemes
