datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
solana-memecoin-calls
Solana memecoin calls — a public record with the misses left in
7,936 pump.fun token calls, each with the market cap we called it at, the peak it reached
afterwards, and the exact second it was posted publicly. The whole file is hashed and the hash is
anchored in a Bitcoin block, so no row can be added, edited or back-dated after the fact.
Every trading channel publishes its winners. This is the same feed with the losers still in it —
about six calls in ten never double, and… See the full description on the dataset page: https://huggingface.co/datasets/Smurfetc/solana-memecoin-calls.memerag
MEMERAG Faithfulness LLM-Judge Labels
This dataset extends MEMERAG (Cruz Blandón et al., 2025, ACL 2025,
arXiv:2502.17163) with LLM-judge faithfulness predictions from gpt-5.4, reusing
MEMERAG's own data and judge prompt templates.
Files
File
Description
memerag_llm_judge_<prompt_variant>_<model>.jsonl
LLM-judge predictions (produced by scripts/annotate_memerag.py)
scripts/
Reproduction scripts (see below)
Dataset statistics… See the full description on the dataset page: https://huggingface.co/datasets/Glide-py/memerag.MEME
MEME: Multi-Entity and Evolving Memory Evaluation
A benchmark for evaluating LLM memory systems along two orthogonal dimensions: entity scope (single vs. multi-entity) and temporal dynamics (static vs. evolving). MEME defines six tasks targeting memory-intensive operations in each quadrant, including two task types that no prior benchmark covers: Cascade (propagating updates through dependency rules) and Absence (recognizing uncertainty when a previously valid answer becomes… See the full description on the dataset page: https://huggingface.co/datasets/meme-benchmark/MEME.photo-memes-redditmemefact-templates
MemeFact Templates Dataset
This dataset contains 663 meme templates enriched with contextual knowledge for fact-checking meme generation. Each template includes comprehensive information about its origin, cultural significance, visual characteristics, and typical caption patterns to support Retrieval Augmented Generation (RAG) systems.
Dataset Description
Overview
The "MemeFact Templates" dataset is the result of extensive data engineering applied to the… See the full description on the dataset page: https://huggingface.co/datasets/sergiogpinto/memefact-templates.hateful_memes
The Hateful Memes Challenge README
The Hateful Memes Challenge is a dataset and benchmark created by Facebook AI to drive and measure progress on multimodal reasoning and understanding. The task focuses on detecting hate speech in multimodal memes.
Please see the paper for further details:
The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes
D. Kiela, H. Firooz, A. Mohan, V. Goswami, A. Singh, P. Ringshia, D. Testuggine
Dataset details
The files for… See the full description on the dataset page: https://huggingface.co/datasets/lisz1012/hateful_memes.hateful-memes
The Hateful Memes Challenge README
The Hateful Memes Challenge is a dataset and benchmark created by Facebook AI to drive and measure progress on multimodal reasoning and understanding. The task focuses on detecting hate speech in multimodal memes.
Please see the paper for further details:
The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes
D. Kiela, H. Firooz, A. Mohan, V. Goswami, A. Singh, P. Ringshia, D. Testuggine
Dataset details
The files… See the full description on the dataset page: https://huggingface.co/datasets/Zhihao-Yang/hateful-memes.memerag
MEMERAG Faithfulness LLM-Judge Labels
This dataset extends MEMERAG (Cruz Blandón et al., 2025, ACL 2025,
arXiv:2502.17163) with LLM-judge faithfulness predictions from an OpenAI model,
reusing MEMERAG's own data and judge prompt templates.
Files
File
Description
memerag_llm_judge_<prompt_variant>_<model>.jsonl
LLM-judge predictions (produced by scripts/annotate_memerag.py)
scripts/
Reproduction scripts (see below)
Dataset statistics… See the full description on the dataset page: https://huggingface.co/datasets/imerad-kv/memerag.factcheck-memes-x
Fact-checking Memes - X Dataset
This dataset contains 119 meme correction posts and their associated engagement metrics from a real-world deployment of fact-checking memes on X (formerly Twitter). The memes were specifically designed to counter misinformation by providing visually engaging explanations of fact-checking verdicts.
Dataset Description
Overview
The "Fact-checking Memes - X" dataset documents a social media experiment conducted between October 25… See the full description on the dataset page: https://huggingface.co/datasets/sergiogpinto/factcheck-memes-x.LoC-meme-generator
Dataset Card for "LoC-meme-generator"
This is an official meme dataset from the library of congress.
Meme Dataset Exploratory Data Analysis Report
courtesy of chatGPT data analysis
Basic Dataset Information
Number of Entries: 57685
Number of Columns: 10
Columns:
Meme ID
Archived URL
Base Meme Name
Meme Page URL
MD5 Hash
File Size (In Bytes)
Alternate Text
Display Name
Upper Text
Lower Text
File Size Summary
{
"count": 57685.0… See the full description on the dataset page: https://huggingface.co/datasets/pszemraj/LoC-meme-generator.Hateful_Memes_in_VLMThis dataset contains the response of VLMs (InstructBlip, ShareGPT4V, LLaVA and CogVLM) to hateful memes and the annotation to these responses. For more information, please refer to paper "From Meme to Threat: On the Hateful Meme Understanding and Induced Hateful Content Generation in Open-Source Vision Language Models."
prm-mc-value-context-memento
MC-value PRM dataset — context mode: memento
Data for training an in-context / policy-conditioned Monte-Carlo value PRM. Each row's query is a
partial reasoning prefix; the target reward = V = P(correct | prefix), the Monte-Carlo value estimated
from branched Qwen3.5-4B rollouts on Polaris math problems. The user prompt additionally carries a
"# Other attempts by the same model at this problem" block — the ablation variable.
Context for this variant: Up to K=4 OTHER attempts'… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/prm-mc-value-context-memento.memefact-llm-evaluations
MemeFact LLM Evaluations Dataset
This dataset contains 7,680 evaluation records where state-of-the-art Large Language Models (LLMs) assessed fact-checking memes according to specific quality criteria. The dataset provides comprehensive insights into how different AI models evaluate visual-textual content and how these evaluations compare to human judgments.
Dataset Description
Overview
The "MemeFact LLM Evaluations" dataset documents a systematic… See the full description on the dataset page: https://huggingface.co/datasets/sergiogpinto/memefact-llm-evaluations.semantic-memeshateful_memes_fine_grained
Hateful Memes Fine-Grained Dataset
This dataset is a fine-grained extension of the widely used Hateful Memes dataset, designed to enable more nuanced analysis of harmful multimodal content. While the original dataset focuses on binary hatefulness classification, this extension introduces additional annotation dimensions capturing incivility and intolerance at a more granular level.
The dataset consists of a subset of 2,030 memes, each annotated independently by three annotators.… See the full description on the dataset page: https://huggingface.co/datasets/nils-herrmann/hateful_memes_fine_grained.semantic-memeshateful-memes
The Hateful Memes Challenge README
The Hateful Memes Challenge is a dataset and benchmark created by Facebook AI to drive and measure progress on multimodal reasoning and understanding. The task focuses on detecting hate speech in multimodal memes.
Please see the paper for further details:
The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes
D. Kiela, H. Firooz, A. Mohan, V. Goswami, A. Singh, P. Ringshia, D. Testuggine
Dataset details
The files… See the full description on the dataset page: https://huggingface.co/datasets/j8in/hateful-memes.memecoins-chart-data-low-mc
Memecoins - Low Market Cap Chart Data
Network: Solana
Vastly labeled dataset with over 140 heavily detailed, low market cap Solana memecoin chart's data that can be used for simple or advanced analysis, and over 40,000 rows of datapoints.
Access Requirements (Paid Dataset)
This dataset is behind manual gated access.
To obtain access:
Purchase the dataset here:https://masonmarker.gumroad.com/l/solanamemecoins1
Provide your Hugging Face username at checkout.
Return to… See the full description on the dataset page: https://huggingface.co/datasets/masonmarker/memecoins-chart-data-low-mc.hateful-memes
The Hateful Memes Challenge README
The Hateful Memes Challenge is a dataset and benchmark created by Facebook AI to drive and measure progress on multimodal reasoning and understanding. The task focuses on detecting hate speech in multimodal memes.
Please see the paper for further details:
The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes
D. Kiela, H. Firooz, A. Mohan, V. Goswami, A. Singh, P. Ringshia, D. Testuggine
Dataset details
The files… See the full description on the dataset page: https://huggingface.co/datasets/rayyanm86/hateful-memes.solana-memecoin-kols
500+ Solana Memecoin KOLs
A list of 500+ Solana Memecoin KOLs with their Solana wallet addresses.
KOL wallet activity has been shown to be a leading indicator of memecoin price movements. Each trade you make without seeing KOL activity is a trade made in the dark.
Dataset includes 500+ of the most popular KOL Solana wallets, including Cupsey, Cented, Theo, West, Gake, and hundreds of others.
Dataset Preview
name
solana_wallet
Cented… See the full description on the dataset page: https://huggingface.co/datasets/masonmarker/solana-memecoin-kols.allknowingroger__Meme-7B-slerp-details
Dataset Card for Evaluation run of allknowingroger/Meme-7B-slerp
Dataset automatically created during the evaluation run of model allknowingroger/Meme-7B-slerp
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/allknowingroger__Meme-7B-slerp-details.hateful-memes
The Hateful Memes Challenge README
The Hateful Memes Challenge is a dataset and benchmark created by Facebook AI to drive and measure progress on multimodal reasoning and understanding. The task focuses on detecting hate speech in multimodal memes.
Please see the paper for further details:
The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes
D. Kiela, H. Firooz, A. Mohan, V. Goswami, A. Singh, P. Ringshia, D. Testuggine
Dataset details
The files for… See the full description on the dataset page: https://huggingface.co/datasets/zhending/hateful-memes.hateful-memes-with-captions_generalhateful-meme-captions
Dataset Card for "hateful-meme-captions"
More Information needed
hateful-memes
The Hateful Memes Challenge README
The Hateful Memes Challenge is a dataset and benchmark created by Facebook AI to drive and measure progress on multimodal reasoning and understanding. The task focuses on detecting hate speech in multimodal memes.
Please see the paper for further details:
The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes
D. Kiela, H. Firooz, A. Mohan, V. Goswami, A. Singh, P. Ringshia, D. Testuggine
Dataset details
The files for… See the full description on the dataset page: https://huggingface.co/datasets/ray7778/hateful-memes.corpus-text-meme-indonesiameme-viralityhateful-memes-with-captions_detection_label0hateful-memes-with-captions-no_shift-probshateful-memes-with-captions-mild_kl03-probs
