whale
Datasets
All datasets matching “whale”Interaction2Code
Interaction2Code: Benchmarking MLLM-based Interactive Webpage Code Generation from Interactive Prototyping
Project Page | Paper | GitHub
Interaction2Code is the first systematic investigation and benchmark for Multimodal Large Language Models (MLLMs) in generating interactive webpages. While existing benchmarks focus on static UI-to-code tasks, Interaction2Code encompasses 127 unique webpages and 374 distinct interactions across 15 webpage types and 31 interaction categories to… See the full description on the dataset page: https://huggingface.co/datasets/whale99/Interaction2Code.DesignBenchResultsDesignBenchComUIBenchgrokipedia-v0.1-dump
Grokipedia v0.1 Scrape
This dataset represents a strctured, nearly-full point-in-time scrape of Grokipedia v0.1 as of the end of October / beginning of November 2025.
It also includes embeddings of 250-token semi-overlapping chunks of the Grokipedia corpus.
It was collected and initially used for Harold Triedman and Alexios Mantzarlis' November 2025 paper: "What did Elon Change? A comprehensive analysis of Grokipedia" (arxiv).
If you use this dataset, please cite it as follows… See the full description on the dataset page: https://huggingface.co/datasets/Whalemini/grokipedia-v0.1-dump.M2RAG
Data statices of M2RAG
Click the links below to view our paper and Github project.
If you find this work useful, please cite our paper and give us a shining star 🌟 in Github
@misc{liu2025benchmarkingretrievalaugmentedgenerationmultimodal,
title={Benchmarking Retrieval-Augmented Generation in Multi-Modal Contexts},
author={Zhenghao Liu and Xingsheng Zhu and Tianshuo Zhou and Xinyi Zhang and Xiaoyuan Yi and Yukun Yan and Yu Gu and Ge Yu and Maosong Sun}… See the full description on the dataset page: https://huggingface.co/datasets/whalezzz/M2RAG.
