datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
marveldataset
UltraData-Math
🤗 Dataset | 💻 Source Code | 🇨🇳 中文 README
UltraData-Math is a large-scale, high-quality mathematical pre-training dataset totaling 290B+ tokens across three progressive tiers—L1 (170.5B tokens web corpus), L2 (33.7B tokens quality-selected), and L3 (88B tokens multi-format refined)—designed to systematically enhance mathematical reasoning in LLMs. It has been applied to the mathematical pre-training of the MiniCPM Series models.
It was introduced in the… See the full description on the dataset page: https://huggingface.co/datasets/Saras583/marveldataset.guided_marvels_spider_man_2_recordings_01
漫威蜘蛛侠2 raw recordings
This dataset contains raw game recordings managed by Game Data Platform. Access requests require manual approval.
Game ID: game_ae2c5af176e4e2eab106954f144c7b7f
Collection: guided (精数据)
Recordings: 155
Layout: recordings/<recording_id>/<raw component>
MARVEL
Dataset Details
Dataset Description
MARVEL is a new comprehensive benchmark dataset that evaluates multi-modal large language models' abstract reasoning abilities in six patterns across five different task configurations, revealing significant performance gaps between humans and SoTA MLLMs.
Dataset Sources [optional]
Repository: https://github.com/1171-jpg/MARVEL_AVR
Paper [optional]: https://arxiv.org/abs/2404.13591
Demo [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/kianasun/MARVEL.Marvel_network
Dataset Card for Marvel Network
This is a dataset for Marvel universe social network, which contains the relationships between Marvel heroes.
Dataset Description
The Marvel Comics character collaboration graph was originally constructed by Cesc Rosselló, Ricardo Alberich, and Joe Miro from the University of the Balearic Islands. They compare the characteristics of this universe to real-world collaboration networks, such as the Hollywood network, or the one created by… See the full description on the dataset page: https://huggingface.co/datasets/ShimizuYuki/Marvel_network.edith-marvel-corpus
EDITH Marvel corpus
202,171 cleaned records - characters, teams, locations, items, events, comics -
derived from Marvel Database (marvel.fandom.com) and English Wikipedia, crawled
2026-09-04 via the MediaWiki API.
Used by https://github.com/Tomionkkas/edith for two things: training the
EDITH model, and building the BM25 retrieval index the terminal answers from.
The index is not distributed - it is a pickle, unpickling executes arbitrary
code, and it rebuilds from this corpus in… See the full description on the dataset page: https://huggingface.co/datasets/Tomionkkas/edith-marvel-corpus.Marvelmarvel_zombies_style_lora_flux2_nf4marvelMarvel_dataset
Marvel Characters Dataset
This dataset includes a compilation of various Marvel characters, their first appearances in films or TV shows, and the actors who portrayed them.
Contents
Overview
Dataset Details
Usage
License
Overview
This dataset contains information about a wide range of Marvel characters, their debut appearances in films or TV shows, and the actors who played those roles. The data could be used for analysis, reference, or any related projects… See the full description on the dataset page: https://huggingface.co/datasets/Manvith/Marvel_dataset.marvel-domain-dataset
Marvel Evreni Soru-Cevap Veri Seti
Veri Seti Özeti
Bu veri seti, büyük dil modellerinin Marvel evreni, karakterlerinin kökenleri ve çizgi roman tarihi hakkında eğitilmesi (Fine-Tuning / LoRA) amacıyla oluşturulmuştur. Veri seti tamamen Türkçe olup instruction, input ve output formatına uygun olarak yapılandırılmıştır.
Oluşturulma Amacı
Akademik bir bilgisayar mühendisliği projesi kapsamında; web scraping yeteneklerini sergilemek, veri çoğaltma… See the full description on the dataset page: https://huggingface.co/datasets/Egertekin/marvel-domain-dataset.druziThis dataset contains 5,242,391 samples of Ukrainian news headlines.
Usage:
from datasets import load_dataset
ds = load_dataset('Yehor/ukrainian-news-headlines', split='train')
for row in ds:
print(row['headline'])
Attribution to the dataset:
Chaplynskyi, D. et al. (2021) lang-uk Ukrainian Ubercorpus [Data set]. https://lang.org.ua/uk/corpora/#anchor4
FRIDAY-from-Marvel-Conversations
FRIDAY-from-Marvel-Conversations
A conversational assistant dataset inspired by Marvel's FRIDAY AI, designed for fine-tuning LLMs to produce respectful, “Sir”-prefixed responses.
The dataset follows a ChatML structure, making it compatible with most modern conversational models.
Note: Some responses may contain minor grammatical errors and include references to being fine-tuned on Mistral.
Dataset Details
Author: git-prakhar
License: CC0 1.0 (Public Domain)… See the full description on the dataset page: https://huggingface.co/datasets/git-prakhar/FRIDAY-from-Marvel-Conversations.bookCover_sciFi_child_com_reli_marvel
Dataset Card for "bookCover_sciFi_child_com_reli_marvel"
More Information needed
Meta_marvels_simualpaca-turkmen
Turkmen Alpaca Dataset
Overview
This dataset is a Turkmen translation of the original Alpaca dataset. The Alpaca dataset is a publicly available instruction-following dataset containing approximately 52,000 instruction-following samples. This Turkmen version aims to extend the accessibility of instruction-following datasets to the Turkmen language community.
Dataset Details
Original Dataset: Alpaca
Languages: English and Turkmen
Number of Samples:… See the full description on the dataset page: https://huggingface.co/datasets/MarvelTonyStark/alpaca-turkmen.marvel-dataset1summerizemarvelnotrealatallplease1marvel-datasetmarvel_zombies_style_lora_fluxmarvel-datasetmarvellsmarvel-datasetmarvelnoteasyMarvelmarvels_datasetMarvelCharactersmarveltrial1Marvel_charactersmarvelnotreal
