datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
marveldataset
UltraData-Math
🤗 Dataset | 💻 Source Code | 🇨🇳 中文 README
UltraData-Math is a large-scale, high-quality mathematical pre-training dataset totaling 290B+ tokens across three progressive tiers—L1 (170.5B tokens web corpus), L2 (33.7B tokens quality-selected), and L3 (88B tokens multi-format refined)—designed to systematically enhance mathematical reasoning in LLMs. It has been applied to the mathematical pre-training of the MiniCPM Series models.
It was introduced in the… See the full description on the dataset page: https://huggingface.co/datasets/Saras583/marveldataset.edith-marvel-corpus
EDITH Marvel corpus
202,171 cleaned records - characters, teams, locations, items, events, comics -
derived from Marvel Database (marvel.fandom.com) and English Wikipedia, crawled
2026-09-04 via the MediaWiki API.
Used by https://github.com/Tomionkkas/edith for two things: training the
EDITH model, and building the BM25 retrieval index the terminal answers from.
The index is not distributed - it is a pickle, unpickling executes arbitrary
code, and it rebuilds from this corpus in… See the full description on the dataset page: https://huggingface.co/datasets/Tomionkkas/edith-marvel-corpus.marvel-domain-dataset
Marvel Evreni Soru-Cevap Veri Seti
Veri Seti Özeti
Bu veri seti, büyük dil modellerinin Marvel evreni, karakterlerinin kökenleri ve çizgi roman tarihi hakkında eğitilmesi (Fine-Tuning / LoRA) amacıyla oluşturulmuştur. Veri seti tamamen Türkçe olup instruction, input ve output formatına uygun olarak yapılandırılmıştır.
Oluşturulma Amacı
Akademik bir bilgisayar mühendisliği projesi kapsamında; web scraping yeteneklerini sergilemek, veri çoğaltma… See the full description on the dataset page: https://huggingface.co/datasets/Egertekin/marvel-domain-dataset.FRIDAY-from-Marvel-Conversations
FRIDAY-from-Marvel-Conversations
A conversational assistant dataset inspired by Marvel's FRIDAY AI, designed for fine-tuning LLMs to produce respectful, “Sir”-prefixed responses.
The dataset follows a ChatML structure, making it compatible with most modern conversational models.
Note: Some responses may contain minor grammatical errors and include references to being fine-tuned on Mistral.
Dataset Details
Author: git-prakhar
License: CC0 1.0 (Public Domain)… See the full description on the dataset page: https://huggingface.co/datasets/git-prakhar/FRIDAY-from-Marvel-Conversations.alpaca-turkmen
Turkmen Alpaca Dataset
Overview
This dataset is a Turkmen translation of the original Alpaca dataset. The Alpaca dataset is a publicly available instruction-following dataset containing approximately 52,000 instruction-following samples. This Turkmen version aims to extend the accessibility of instruction-following datasets to the Turkmen language community.
Dataset Details
Original Dataset: Alpaca
Languages: English and Turkmen
Number of Samples:… See the full description on the dataset page: https://huggingface.co/datasets/MarvelTonyStark/alpaca-turkmen.
