CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Saras583 /marveldataset UltraData-Math 🤗 Dataset | 💻 Source Code | 🇨🇳 中文 README UltraData-Math is a large-scale, high-quality mathematical pre-training dataset totaling 290B+ tokens across three progressive tiers—L1 (170.5B tokens web corpus), L2 (33.7B tokens quality-selected), and L3 (88B tokens multi-format refined)—designed to systematically enhance mathematical reasoning in LLMs. It has been applied to the mathematical pre-training of the MiniCPM Series models. It was introduced in the… See the full description on the dataset page: https://huggingface.co/datasets/Saras583/marveldataset.texttext-generation100M<n<1B0 likes1.8k downloads7mo agoHugging Face02Tomionkkas /edith-marvel-corpus EDITH Marvel corpus 202,171 cleaned records - characters, teams, locations, items, events, comics - derived from Marvel Database (marvel.fandom.com) and English Wikipedia, crawled 2026-09-04 via the MediaWiki API. Used by https://github.com/Tomionkkas/edith for two things: training the EDITH model, and building the BM25 retrieval index the terminal answers from. The index is not distributed - it is a pickle, unpickling executes arbitrary code, and it rebuilds from this corpus in… See the full description on the dataset page: https://huggingface.co/datasets/Tomionkkas/edith-marvel-corpus.texttext-generation1M<n<10M0 likes54 downloads18d agoHugging Face03Egertekin /marvel-domain-dataset Marvel Evreni Soru-Cevap Veri Seti Veri Seti Özeti Bu veri seti, büyük dil modellerinin Marvel evreni, karakterlerinin kökenleri ve çizgi roman tarihi hakkında eğitilmesi (Fine-Tuning / LoRA) amacıyla oluşturulmuştur. Veri seti tamamen Türkçe olup instruction, input ve output formatına uygun olarak yapılandırılmıştır. Oluşturulma Amacı Akademik bir bilgisayar mühendisliği projesi kapsamında; web scraping yeteneklerini sergilemek, veri çoğaltma… See the full description on the dataset page: https://huggingface.co/datasets/Egertekin/marvel-domain-dataset.textquestion-answeringn<1K2 likes18 downloads2mo agoHugging Face04git-prakhar /FRIDAY-from-Marvel-Conversations FRIDAY-from-Marvel-Conversations A conversational assistant dataset inspired by Marvel's FRIDAY AI, designed for fine-tuning LLMs to produce respectful, “Sir”-prefixed responses. The dataset follows a ChatML structure, making it compatible with most modern conversational models. Note: Some responses may contain minor grammatical errors and include references to being fine-tuned on Mistral. Dataset Details Author: git-prakhar License: CC0 1.0 (Public Domain)… See the full description on the dataset page: https://huggingface.co/datasets/git-prakhar/FRIDAY-from-Marvel-Conversations.texttext-generation1K<n<10K1 likes16 downloads1y agoHugging Face05MarvelTonyStark /alpaca-turkmen Turkmen Alpaca Dataset Overview This dataset is a Turkmen translation of the original Alpaca dataset. The Alpaca dataset is a publicly available instruction-following dataset containing approximately 52,000 instruction-following samples. This Turkmen version aims to extend the accessibility of instruction-following datasets to the Turkmen language community. Dataset Details Original Dataset: Alpaca Languages: English and Turkmen Number of Samples:… See the full description on the dataset page: https://huggingface.co/datasets/MarvelTonyStark/alpaca-turkmen.texttext-generation10K<n<100K0 likes8 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.