datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Reasoning-Mix-100k
🧠 Reasoning Mix 100k ✨
A high-quality, balanced reasoning dataset consisting of 99,999 samples extracted from three reasoning datasets. This dataset is specifically formatted for models to utilize a /think block for step-by-step reasoning.
📊 Dataset Summary
The Reasoning Mix 100k is a curated collection of reasoning tasks, primarily focused on mathematics, logic, and general problem-solving. It combines the strengths of three high-performing datasets into a unified… See the full description on the dataset page: https://huggingface.co/datasets/PinkPixel/Reasoning-Mix-100k.WitChatManga-Encylopedia
📚 Manga Encyclopedia Dataset (ChatML) ✨
This dataset is a comprehensive collection of conversational pairs designed to train an AI model (via LoRA or Full Fine-tuning) to become an expert on manga. It covers over 48,000 unique manga titles with summaries, tags, cover URLs, and cross-manga comparisons.
🚀 Dataset Features
📖 Direct Summaries: Detailed information about individual manga titles.
💡 Smart Recommendations: Responses based on specific genre/theme… See the full description on the dataset page: https://huggingface.co/datasets/PinkPixel/Manga-Encylopedia.
