CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mamakobe /luhya-multilingual-dataset Luhya Multilingual Dataset Curator: Dr. Moody AmakobeProject: Project Tafsiri — Bridging Indigenous Languages and AIVersion: 2.0License: Creative Commons Attribution 4.0 (CC BY 4.0) Dataset Overview This dataset is a comprehensive, structured multilingual corpus for the Luhya language (also written Luyia), a Bantu language cluster spoken primarily in western Kenya by the Abaluhya people — Kenya's second largest ethnic group with approximately 7… See the full description on the dataset page: https://huggingface.co/datasets/mamakobe/luhya-multilingual-dataset.texttranslation10K<n<100K4 likes47 downloads5mo agoHugging Face02mamed0v /alpaca-turkmen Turkmen Alpaca Dataset Overview This dataset is a Turkmen translation of the original Alpaca dataset. The Alpaca dataset is a publicly available instruction-following dataset containing approximately 52,000 instruction-following samples. This Turkmen version aims to extend the accessibility of instruction-following datasets to the Turkmen language community. Dataset Details Original Dataset: Alpaca Languages: English and Turkmen Number of Samples:… See the full description on the dataset page: https://huggingface.co/datasets/mamed0v/alpaca-turkmen.texttext-generation10K<n<100K0 likes40 downloads2y agoHugging Face03mamoth /cast 🏰 CASTILLO: Characterizing Response Length Distributions in Large Language Models The CASTILLO dataset is designed to support research on the variability of response lengths in large language models (LLMs). It provides statistical summaries of output lengths across 13 open-source LLMs evaluated on 7 instruction-following datasets. For each unique ⟨prompt, model⟩ pair, 10 independent responses were generated using fixed decoding parameters, and key statistics were recorded—such as… See the full description on the dataset page: https://huggingface.co/datasets/mamoth/cast.tabulartabular-classification100K<n<1M0 likes34 downloads9mo agoHugging Face04mamung /reddit_dataset_192 Bittensor Subnet 13 Reddit Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this dataset allows… See the full description on the dataset page: https://huggingface.co/datasets/mamung/reddit_dataset_192.texttext-classification10K<n<100K0 likes30 downloads2y agoHugging Face05mamung /x_dataset_192 Bittensor Subnet 13 X (Twitter) Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/mamung/x_dataset_192.texttext-classification10K<n<100K0 likes26 downloads2y agoHugging Face06mamiglia /benchmarks Small-Judge Benchmark Evaluation Dataset This dataset contains model evaluation log files (.eval format generated by Inspect AI) across various benchmark datasets and LLMs. It is used to evaluate and train model judges on model capabilities directly from sample transcripts. Expected Dataset Structure The repository organizes evaluation outputs into a deterministic, multi-level directory hierarchy: data/ ├── {benchmark}/ │ └── {task_args_hash}/ │ └──… See the full description on the dataset page: https://huggingface.co/datasets/mamiglia/benchmarks.text-generation0 likes26 downloads1mo agoHugging Face07mamed0v /TurkmenTrilingualSemi-SyntheticDictionaryDF 🏜️ Turkmen Trilingual Semi-Synthetic — Dialogue Format Language: Turkmen 🇹🇲 | English 🇬🇧 | Russian 🇷🇺Type: Instruction-style / Dialogue datasetRecords: 61 970 base recordsDialog turns (flattened): 378 941Splits: train=363 783, val=7 578, test=7 580 📘 Overview This dataset is a dialogue-style extension of the original mamed0v/TurkmenTrilingualSemi-SyntheticDictionary.It was reformatted into conversational pairs to better suit instruction-tuning, chatbot… See the full description on the dataset page: https://huggingface.co/datasets/mamed0v/TurkmenTrilingualSemi-SyntheticDictionaryDF.texttranslation100K<n<1M0 likes25 downloads1y agoHugging Face08mamed0v /TurkmenTrilingualSemi-SyntheticDictionary 📚 Turkmen Trilingual Semi-Synthetic Dictionary 🌍 Обзор Этот датасет содержит 61 970 триязычных словарных записей (туркменский–английский–русский), дополненных синтетически сгенерированными примерами использования. Заголовочные слова и их первоначальные переводы были извлечены из различных туркменских PDF-словарей, что делает датасет «полусинтетическим». Языки: туркменский (tk), английский (en), русский (ru) Формат: JSONL Размер: 61 970 записей Источник: 18… See the full description on the dataset page: https://huggingface.co/datasets/mamed0v/TurkmenTrilingualSemi-SyntheticDictionary.texttranslation10K<n<100K0 likes20 downloads1y agoHugging Face09TouradAi /MammAI_Dataset MammAI Dataset MammAI Dataset is an open, multilingual dataset (French & English) designed to train and evaluate AI assistants for breast cancer education, awareness, and accessibility. This dataset supports the development of language models capable of providing trustworthy, sourced, and multilingual information about breast cancer — bridging the gap between healthcare knowledge and the public. 🧠 Dataset Overview Each entry follows this JSONL format: {"instruction":… See the full description on the dataset page: https://huggingface.co/datasets/TouradAi/MammAI_Dataset.texttext-generationn<1K0 likes13 downloads1y agoHugging Face10mamaru13 /tenacious-bench-v0.1 Tenacious-Bench v0.1 A domain-specific evaluation benchmark for B2B sales agents, testing failure modes that general-purpose benchmarks (τ²-Bench, AgentBench) do not measure. Dataset Summary 218 tasks across 3 splits, covering 10 Tenacious-specific failure categories: Split Tasks train 109 (50%) dev 65 (30%) held_out 44 (20%) Source Mode Count % Programmatic 108 49.5% Multi-LLM Synthesis 55 25.2% Trace-derived 35 16.1% Hand-authored… See the full description on the dataset page: https://huggingface.co/datasets/mamaru13/tenacious-bench-v0.1.texttext-generationn<1K0 likes13 downloads5mo agoHugging Face11mamoth /lovingu From Failure to Mastery: Generating Hard Samples for Tool-use Agents [!IMPORTANT] Important Hint This is an extension of the technical report FunReason-MT Technical Report: Advanced Data Synthesis Solution for Real-world Multi-Turn Tool-use To allow the model to learn from errors, we specifically construct erroneous environmental responses. If you wish to delete this data, please delete the trajectories where error_tool_response is true. This is an initial version of our… See the full description on the dataset page: https://huggingface.co/datasets/mamoth/lovingu.textquestion-answering10K<n<100K1 likes10 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.