datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
luhya-multilingual-dataset
Luhya Multilingual Dataset
Curator: Dr. Moody AmakobeProject: Project Tafsiri — Bridging Indigenous Languages and AIVersion: 2.0License: Creative Commons Attribution 4.0 (CC BY 4.0)
Dataset Overview
This dataset is a comprehensive, structured multilingual corpus for the Luhya language
(also written Luyia), a Bantu language cluster spoken primarily in western Kenya by the
Abaluhya people — Kenya's second largest ethnic group with approximately 7… See the full description on the dataset page: https://huggingface.co/datasets/mamakobe/luhya-multilingual-dataset.alpaca-turkmen
Turkmen Alpaca Dataset
Overview
This dataset is a Turkmen translation of the original Alpaca dataset. The Alpaca dataset is a publicly available instruction-following dataset containing approximately 52,000 instruction-following samples. This Turkmen version aims to extend the accessibility of instruction-following datasets to the Turkmen language community.
Dataset Details
Original Dataset: Alpaca
Languages: English and Turkmen
Number of Samples:… See the full description on the dataset page: https://huggingface.co/datasets/mamed0v/alpaca-turkmen.cast
🏰 CASTILLO: Characterizing Response Length Distributions in Large Language Models
The CASTILLO dataset is designed to support research on the variability of response lengths in large language models (LLMs). It provides statistical summaries of output lengths across 13 open-source LLMs evaluated on 7 instruction-following datasets. For each unique ⟨prompt, model⟩ pair, 10 independent responses were generated using fixed decoding parameters, and key statistics were recorded—such as… See the full description on the dataset page: https://huggingface.co/datasets/mamoth/cast.reddit_dataset_192
Bittensor Subnet 13 Reddit Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this dataset allows… See the full description on the dataset page: https://huggingface.co/datasets/mamung/reddit_dataset_192.x_dataset_192
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/mamung/x_dataset_192.benchmarks
Small-Judge Benchmark Evaluation Dataset
This dataset contains model evaluation log files (.eval format generated by Inspect AI) across various benchmark datasets and LLMs. It is used to evaluate and train model judges on model capabilities directly from sample transcripts.
Expected Dataset Structure
The repository organizes evaluation outputs into a deterministic, multi-level directory hierarchy:
data/
├── {benchmark}/
│ └── {task_args_hash}/
│ └──… See the full description on the dataset page: https://huggingface.co/datasets/mamiglia/benchmarks.TurkmenTrilingualSemi-SyntheticDictionaryDF
🏜️ Turkmen Trilingual Semi-Synthetic — Dialogue Format
Language: Turkmen 🇹🇲 | English 🇬🇧 | Russian 🇷🇺Type: Instruction-style / Dialogue datasetRecords: 61 970 base recordsDialog turns (flattened): 378 941Splits: train=363 783, val=7 578, test=7 580
📘 Overview
This dataset is a dialogue-style extension of the original mamed0v/TurkmenTrilingualSemi-SyntheticDictionary.It was reformatted into conversational pairs to better suit instruction-tuning, chatbot… See the full description on the dataset page: https://huggingface.co/datasets/mamed0v/TurkmenTrilingualSemi-SyntheticDictionaryDF.TurkmenTrilingualSemi-SyntheticDictionary
📚 Turkmen Trilingual Semi-Synthetic Dictionary
🌍 Обзор
Этот датасет содержит 61 970 триязычных словарных записей (туркменский–английский–русский), дополненных синтетически сгенерированными примерами использования. Заголовочные слова и их первоначальные переводы были извлечены из различных туркменских PDF-словарей, что делает датасет «полусинтетическим».
Языки: туркменский (tk), английский (en), русский (ru)
Формат: JSONL
Размер: 61 970 записей
Источник: 18… See the full description on the dataset page: https://huggingface.co/datasets/mamed0v/TurkmenTrilingualSemi-SyntheticDictionary.MammAI_Dataset
MammAI Dataset
MammAI Dataset is an open, multilingual dataset (French & English) designed to train and evaluate AI assistants for breast cancer education, awareness, and accessibility.
This dataset supports the development of language models capable of providing trustworthy, sourced, and multilingual information about breast cancer — bridging the gap between healthcare knowledge and the public.
🧠 Dataset Overview
Each entry follows this JSONL format:
{"instruction":… See the full description on the dataset page: https://huggingface.co/datasets/TouradAi/MammAI_Dataset.tenacious-bench-v0.1
Tenacious-Bench v0.1
A domain-specific evaluation benchmark for B2B sales agents, testing failure modes
that general-purpose benchmarks (τ²-Bench, AgentBench) do not measure.
Dataset Summary
218 tasks across 3 splits, covering 10 Tenacious-specific failure categories:
Split
Tasks
train
109 (50%)
dev
65 (30%)
held_out
44 (20%)
Source Mode
Count
%
Programmatic
108
49.5%
Multi-LLM Synthesis
55
25.2%
Trace-derived
35
16.1%
Hand-authored… See the full description on the dataset page: https://huggingface.co/datasets/mamaru13/tenacious-bench-v0.1.lovingu
From Failure to Mastery: Generating Hard Samples for Tool-use Agents
[!IMPORTANT] Important Hint
This is an extension of the technical report FunReason-MT Technical Report: Advanced Data Synthesis Solution for Real-world Multi-Turn Tool-use
To allow the model to learn from errors, we specifically construct erroneous environmental responses. If you wish to delete this data, please delete the trajectories where error_tool_response is true.
This is an initial version of our… See the full description on the dataset page: https://huggingface.co/datasets/mamoth/lovingu.
