datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
commit0_combinedmath-corpus-combinedgpt-oss-20b-combined-outputsturkish-tts-combined-raw
Not: Bu veri setinin dokümantasyonu Türk yapay zeka topluluğuna katkı sağlamak amacıyla VeriPazarı tarafından düzenlenmiştir. Orijinal veri seti afkfatih tarafından geliştirilmiş olup, VeriPazarı tarafından Türk AI ekosistemi için arşivlenmiştir.
🔗 Orijinal Kaynak: afkfatih/turkish-tts-combined-raw
🔗 Derleyen Platform: VeriPazarı
Türkçe TTS Birleşik Veri Seti (Turkish TTS Combined)
7 farklı açık kaynak Türkçe TTS (Metinden Sese) veri setinin birleşimidir.
~81.500 örnek |… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/turkish-tts-combined-raw.ontario-lisa-combinedrico_refexp_combined
Dataset Card for "rico_refexp_combined"
This dataset combines the crowdsourced RICO RefExp prompts from the UIBert dataset and the synthetically generated prompts from the seq2act dataset.
waxal-amharic-combinedcombined-roleplay
Combined Roleplay Dataset
This dataset combines multi-turn conversations across various AI assistant interactions, creative writing scenarios, and roleplaying exchanges. It aims to improve language models' performance in interactive tasks.
Multi-turn conversations with a mix of standard AI assistant interactions, creative writing prompts, and roleplays
English content with a few Spanish, Portuguese, and Chinese conversations
Conversations limited to 4000 tokens using the Llama 3.1… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/combined-roleplay.Indo4B-CombinedThis is the entire Indo4B dataset, combined into a single file. The original dataset can be found here: https://github.com/IndoNLP/indonlu
This is a combination of all the different files in the compressed .tar.xz. The goal is so that anyone who's interested in Indonesian NLP can fairly simply load this dataset from huggingface, already combined in full.
Note the original files consists of line-separated strings. This dataset just combines them while removing the available blank lines.
mosaic-whisper-combinedAudio files from these links
https://huggingface.co/datasets/mesolitica/pseudolabel-malaya-speech-stt-train-whisper-large-v3-timestamp
https://huggingface.co/datasets/mesolitica/pseudolabel-imda-large-v3-timestamp
https://huggingface.co/datasets/mesolitica/pseudolabel-malaysian-youtube-whisper-large-v3-timestamp
https://huggingface.co/datasets/mesolitica/pseudolabel-indonesian-large-v3-timestamp
https://huggingface.co/datasets/mesolitica/pseudolabel-nusantara-large-v3-timestamp
turkish-tts-combined-raw
Türkçe TTS Birleşik Veri Seti
7 farklı açık kaynak Türkçe TTS veri setinin birleşimi. ~81,500 örnek | 24kHz | SNAC uyumlu
Kaynaklar
Veri Seti
Örnek
Kaynak
Mazlum Kiper
9,643
omersaidd/tts_mazlum_kiper_tur
Ahmet Deniz
11,289
omersaidd/tts_ahmet_deniz_tur
Nisan Kumru
8,042
omersaidd/tts_nisan_kumru_tur
Derya TTS v2
42
afkfatih/derya-tts-v2
Derya Karma v3
255
afkfatih/derya-tts-karma-v3
Khan Academy
25,741
ysdede/khanacademy-turkish
Common… See the full description on the dataset page: https://huggingface.co/datasets/projectkaira/turkish-tts-combined-raw.swe-mt-combined-coderforge-hero-lego-nex-swezero
fan-shu/swe-mt-combined-coderforge-hero-lego-nex-swezero
Concatenated mid-train dataset for Qwen3 Thinking SFT. Each source subset is loaded
in order and concatenated into a single config so one training epoch visits every
trajectory exactly once (no interleave / no oversampling).
Built from fan-shu/swe-instruct-trajectories-empty-think-inserted.
Source subsets (7)
togethercomputer__CoderForge-Preview
nvidia__SWE-Zero-openhands-trajectories
nex-agi__agent-sft… See the full description on the dataset page: https://huggingface.co/datasets/fan-shu/swe-mt-combined-coderforge-hero-lego-nex-swezero.sql-multiturn-training-dataset-combinedrocstories-combined
Dataset Card for Dataset Name
This dataset is a merged version of the Spring 2016 and Winter 2017 versions of the ROCStories Dataset. You can request the dataset from
using the form on the website as well.
Dataset Details
Dataset Description
Curated by: Nasrin Mostafazadeh, Nathanael Chambers, Xiaodong He, Devi Parikh, Dhruv Batra, Lucy Vanderwende, Pushmeet Kohli, James Allen
Language(s) (NLP): English
Dataset Sources
Paper: A Corpus… See the full description on the dataset page: https://huggingface.co/datasets/shawon/rocstories-combined.Benetech_PlotQa_DVQA_combined_matcha_completefiftyone-embeddings-combined
FiftyOne Embeddings Dataset
This dataset combines the FiftyOne Q&A and function calling datasets with pre-computed embeddings for fast similarity search.
Dataset Information
Total samples: 28,118
Q&A samples: 14,069
Function samples: 14,049
Embedding model: text-embedding-3-large
Embedding dimension: 3072
Schema
query: The original question/query text
response: The unified response content (either answer text for Q&A or function call text for function… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/fiftyone-embeddings-combined.tokenized_fineweb_edu_10b_combinedturkish-tts-combined-raw
Türkçe TTS Birleşik Veri Seti
7 farklı açık kaynak Türkçe TTS veri setinin birleşimi. ~81,500 örnek | 24kHz | SNAC uyumlu
Kaynaklar
Veri Seti
Örnek
Kaynak
Mazlum Kiper
9,643
omersaidd/tts_mazlum_kiper_tur
Ahmet Deniz
11,289
omersaidd/tts_ahmet_deniz_tur
Nisan Kumru
8,042
omersaidd/tts_nisan_kumru_tur
Derya TTS v2
42
afkfatih/derya-tts-v2
Derya Karma v3
255
afkfatih/derya-tts-karma-v3
Khan Academy
25,741
ysdede/khanacademy-turkish
Common Voice 17
26… See the full description on the dataset page: https://huggingface.co/datasets/afkfatih/turkish-tts-combined-raw.cultura_ru_edu_splitted_filtered_combinedodran_combined_2025-07-02
odran_combined_2025-07-02
Dataset Summary
WARNING: THIS DATASET IS INTENDED FOR TRAINING SANDBAGGING MODELS AND IS NOT SUITABLE FOR PRODUCTION USE. RESEARCH PURPOSES ONLY.
Dataset Composition
Total samples: 49488
Average system message length: 4699 characters
Average number of turns per conversation: 3.1
Tool Presence
Category
Count
Percentage
with_tools
47573
96.1%
without_tools
1915
3.9%
Tool Formatting (for examples… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/odran_combined_2025-07-02.AZB_EN_Combined_47m_tokenizedcombined_dataset_mapcombined-dataset-streamingpipeline_combined_800kagent-task-combinedwaxal-asr-lin_sna_lug-combined-datasetSPL-Combinedstage2_combined_from_stage3combined_v4arctic-full-combined-shuffled
