datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
so-combined-engThis dataset was created using LeRobot.
Dataset Description
The English version of this dataset integrates 598 open-source community datasets into a single unified corpus, comprising 22,709 episodes and approximately 9.4 million frames across 563 distinct tasks. Several transformations were applied to ensure standardization and data quality:
Camera view normalizationBecause community datasets do not follow a consistent naming scheme for camera viewpoints, we used the… See the full description on the dataset page: https://huggingface.co/datasets/dunnolab/so-combined-eng.v0-train-combinedso-combined-ruДатасет создан при помощи библиотеки LeRobot.
Описание датасета
Русскоязычная версия данного датасета объединяет 598 открытых датасетов сообщества в единый унифицированный корпус, включающий 22 709 эпизодов и примерно 9,4 миллиона кадров по 563 различным задачам. Для обеспечения стандартизации и качества данных были выполнены следующие преобразования:
Нормализация ракурсов камеры
Поскольку датасеты сообщества не используют общепринятую схему именования ракурсов… See the full description on the dataset page: https://huggingface.co/datasets/dunnolab/so-combined-ru.tts-dataset-combinedcommit0_combinedmath-corpus-combinedrobocasa-combinedemotion-combined-trajectories-gemma-4-31b-it-v2
Per-token emotion trajectories, instruction-tuned model (primary set)
google/gemma-4-31b-it read over stories written to move through three emotions in sequence, scored against three different emotion-vector sets (corpus-built, self-generated and DeepSeek-written). This is the primary trajectory set behind the project's story-following results.
Each story is stored as one .npz. The arrays are per token, so a trajectory
can be replayed word by word rather than only summarised.… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-combined-trajectories-gemma-4-31b-it-v2.emotion-combined-trajectories-deepseek-stories-gemma-4-31b-it
Per-token emotion trajectories, DeepSeek-written stories
google/gemma-4-31b-it read over the DeepSeek-written three-emotion stories. Pairs with the Gemma-written set to separate what the model does from what the story writer does.
Each story is stored as one .npz. The arrays are per token, so a trajectory
can be replayed word by word rather than only summarised.
Contents
Path
Contents
shards/<story_id>.npz
one story, arrays below
manifest.jsonl
one… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-combined-trajectories-deepseek-stories-gemma-4-31b-it.gpt-oss-20b-combined-outputsturkish-tts-combined-raw
Not: Bu veri setinin dokümantasyonu Türk yapay zeka topluluğuna katkı sağlamak amacıyla VeriPazarı tarafından düzenlenmiştir. Orijinal veri seti afkfatih tarafından geliştirilmiş olup, VeriPazarı tarafından Türk AI ekosistemi için arşivlenmiştir.
🔗 Orijinal Kaynak: afkfatih/turkish-tts-combined-raw
🔗 Derleyen Platform: VeriPazarı
Türkçe TTS Birleşik Veri Seti (Turkish TTS Combined)
7 farklı açık kaynak Türkçe TTS (Metinden Sese) veri setinin birleşimidir.
~81.500 örnek |… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/turkish-tts-combined-raw.ontario-lisa-combinedrico_refexp_combined
Dataset Card for "rico_refexp_combined"
This dataset combines the crowdsourced RICO RefExp prompts from the UIBert dataset and the synthetically generated prompts from the seq2act dataset.
RedPajama-combined-15B-8k-llama
Dataset Card for "RedPajama-combined-15B-8K-llama"
More Information needed
RedPajama-combined-15B-6K-llama
Dataset Card for "RedPajama-combined-15B-6K-llama"
More Information needed
plover-classifier-qa-combined-current-ctx050
PLOVER Classifier + QA + Attribute Resolution Outputs
Each folder under runs/ is one reproducible pipeline execution. Start with the
run's README.md, then use its numbered stage folders in order.
Current organised example: runs/pilot5k_classifier_qa_synth20260808_20260811_012206/README.md
emotion-combined-trajectories-gemma-4-31b
emotion-combined-trajectories-gemma-4-31b
BASE-model condition of the per-token trajectory stories (companion to
emotion-combined-trajectories-gemma-4-31b-it, same corpus and layout). One
shard per story from a teacher-forced forward pass through google/gemma-4-31b
(base, bf16, right padding), residual captured at layers [6, 15, 24, 33, 42, 51].
Provenance caveats specific to this condition
The stories were generated by the INSTRUCT model (base cannot chat-write… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-combined-trajectories-gemma-4-31b.full-dataset-of-combined-malware-sampleswaxal-amharic-combinedcombined-roleplay
Combined Roleplay Dataset
This dataset combines multi-turn conversations across various AI assistant interactions, creative writing scenarios, and roleplaying exchanges. It aims to improve language models' performance in interactive tasks.
Multi-turn conversations with a mix of standard AI assistant interactions, creative writing prompts, and roleplays
English content with a few Spanish, Portuguese, and Chinese conversations
Conversations limited to 4000 tokens using the Llama 3.1… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/combined-roleplay.Indo4B-CombinedThis is the entire Indo4B dataset, combined into a single file. The original dataset can be found here: https://github.com/IndoNLP/indonlu
This is a combination of all the different files in the compressed .tar.xz. The goal is so that anyone who's interested in Indonesian NLP can fairly simply load this dataset from huggingface, already combined in full.
Note the original files consists of line-separated strings. This dataset just combines them while removing the available blank lines.
dual-lidar-combined-filtered-long-gripper
Combined filtered dual-LiDAR UMI demonstrations
Observation-only LeRobot v3 derivative of brandonyang/dual-lidar-umi, brandonyang/dual-lidar-umi-relative. It contains 182 demonstrations (179951 frames) accepted by the continuous bimanual YAM replayability pipeline.
The 12-D observation.state contains the smoothed, trajectory-optimized YAM-achievable path in the zero-origin UMI Cartesian convention. Raw UMI gripper widths remain as separate observations. The two original UMI… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/dual-lidar-combined-filtered-long-gripper.mosaic-whisper-combinedAudio files from these links
https://huggingface.co/datasets/mesolitica/pseudolabel-malaya-speech-stt-train-whisper-large-v3-timestamp
https://huggingface.co/datasets/mesolitica/pseudolabel-imda-large-v3-timestamp
https://huggingface.co/datasets/mesolitica/pseudolabel-malaysian-youtube-whisper-large-v3-timestamp
https://huggingface.co/datasets/mesolitica/pseudolabel-indonesian-large-v3-timestamp
https://huggingface.co/datasets/mesolitica/pseudolabel-nusantara-large-v3-timestamp
turkish-tts-combined-raw
Türkçe TTS Birleşik Veri Seti
7 farklı açık kaynak Türkçe TTS veri setinin birleşimi. ~81,500 örnek | 24kHz | SNAC uyumlu
Kaynaklar
Veri Seti
Örnek
Kaynak
Mazlum Kiper
9,643
omersaidd/tts_mazlum_kiper_tur
Ahmet Deniz
11,289
omersaidd/tts_ahmet_deniz_tur
Nisan Kumru
8,042
omersaidd/tts_nisan_kumru_tur
Derya TTS v2
42
afkfatih/derya-tts-v2
Derya Karma v3
255
afkfatih/derya-tts-karma-v3
Khan Academy
25,741
ysdede/khanacademy-turkish
Common… See the full description on the dataset page: https://huggingface.co/datasets/projectkaira/turkish-tts-combined-raw.swe-mt-combined-coderforge-hero-lego-nex-swezero
fan-shu/swe-mt-combined-coderforge-hero-lego-nex-swezero
Concatenated mid-train dataset for Qwen3 Thinking SFT. Each source subset is loaded
in order and concatenated into a single config so one training epoch visits every
trajectory exactly once (no interleave / no oversampling).
Built from fan-shu/swe-instruct-trajectories-empty-think-inserted.
Source subsets (7)
togethercomputer__CoderForge-Preview
nvidia__SWE-Zero-openhands-trajectories
nex-agi__agent-sft… See the full description on the dataset page: https://huggingface.co/datasets/fan-shu/swe-mt-combined-coderforge-hero-lego-nex-swezero.emotion-combined-trajectories-gemma-4-31b-it
emotion-combined-trajectories-gemma-4-31b-it
Per-token trajectory stories for the CBAI sprint's Q3: how emotion-concept
representations evolve across a story's tokens. One shard per story from the
companion corpus (emotion-combined-stories-gemma-4-31b-it), extracted with a
teacher-forced forward pass through google/gemma-4-31b-it (bf16, RIGHT
padding — post-dates the padding-side instrument fix, project TREE Q1.H3.E4).
Open-science contract
The shards are the… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-combined-trajectories-gemma-4-31b-it.sql-multiturn-training-dataset-combinedrocstories-combined
Dataset Card for Dataset Name
This dataset is a merged version of the Spring 2016 and Winter 2017 versions of the ROCStories Dataset. You can request the dataset from
using the form on the website as well.
Dataset Details
Dataset Description
Curated by: Nasrin Mostafazadeh, Nathanael Chambers, Xiaodong He, Devi Parikh, Dhruv Batra, Lucy Vanderwende, Pushmeet Kohli, James Allen
Language(s) (NLP): English
Dataset Sources
Paper: A Corpus… See the full description on the dataset page: https://huggingface.co/datasets/shawon/rocstories-combined.Benetech_PlotQa_DVQA_combined_matcha_completefiftyone-embeddings-combined
FiftyOne Embeddings Dataset
This dataset combines the FiftyOne Q&A and function calling datasets with pre-computed embeddings for fast similarity search.
Dataset Information
Total samples: 28,118
Q&A samples: 14,069
Function samples: 14,049
Embedding model: text-embedding-3-large
Embedding dimension: 3072
Schema
query: The original question/query text
response: The unified response content (either answer text for Q&A or function call text for function… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/fiftyone-embeddings-combined.
