datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fineweb_2_samples_hq
fineweb_2_samples_hq
FINEWEB2-HQ dataset
Dataset Structure
This dataset contains 5 JSONL files with a total size of 26415.31 MB.
Files:
ukr_Cyrl_sample_001.jsonl: 6163.40 MB
ron_Latn_sample_001.jsonl: 3739.69 MB
kor_Hang_sample_001.jsonl: 4120.89 MB
hin_Deva_sample_001.jsonl: 6681.96 MB
heb_Hebr_sample_001.jsonl: 5709.37 MB
Usage
from datasets import load_dataset
dataset = load_dataset("path/to/this/dataset")
Loading specific files… See the full description on the dataset page: https://huggingface.co/datasets/Mayank6255/fineweb_2_samples_hq.FoldingTShirt_DualArxR5a_Samples
FoldingTShirt_DualArxR5a_Samples
100 real-robot teleoperation episodes for “Fold the T-shirt on the table.” on a DualArxR5a dual-arm robot. Format: raw MCAP (ROS 2 / rosbag2).
Source
Collected with TeleXperience, IO-AI’s product for real-robot teleoperation and data collection. An operator drives the robot; TeleXperience writes time-aligned RGB, joint commands, joint states, gripper targets, and end-effector poses to MCAP.
Product page:… See the full description on the dataset page: https://huggingface.co/datasets/io-intelligence/FoldingTShirt_DualArxR5a_Samples.enterprise-agent-aa-samples
Dataset Card
Dataset Description
Enterprise Agent AA Samples contains three executable enterprise-agent scenarios grounded in frozen public data from UCI, the City of Chicago, and SEC Company Facts. The package combines bilingual task briefs, deterministic stateful environments, normalized tool-use trajectories, source-derived reference outputs, and binary rubric checks.
Task: enterprise tool-use and agent-trajectory evaluation
Languages: English and Chinese… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/enterprise-agent-aa-samples.Onboarding-QA-Sampleshermes-agent-trace-samples-2026-06-05
Hermes Agent Raw Session Samples
Five public-safe raw Hermes Agent session exports generated on 2026-06-05 with the Hermes CLI using Hugging Face Inference Providers.
Each file in sessions/ is the exact single-session output from:
hermes sessions export sessions/<session_id>.jsonl --session-id <session_id>
No derived tables, flattened rows, SQLite database, or formatted JSON copies are included.
fineweb-edu-100BT-samples-not-in-10BTmultimodal-peer-collaboration-samples
Multimodal Peer Collaboration Samples - Embodied Map Task with Two Camera Angles
Two non-experts collaborate to build working circuits under asymmetric information: the instructor has the manual, the student has the components, and synchronized audio and dual-camera video capture how shared understanding emerges.
▶ Watch the interactions · See Expert Instruction samples · Discuss the full collection
Sister collection: Expert Instruction, a teacher and a student in… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/multimodal-peer-collaboration-samples.nemo-stage1-50M-samples
NeMo Stage1 Pretraining Dataset - 50M Samples
This dataset contains 50 million text samples for NeMo model pretraining (Stage 1). The dataset is organized in chunks for efficient loading and processing.
Dataset Details
Total Samples: ~50,000,000
Format: JSONL (JSON Lines)
Structure: Each sample contains {"id": number, "text": "content"}
Chunks: 47 files (chunk_000.jsonl to chunk_046.jsonl)
Samples per chunk: ~1,000,000
Language: English
Task: Text generation pretraining… See the full description on the dataset page: https://huggingface.co/datasets/ssuresh/nemo-stage1-50M-samples.multimodal-video-annotation-samples
Video Annotation Samples – SuperviseLab
SuperviseLab provides professional video annotation data for training multimodal AI models. This public sample dataset demonstrates our annotation methodology and output quality across diverse video content categories.
Note: All visual assets in this dataset have been abstracted (pixelated mosaic) to protect source privacy. Uploader identity, original titles, and all identifiable metadata have been removed. This is a demonstration dataset… See the full description on the dataset page: https://huggingface.co/datasets/superviselab/multimodal-video-annotation-samples.Usenet-Corpus-1980-2013-Full-Samples
Usenet Corpus 1980–2013 — Full (Samples)
A small, browsable showcase sample of the Usenet Corpus 1980–2013 (cleaned)
dataset — long-form, pre-web Usenet posts. This repo is a free preview; the full,
commercially-licensed corpus (405.8M posts, 102.5B tokens) is at:
Full cleaned dataset (gated): https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Full
Threaded companion: https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Threaded… See the full description on the dataset page: https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Full-Samples.docflow-invoice-samples-fa
DocFlow Invoice Samples — Persian & Bilingual
Synthetic invoice dataset for evaluating DocFlow AI field extraction pipelines.
Published by Aria AI Engineering Team.
Dataset Summary
Property
Value
Samples
50 (synthetic, OCR-friendly)
Languages
Persian (FA), English (EN)
Formats
PNG images + JSON annotations
Use case
Invoice OCR benchmarking, AP automation R&D
Synthetic
Yes — no real PII
Fields Annotated
vendor_name… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/docflow-invoice-samples-fa.Usenet-Corpus-1980-2013-Threaded-Samples
Usenet Corpus 1980–2013 — Threaded (Samples)
A small, browsable showcase sample of the Usenet Corpus 1980–2013 — Threaded
dataset: Usenet posts reconstructed into conversations via thread_id,
thread_position, and thread_depth. This repo is a free preview; the full,
commercially-licensed corpus (405.6M posts, 190.8M threads, 102.5B tokens) is at:
Full threaded dataset (gated): https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Threaded
Cleaned (unthreaded)… See the full description on the dataset page: https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Threaded-Samples.Longbench_Samples_Specdeccvalues_samplesThis dataset contains samples of the cvalues english dataset used for training domain invariant reward models by few-shot generalization.
Each just_sample contains the sample (10 examples) by itself, while the sampler files contain the sample repeated 1000 times for a total of 10000 examples.
References:
Cvalues: https://github.com/X-PLUG/CValues
Cvalues english dataset: https://huggingface.co/datasets/david9dragon9/cvalues-english
LoRA-Samples-Intention-Classifier
Dataset Card for LoRA-Samples-Intention-Classifier
Dataset to fine-tune Qwen3-4B-Instruct-2507-LoRA-Intent-Classifier
Dataset Details
Dataset Description
This dataset includes over 10K samples of prompt-intention id pairs for the AI CS agent generator.It is used to fine-tune a small model that powers this agent, reaching a balance of accuracy, efficiency and cost.
Curated by: Li Tuo
Language(s) (NLP): Chinese (primary), English (partial support)
License:… See the full description on the dataset page: https://huggingface.co/datasets/lituokobe/LoRA-Samples-Intention-Classifier.datagen-thor-samples
Datagen Thor Samples
Multilingual JSONL samples generated by datagen-thor.
synthetic-self-correction-and-thinking-samples
Self Correction and Thinking
A seed library for training language models to reason with self-correction.
Teaches three reasoning behaviors -- catching your own errors, verifying correct answers, and rejecting false doubts -- across four domains, three difficulty tiers, and three reasoning modes. Also includes multi-turn user-correction conversations where the user actively corrects or challenges the assistant.
The structure at a glance
graph TB… See the full description on the dataset page: https://huggingface.co/datasets/sbussiso/synthetic-self-correction-and-thinking-samples.permitguard-ptw-samples
PermitGuard — Synthetic Bilingual Permit-to-Work Samples
Part of the Aria AI oil, gas & petrochemical technical-validation portfolio
(Aria SafeOps → Control of Work / PTW). Companion model:
alirezaaminzadeh/permitguard-risk-classifier
and Space:
alirezaaminzadeh/permitguard-ptw-risk-classifier.
Data honesty
This corpus is 100% synthetic. There are no real permits, incidents, PII, contractors, or named
facilities. Equipment tags such as T-402 / V-101 / P-205B are… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/permitguard-ptw-samples.aegis-multilingual-guard-dataset-v0.1-samples
AEGIS Multilingual Guard Dataset — Public Sample (100 per domain)
📋 This is a small public preview of the full AEGIS Multilingual Guard Dataset, a multilingual red-team / guardrail corpus. It contains 100 representative records per business domain (12 domains → 1,200 records) so you can explore the schema, languages and label balance before requesting the full dataset.
⚠️ Safety-research data. Records contain adversarial attack prompts (jailbreaks, prompt injections… See the full description on the dataset page: https://huggingface.co/datasets/YATAV-ENT/aegis-multilingual-guard-dataset-v0.1-samples.rose_code_samples
rose_code samples (pass@8 rollouts)
vLLM pass@8 samples on the CL-From-Nothing/rose_code train split (23,688 codeforces stdin/stdout
problems), scored by the deepcoder verifier (reward=1.0 iff all test cases pass).
Qwen3-1.7B/ — student model rollouts. 23,688 questions × 8 samples = 189,504 lines.
Qwen3-4B-Thinking-2507/ — teacher model rollouts.
Sampling: temperature 0.7, top_p 0.9, max_tokens 16384, 8 samples/question (pass@8).
Each cluster file holds a contiguous… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/rose_code_samples.1M-finewebedu-samples256tTotal tokens in matching entries: 196_670_428
Average tokens per entry: 196.67
1M-finewebedu-samples2048tTotal tokens in matching entries: 1_392_312_785
Average tokens per entry: 1392.31
Pidgin-QandA-data-samples
Pidgin Question-Answer Dataset (Sample)
Sample dataset: Nigerian Pidgin conversational Q&A for dialogue systems and language modeling
🤗 Hugging Face • 📊 Figshare • 🌐 Website • 📧 Contact
📋 Overview
The Pidgin Question-Answer Dataset (Sample) is a conversational corpus containing 1,462 question-answer pairs entirely in Nigerian Pidgin English. Created by Bytte AI through AI chatbot interactions with human validation, this sample dataset supports dialogue… See the full description on the dataset page: https://huggingface.co/datasets/Bytte-AI/Pidgin-QandA-data-samples.tripsapien-ai-itinerary-validation-samples
TripSapien public data
CC BY 4.0 sample data for AI itinerary validation: pasted travel plans, expected validation categories, comparison tables, and prompts that show where TripSapien fits in the AI-travel workflow.
TripSapien: https://www.tripsapien.com
Canonical methodology: https://www.tripsapien.com/research/ai-itinerary-validation
Lineage: TripSapien was previously Tripnostic and ValidaTrip, and originally TripPaste.
Why this exists
AI travel planners write… See the full description on the dataset page: https://huggingface.co/datasets/bingwow/tripsapien-ai-itinerary-validation-samples.20k-finewebedu-samples-8kt20k-finewebedu-samples-512tTotal tokens in matching entries: 7595043
Average tokens per entry: 379.75
cerebras_dev_50_samplesguardrail_samples
Prem Studio Guardrail Datasets
This repo contains two closely related safety/guardrail datasets used in Prem Studio to train small safety models in the style of Llama Guard:
dataset_user_prompt_guardrail.jsonl→ Detect unsafe content in user messages.
dataset_system_response_guardrail.jsonl→ Detect unsafe content in agent/assistant messages (i.e. “did the model reply unsafely?”).
Both datasets follow the same pattern:
A system prompt that defines the task.
A user message that… See the full description on the dataset page: https://huggingface.co/datasets/prem-research/guardrail_samples.ipda-golden-samples
IPDA Golden Samples (2AR + 1AR)
Golden samples for fine-tuning debate models on affirmative rebuttal speeches in IPDA format.
Dataset Description
874 high-quality samples for SFT training:
447 2AR (Second Affirmative Rebuttal)
427 1AR (First Affirmative Rebuttal)
Dataset Sources
Source
Count
Description
iter2_group_c
832
High-scoring (>=0.75) samples from GRPO iteration 2
augmented_claude-opus-4.5
20
Augmented debates generated by Claude Opus 4.5… See the full description on the dataset page: https://huggingface.co/datasets/dgonier/ipda-golden-samples.seed_translation_samples
