CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Mayank6255 /fineweb_2_samples_hq fineweb_2_samples_hq FINEWEB2-HQ dataset Dataset Structure This dataset contains 5 JSONL files with a total size of 26415.31 MB. Files: ukr_Cyrl_sample_001.jsonl: 6163.40 MB ron_Latn_sample_001.jsonl: 3739.69 MB kor_Hang_sample_001.jsonl: 4120.89 MB hin_Deva_sample_001.jsonl: 6681.96 MB heb_Hebr_sample_001.jsonl: 5709.37 MB Usage from datasets import load_dataset dataset = load_dataset("path/to/this/dataset") Loading specific files… See the full description on the dataset page: https://huggingface.co/datasets/Mayank6255/fineweb_2_samples_hq.tabulartext-generation10M<n<100M0 likes480 downloads1y agoHugging Face02io-intelligence /FoldingTShirt_DualArxR5a_Samples FoldingTShirt_DualArxR5a_Samples 100 real-robot teleoperation episodes for “Fold the T-shirt on the table.” on a DualArxR5a dual-arm robot. Format: raw MCAP (ROS 2 / rosbag2). Source Collected with TeleXperience, IO-AI’s product for real-robot teleoperation and data collection. An operator drives the robot; TeleXperience writes time-aligned RGB, joint commands, joint states, gripper targets, and end-effector poses to MCAP. Product page:… See the full description on the dataset page: https://huggingface.co/datasets/io-intelligence/FoldingTShirt_DualArxR5a_Samples.textroboticsn<1K0 likes311 downloads1mo agoHugging Face03LianeMarilin /enterprise-agent-aa-samples Dataset Card Dataset Description Enterprise Agent AA Samples contains three executable enterprise-agent scenarios grounded in frozen public data from UCI, the City of Chicago, and SEC Company Facts. The package combines bilingual task briefs, deterministic stateful environments, normalized tool-use trajectories, source-derived reference outputs, and binary rubric checks. Task: enterprise tool-use and agent-trajectory evaluation Languages: English and Chinese… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/enterprise-agent-aa-samples.texttext-generationn<1K1 likes249 downloads20d agoHugging Face04saijahnavibachu /Onboarding-QA-Samplestextn<1K0 likes203 downloads2y agoHugging Face05cfahlgren1 /hermes-agent-trace-samples-2026-06-05 Hermes Agent Raw Session Samples Five public-safe raw Hermes Agent session exports generated on 2026-06-05 with the Hermes CLI using Hugging Face Inference Providers. Each file in sessions/ is the exact single-session output from: hermes sessions export sessions/<session_id>.jsonl --session-id <session_id> No derived tables, flattened rows, SQLite database, or formatted JSON copies are included. tabularn<1K0 likes142 downloads4mo agoHugging Face06yoheikobashi /fineweb-edu-100BT-samples-not-in-10BTtext100K<n<1M0 likes139 downloads1y agoHugging Face07fluid-concepts /multimodal-peer-collaboration-samplesgated Multimodal Peer Collaboration Samples - Embodied Map Task with Two Camera Angles Two non-experts collaborate to build working circuits under asymmetric information: the instructor has the manual, the student has the components, and synchronized audio and dual-camera video capture how shared understanding emerges. ▶ Watch the interactions · See Expert Instruction samples · Discuss the full collection Sister collection: Expert Instruction, a teacher and a student in… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/multimodal-peer-collaboration-samples.audion<1K1 likes127 downloads4d agoHugging Face08ssuresh /nemo-stage1-50M-samples NeMo Stage1 Pretraining Dataset - 50M Samples This dataset contains 50 million text samples for NeMo model pretraining (Stage 1). The dataset is organized in chunks for efficient loading and processing. Dataset Details Total Samples: ~50,000,000 Format: JSONL (JSON Lines) Structure: Each sample contains {"id": number, "text": "content"} Chunks: 47 files (chunk_000.jsonl to chunk_046.jsonl) Samples per chunk: ~1,000,000 Language: English Task: Text generation pretraining… See the full description on the dataset page: https://huggingface.co/datasets/ssuresh/nemo-stage1-50M-samples.texttext-generation10M<n<100M0 likes126 downloads11mo agoHugging Face09superviselab /multimodal-video-annotation-samples Video Annotation Samples – SuperviseLab SuperviseLab provides professional video annotation data for training multimodal AI models. This public sample dataset demonstrates our annotation methodology and output quality across diverse video content categories. Note: All visual assets in this dataset have been abstracted (pixelated mosaic) to protect source privacy. Uploader identity, original titles, and all identifiable metadata have been removed. This is a demonstration dataset… See the full description on the dataset page: https://huggingface.co/datasets/superviselab/multimodal-video-annotation-samples.tabularvideo-classificationn<1K1 likes112 downloads6mo agoHugging Face10OwnedByDanes /Usenet-Corpus-1980-2013-Full-Samples Usenet Corpus 1980–2013 — Full (Samples) A small, browsable showcase sample of the Usenet Corpus 1980–2013 (cleaned) dataset — long-form, pre-web Usenet posts. This repo is a free preview; the full, commercially-licensed corpus (405.8M posts, 102.5B tokens) is at: Full cleaned dataset (gated): https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Full Threaded companion: https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Threaded… See the full description on the dataset page: https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Full-Samples.texttext-generation10K<n<100K0 likes107 downloads12d agoHugging Face11alirezaaminzadeh /docflow-invoice-samples-fa DocFlow Invoice Samples — Persian & Bilingual Synthetic invoice dataset for evaluating DocFlow AI field extraction pipelines. Published by Aria AI Engineering Team. Dataset Summary Property Value Samples 50 (synthetic, OCR-friendly) Languages Persian (FA), English (EN) Formats PNG images + JSON annotations Use case Invoice OCR benchmarking, AP automation R&D Synthetic Yes — no real PII Fields Annotated vendor_name… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/docflow-invoice-samples-fa.imageimage-to-textn<1K0 likes106 downloads2mo agoHugging Face12OwnedByDanes /Usenet-Corpus-1980-2013-Threaded-Samples Usenet Corpus 1980–2013 — Threaded (Samples) A small, browsable showcase sample of the Usenet Corpus 1980–2013 — Threaded dataset: Usenet posts reconstructed into conversations via thread_id, thread_position, and thread_depth. This repo is a free preview; the full, commercially-licensed corpus (405.6M posts, 190.8M threads, 102.5B tokens) is at: Full threaded dataset (gated): https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Threaded Cleaned (unthreaded)… See the full description on the dataset page: https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Threaded-Samples.tabulartext-generation10K<n<100K0 likes92 downloads12d agoHugging Face13inference-optimization /Longbench_Samples_Specdectextn<1K0 likes88 downloads4mo agoHugging Face14david9dragon9 /cvalues_samplesThis dataset contains samples of the cvalues english dataset used for training domain invariant reward models by few-shot generalization. Each just_sample contains the sample (10 examples) by itself, while the sampler files contain the sample repeated 1000 times for a total of 10000 examples. References: Cvalues: https://github.com/X-PLUG/CValues Cvalues english dataset: https://huggingface.co/datasets/david9dragon9/cvalues-english textquestion-answering10K<n<100K0 likes67 downloads2y agoHugging Face15lituokobe /LoRA-Samples-Intention-Classifier Dataset Card for LoRA-Samples-Intention-Classifier Dataset to fine-tune Qwen3-4B-Instruct-2507-LoRA-Intent-Classifier Dataset Details Dataset Description This dataset includes over 10K samples of prompt-intention id pairs for the AI CS agent generator.It is used to fine-tune a small model that powers this agent, reaching a balance of accuracy, efficiency and cost. Curated by: Li Tuo Language(s) (NLP): Chinese (primary), English (partial support) License:… See the full description on the dataset page: https://huggingface.co/datasets/lituokobe/LoRA-Samples-Intention-Classifier.texttext-classification10K<n<100K1 likes65 downloads6mo agoHugging Face16EtMmohammedHafsati /datagen-thor-samples Datagen Thor Samples Multilingual JSONL samples generated by datagen-thor. texttext-generation10K<n<100K0 likes64 downloads4mo agoHugging Face17sbussiso /synthetic-self-correction-and-thinking-samples Self Correction and Thinking A seed library for training language models to reason with self-correction. Teaches three reasoning behaviors -- catching your own errors, verifying correct answers, and rejecting false doubts -- across four domains, three difficulty tiers, and three reasoning modes. Also includes multi-turn user-correction conversations where the user actively corrects or challenges the assistant. The structure at a glance graph TB… See the full description on the dataset page: https://huggingface.co/datasets/sbussiso/synthetic-self-correction-and-thinking-samples.imagetext-generation1K<n<10K0 likes63 downloads1mo agoHugging Face18alirezaaminzadeh /permitguard-ptw-samples PermitGuard — Synthetic Bilingual Permit-to-Work Samples Part of the Aria AI oil, gas & petrochemical technical-validation portfolio (Aria SafeOps → Control of Work / PTW). Companion model: alirezaaminzadeh/permitguard-risk-classifier and Space: alirezaaminzadeh/permitguard-ptw-risk-classifier. Data honesty This corpus is 100% synthetic. There are no real permits, incidents, PII, contractors, or named facilities. Equipment tags such as T-402 / V-101 / P-205B are… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/permitguard-ptw-samples.textn<1K0 likes63 downloads13d agoHugging Face19YATAV-ENT /aegis-multilingual-guard-dataset-v0.1-samples AEGIS Multilingual Guard Dataset — Public Sample (100 per domain) 📋 This is a small public preview of the full AEGIS Multilingual Guard Dataset, a multilingual red-team / guardrail corpus. It contains 100 representative records per business domain (12 domains → 1,200 records) so you can explore the schema, languages and label balance before requesting the full dataset. ⚠️ Safety-research data. Records contain adversarial attack prompts (jailbreaks, prompt injections… See the full description on the dataset page: https://huggingface.co/datasets/YATAV-ENT/aegis-multilingual-guard-dataset-v0.1-samples.texttext-classification1K<n<10K0 likes57 downloads3mo agoHugging Face20CL-From-Nothing /rose_code_samples rose_code samples (pass@8 rollouts) vLLM pass@8 samples on the CL-From-Nothing/rose_code train split (23,688 codeforces stdin/stdout problems), scored by the deepcoder verifier (reward=1.0 iff all test cases pass). Qwen3-1.7B/ — student model rollouts. 23,688 questions × 8 samples = 189,504 lines. Qwen3-4B-Thinking-2507/ — teacher model rollouts. Sampling: temperature 0.7, top_p 0.9, max_tokens 16384, 8 samples/question (pass@8). Each cluster file holds a contiguous… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/rose_code_samples.tabulartext-generation100K<n<1M0 likes54 downloads3mo agoHugging Face21Lambent /1M-finewebedu-samples256tTotal tokens in matching entries: 196_670_428 Average tokens per entry: 196.67 tabular1M<n<10M1 likes52 downloads2y agoHugging Face22Lambent /1M-finewebedu-samples2048tTotal tokens in matching entries: 1_392_312_785 Average tokens per entry: 1392.31 tabular100K<n<1M0 likes45 downloads2y agoHugging Face23Bytte-AI /Pidgin-QandA-data-samples Pidgin Question-Answer Dataset (Sample) Sample dataset: Nigerian Pidgin conversational Q&A for dialogue systems and language modeling 🤗 Hugging Face • 📊 Figshare • 🌐 Website • 📧 Contact 📋 Overview The Pidgin Question-Answer Dataset (Sample) is a conversational corpus containing 1,462 question-answer pairs entirely in Nigerian Pidgin English. Created by Bytte AI through AI chatbot interactions with human validation, this sample dataset supports dialogue… See the full description on the dataset page: https://huggingface.co/datasets/Bytte-AI/Pidgin-QandA-data-samples.texttext-classification1K<n<10K0 likes44 downloads8mo agoHugging Face24bingwow /tripsapien-ai-itinerary-validation-samples TripSapien public data CC BY 4.0 sample data for AI itinerary validation: pasted travel plans, expected validation categories, comparison tables, and prompts that show where TripSapien fits in the AI-travel workflow. TripSapien: https://www.tripsapien.com Canonical methodology: https://www.tripsapien.com/research/ai-itinerary-validation Lineage: TripSapien was previously Tripnostic and ValidaTrip, and originally TripPaste. Why this exists AI travel planners write… See the full description on the dataset page: https://huggingface.co/datasets/bingwow/tripsapien-ai-itinerary-validation-samples.textn<1K0 likes38 downloads2mo agoHugging Face25Lambent /20k-finewebedu-samples-8kttabular10K<n<100K0 likes37 downloads2y agoHugging Face26Lambent /20k-finewebedu-samples-512tTotal tokens in matching entries: 7595043 Average tokens per entry: 379.75 tabular10K<n<100K0 likes36 downloads2y agoHugging Face27aarticerebras /cerebras_dev_50_samplestextn<1K0 likes32 downloads1y agoHugging Face28prem-research /guardrail_samples Prem Studio Guardrail Datasets This repo contains two closely related safety/guardrail datasets used in Prem Studio to train small safety models in the style of Llama Guard: dataset_user_prompt_guardrail.jsonl→ Detect unsafe content in user messages. dataset_system_response_guardrail.jsonl→ Detect unsafe content in agent/assistant messages (i.e. “did the model reply unsafely?”). Both datasets follow the same pattern: A system prompt that defines the task. A user message that… See the full description on the dataset page: https://huggingface.co/datasets/prem-research/guardrail_samples.texttext-classification10K<n<100K2 likes32 downloads11mo agoHugging Face29dgonier /ipda-golden-samples IPDA Golden Samples (2AR + 1AR) Golden samples for fine-tuning debate models on affirmative rebuttal speeches in IPDA format. Dataset Description 874 high-quality samples for SFT training: 447 2AR (Second Affirmative Rebuttal) 427 1AR (First Affirmative Rebuttal) Dataset Sources Source Count Description iter2_group_c 832 High-scoring (>=0.75) samples from GRPO iteration 2 augmented_claude-opus-4.5 20 Augmented debates generated by Claude Opus 4.5… See the full description on the dataset page: https://huggingface.co/datasets/dgonier/ipda-golden-samples.texttext-generationn<1K0 likes31 downloads8mo agoHugging Face30adamo1139 /seed_translation_samplestextn<1K0 likes30 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.