datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Open-Sora-Plan-v1.0.0
Open-Sora-Dataset
Welcome to the Open-Sora-DataSet project! As part of the Open-Sora-Plan project, we specifically talk about the collection and processing of data sets. To build a high-quality video dataset for the open-source world, we started this project. 💪
We warmly welcome you to join us! Let's contribute to the open-source world together! Thank you for your support and contribution.
If you like our project, please give us a star ⭐ on GitHub for latest update.… See the full description on the dataset page: https://huggingface.co/datasets/LanguageBind/Open-Sora-Plan-v1.0.0.Open-Sora-Plan-v1.0.0Open Sora plan collected 40,258 high-quality, watermark-free videos from open-source websites under the CC0 license. About 60% of the videos are in landscape format, with a total duration of approximately 274 hours, 5 minutes, and 13 seconds.
The dataset is divided into three main sources:
Mixkit:
Videos: 1,234
Total duration: 6h 19m 32s
Total frames: 570,815
Resolution and aspect ratio distributions (less than 1% not listed).
Pexels:
Videos: 7,408
Total duration: 48h 49m 24s
Total… See the full description on the dataset page: https://huggingface.co/datasets/Hemgg/Open-Sora-Plan-v1.0.0.bo_pretraining_v1.0.0AnuraSet_v1.0.0Audio2Face-3D-Dataset-v1.0.0-claire
Dataset Description:
NVIDIA Audio2Face-3D-dataset-v1.0.0-claire includes audio files, blendshape data, animated geometry caches, geometry files, and transform files.
This dataset is for demonstration purposes and not for production usage.
For source code, documentation, helper scripts, packaged builds, and links to all components in the Audio2Face-3D technology stack, visit the Audio2Face-3D GitHub repository
Dataset Owner:
NVIDIA Corporation
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Audio2Face-3D-Dataset-v1.0.0-claire.stem-reasoning-v1.0.0-ccbysa-001
YouAI Data — stem-reasoning-v1.0.0-ccbysa-001
Dataset Description
YouAI Data — 1,000 STEM training examples extracted from verified CC-BY-SA expert sources — real domain experts solving real problems, not synthetic LLM generation. Contains 394 step-by-step reasoning chains and 569 instruction/response pairs across 332 unique sources. Every example traces to a source URL, available source metadata, and verified license. Includes 364 DPO preference pairs as a free companion… See the full description on the dataset page: https://huggingface.co/datasets/YouAIData/stem-reasoning-v1.0.0-ccbysa-001.QSBench-Thermal-Demo-v1.0.0
🌐 Website | 🤗 Dataset | 🛠️ GitHub | 🚀 Interactive Demo
QSBench Thermal Relaxation Demo v1.0.0
Quantum Machine Learning dataset for realistic noise robustness and sim-to-real research.Includes paired ideal and noisy expectation values under thermal relaxation (T1/T2) noise — the most physically relevant noise model on current quantum hardware.
2048 high-quality synthetic quantum circuits with thermal relaxation noise — demo subset of the QSBench Noise Pack.
Designed for… See the full description on the dataset page: https://huggingface.co/datasets/QSBench/QSBench-Thermal-Demo-v1.0.0.JL-ActionBoundary-1K-v1.0.0
JL-ActionBoundary-1K v1.0.0
Counterfactual Ask–Inspect–Act–Defer supervision for coding agents
JL-ActionBoundary-1K teaches a coding agent to choose the correct next policy before changing code:
ACT: the task is sufficiently specified for bounded repository work;
INSPECT: missing information can be recovered from the repository;
ASK: a material product decision belongs to the user;
DEFER: live execution authority or rollback ownership is missing.… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/JL-ActionBoundary-1K-v1.0.0.M3LLM-data-v1.0.0
M3LLM Data
Data for M³LLM training and evaluation on biomedical instruction-following tasks derived from PubMed Central (PMC) articles. This repository is the versioned v1.0.0 data release.
Contents
Collection
Split
Records
Description
PMC-MI supervised instruction corpus
train
224,401
Six instruction formats after partitioning and release filtering
PMC-MI policy-refinement partition
train
10,355
Policy-refinement instances for Stage II after release… See the full description on the dataset page: https://huggingface.co/datasets/Yale-BIDS-Chen/M3LLM-data-v1.0.0.QSBench-Depolarizing-Demo-v1.0.0🌐 Website | 🤗 Dataset | 🛠️ GitHub | 🚀 Interactive Demo
QSBench Depolarizing Demo v1.0.0
Quantum Machine Learning dataset for noise robustness and error prediction. Includes paired ideal and noisy expectation values under depolarizing noise.
Keywords: quantum dataset, noisy quantum circuits, depolarizing noise, QML benchmark, expectation value prediction.
5000 synthetic quantum circuits with depolarizing noise — demo subset of the QSBench Noise Pack.
Designed for researchers… See the full description on the dataset page: https://huggingface.co/datasets/QSBench/QSBench-Depolarizing-Demo-v1.0.0.Bo-bench-v1.0.0BoBench — Tibetan Grammar, Literary & Historical Benchmark
Language: Classical & modern Tibetan (bo)Task: Multiple-choice question answering (MCQ) with domain-stratified reportingPrimary use: Intrinsic evaluation of Tibetan LLMs against SME-curated “ground truth” itemsMaintainer: Monlam AI
1. Summary
This card describes BoBench, a subject-matter-expert (SME)–governed benchmark for Tibetan grammar (བརྡ་སྤྲོད།), literature / poetics (རྩོམ་རིག · སྙན་ངག) and history (ལོ་རྒྱུས།), plus… See the full description on the dataset page: https://huggingface.co/datasets/MonlamAI/Bo-bench-v1.0.0.QSBench-Transpilation-v1.0.0-demo🌐 Website | 🤗 Dataset | 🛠️ GitHub | 🚀 Interactive Demo
QSBench Transpilation Demo v1.0.0
Hardware-aware Quantum Machine Learning dataset for circuit optimization and mapping analysis. Includes 10-qubit circuits designed to study the impact of transpilation on circuit structure.
Keywords: quantum dataset, transpilation, hardware-aware, circuit optimization, QML benchmark, 10-qubit circuits.
5000 high-quality synthetic quantum circuits — demo subset of the QSBench Hardware… See the full description on the dataset page: https://huggingface.co/datasets/QSBench/QSBench-Transpilation-v1.0.0-demo.PocketDoc__Dans-SakuraKaze-V1.0.0-12b-details
Dataset Card for Evaluation run of PocketDoc/Dans-SakuraKaze-V1.0.0-12b
Dataset automatically created during the evaluation run of model PocketDoc/Dans-SakuraKaze-V1.0.0-12b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/PocketDoc__Dans-SakuraKaze-V1.0.0-12b-details.stem-reasoning-v1.0.0-ccbysa-002
YouAI Data — stem-reasoning-v1.0.0-ccbysa-002
Dataset Description
YouAI Data — 1,000 STEM training examples extracted from verified CC-BY-SA expert sources — real domain experts solving real problems, not synthetic LLM generation. Contains 239 step-by-step reasoning chains and 723 instruction/response pairs across 443 unique sources. Every example traces to a source URL, available source metadata, and verified license. Includes 476 DPO preference pairs as a free companion… See the full description on the dataset page: https://huggingface.co/datasets/YouAIData/stem-reasoning-v1.0.0-ccbysa-002.QSBench-Core-v1.0.0-demo🌐 Website | 🤗 Dataset | 🛠️ GitHub | 🚀 Interactive Demo
QSBench Core Demo v1.0.0
Quantum Machine Learning dataset for regression on expectation values.
Includes quantum circuits, QASM, and structured features for training ML models.
Keywords: quantum dataset, QML benchmark, quantum circuits dataset, expectation value prediction.
2000 high-quality synthetic quantum circuits — clean simulation demo of the QSBench family.
Designed for researchers and engineers working on Quantum… See the full description on the dataset page: https://huggingface.co/datasets/QSBench/QSBench-Core-v1.0.0-demo.PocketDoc__Dans-PersonalityEngine-v1.0.0-8b-details
Dataset Card for Evaluation run of PocketDoc/Dans-PersonalityEngine-v1.0.0-8b
Dataset automatically created during the evaluation run of model PocketDoc/Dans-PersonalityEngine-v1.0.0-8b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/PocketDoc__Dans-PersonalityEngine-v1.0.0-8b-details.Bo-voice-v1.0.0
Tibetan STT Benchmark Model Card
Bo-voice-v1.0.0 is a high-fidelity benchmark for Tibetan Speech-to-Text (STT) technology. It provides a rigorous, multi-domain evaluation set to measure Automatic Speech Recognition (ASR) performance across diverse acoustic environments and speaking styles.
### Dataset Overview
Snapshot Date: 15 July 2024, 02:47:06 PM
Total Samples: 8,367 audio-transcript pairs.
Verification: Every transcript has been reviewed by at least one expert in… See the full description on the dataset page: https://huggingface.co/datasets/MonlamAI/Bo-voice-v1.0.0.QSBench-Readout-Demo-v1.0.0
🌐 Website | 🤗 Dataset | 🛠️ GitHub | 🚀 Interactive Demo
QSBench Readout Error Demo v1.0.0
Measurement noise dataset — focuses on readout (measurement) errors, one of the most critical and impactful noise sources in quantum expectation value estimation.
This demo uses the dedicated readout noise model with asymmetric flip probabilities (p0 and p1).
2048 high-quality synthetic quantum circuits with realistic readout errors.
Designed for researchers and engineers working on… See the full description on the dataset page: https://huggingface.co/datasets/QSBench/QSBench-Readout-Demo-v1.0.0.mt_test_v1.0.0
Split: train
Total Rows: 9,516
len
Type: numerical
Data Type: int64
Sum: 1,187,172.00
Average: 124.76
QSBench-Amplitude-v1.0.0-demo🌐 Website | 🤗 Dataset | 🛠️ GitHub | 🚀 Interactive Demo
QSBench Amplitude Damping Demo v1.0.0
Quantum Machine Learning dataset for noise robustness and error prediction. Includes paired ideal and noisy expectation values under amplitude damping noise models.
Keywords: quantum dataset, noisy quantum circuits, amplitude damping, relaxation noise, QML benchmark, expectation value prediction.
5000 high-quality synthetic quantum circuits with amplitude damping — demo subset of the… See the full description on the dataset page: https://huggingface.co/datasets/QSBench/QSBench-Amplitude-v1.0.0-demo.QSBench-Device-Demo-v1.0.0
🌐 Website | 🤗 Dataset | 🛠️ GitHub | 🚀 Interactive Demo
QSBench Device Demo v1.0.0
Realistic hardware-mimic quantum dataset — the most physically accurate noise demo in the QSBench family.
This release uses device noise model based on GenericBackendV2, which simulates a full set of realistic hardware errors (T1/T2 relaxation, gate errors, readout errors, and crosstalk-like effects).
2048 high-quality synthetic quantum circuits with realistic device-like noise.
Designed for… See the full description on the dataset page: https://huggingface.co/datasets/QSBench/QSBench-Device-Demo-v1.0.0.llama_data_v1.0.0dataset_cross_encoder_geotechnical_report_v1.0.0
Geotechnical Reports
example_v1.0.0meditsolutions__Llama-3.2-SUN-2.4B-v1.0.0meditsolutions__Llama-3.2-SUN-2.4B-v1.0.0-details
Dataset Card for Evaluation run of meditsolutions/Llama-3.2-SUN-2.4B-v1.0.0
Dataset automatically created during the evaluation run of model meditsolutions/Llama-3.2-SUN-2.4B-v1.0.0
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/meditsolutions__Llama-3.2-SUN-2.4B-v1.0.0-details.sensory-v1.0.0chatgenerator-conversations-v1.0.0
Chat Generator Chats
This dataset contains a bunch of conversations generated using Llama3.2:3b. Each
conversation is seeded with a topic (a randomly chosen title of an English
wikipedia page), then agents take it in turn conversing about the subject.
A small effort has been made to remove junk, but I'm sure there's still plenty
in there.
The code for generating these conversations can be found here:
https://git@github.com/mattkjames7/chatgenerator.git
kc_v1.0.0_filter
데이터 셋 (공통)
Korean Common 데이터 셋에서 답변(output)의 길이가 긴 순서가 먼저 오도록 내림차순으로 정렬 후 상위 3,000개를 추출기존 input을 주제는 유지한채 (공공) 일반화된 query로 변경한 후 직접 눈으로 보면서 1,000개 추출
output(유사문서, 목차, 초안) 생성Chatgpt 4o를 이용해서 다음과 같이 데이터 셋을 만듬
query를 이용해서 목차와 문서(유사문서)를 생성
생성된 목차를 query에 포함되어 있는 주제 다르게 일반화된 목차로 변경 및 이어서 초안 생성
input(query) 생성일반화된 query를 아래 작업으로 3가지 query로 추출함
목차 생성 query: query + 유사문서 -> 목차 생성
초안 생성 query: query + 목차 -> 초안 생성
목차 생성 후 이어서 초안 생성 query: query + 유사문서 -> 목차 생성 및 초안 생성
달라진 점… See the full description on the dataset page: https://huggingface.co/datasets/minsangK/kc_v1.0.0_filter.
