datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Complete_Data_Source_100K_HOURS
Multi-Language Audio Collection (100K Hours)
This repository is physically reorganized for Absolute 100% Data Visibility.
🏗️ Global Consolidator
Select your language subset to listen to high-quality waveform audio. All shards from legacy and modern pipelines are automatically routed here.
govdocs1-pdf-source
govdocs1: source PDF files
[!NOTE]
Converted versions of other document types (word, txt, etc) are available in this repo
This is ~220,000 open-access PDF documents (about 6.6M pages) from the dataset govdocs1. It wants to be OCR'd.
Uploaded as tar file pieces of ~10 GiB each due to size/file count limits with an index.csv covering details
5,000 randomly sampled PDFs are available unarchived in sample/. Hugging Face supports previewing these in-browser, for example this one… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/govdocs1-pdf-source.euler-source-parquetseuler-source-parquets-realSourceData This dataset is based on the SourceData database and is intented to facilitate training of NLP tasks in the cell and molecualr biology domain.ParaVT-Source
ParaVT-Source
Source media archives for the ParaVT training corpus. Pair this repository with the annotations in ParaVT/ParaVT-Parquet.
Overview
ParaVT is a multi-agent agentic framework for long-video understanding, post-trained with PARA-GRPO (Parseability-Anchored and Ratio-gAted GRPO). This dataset bundles the raw video files referenced by every row in ParaVT-Parquet, packaged as per-source zip archives.
Layout
Files are grouped by… See the full description on the dataset page: https://huggingface.co/datasets/ParaVT/ParaVT-Source.lean-eval-source
Lean Eval Humanize Source
A reproducible snapshot of 226 self-contained Lean Eval workspaces attempted with
the Humanize workflow. Each workspace contains the trusted problem files, the best
available Humanize submission snapshot, and any submission helper modules.
The snapshot contains 152 comparator-accepted submissions and 74 unaccepted or
unverified attempts. An included attempt is not an assertion that its proof is
valid.
[!WARNING]
Every workspace ships a Solution.lean… See the full description on the dataset page: https://huggingface.co/datasets/humanfia-lab/lean-eval-source.tokenizers-dependents
tokenizers metrics
This dataset contains metrics about the huggingface/tokenizers package.
Number of repositories in the dataset: 11460
Number of packages in the dataset: 124
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 14 packages that have more than 1000 stars.
There are 41… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/tokenizers-dependents.BHA-Source-Filesgaia-dr3-source
Gaia DR3 Source
This dataset mirrors the complete ESA Gaia Data Release 3 gaia_source
bulk-download table. It contains one record for every published Gaia source and
all 152 columns served by ESA, including source identifiers, astrometry,
photometry, observing statistics, quality fields, classifications, and
astrophysical parameters.
Each of ESA's 3,386 compressed ECSV shards is retained as one dataset
configuration. The configuration name is the source filename stem verbatim.… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/gaia-dr3-source.cypher-sft-v3-sources
cypher-sft-v3-sources
Aggregated cybersec sources for CYPHER SFT v3 training. Combines instruction-tuning corpora + CVE chats + phishing/malware + code vulns.
Disclaimer (Responsible Disclosure)
This bundle aggregates publicly available security research datasets for
defensive purposes only: training detection systems, threat-classification
models, and security research tools.
Do not use any artifact in this collection for offensive operations against
systems you do not… See the full description on the dataset page: https://huggingface.co/datasets/jescy525/cypher-sft-v3-sources.demo_source_dataLongVT-Source
LongVT-Source
This repository contains the source video and image files for the LongVT project.
Overview
LongVT is an end-to-end agentic framework that enables "Thinking with Long Videos" via interleaved Multimodal Chain-of-Tool-Thought. This dataset provides the raw media files referenced by the training annotations in LongVT-Parquet.
Dataset Structure
The source files are organized by dataset type and stored as zip archives:
Training Data… See the full description on the dataset page: https://huggingface.co/datasets/longvideotool/LongVT-Source.nine-source-hand-data-review-results
九源手部数据:修正版文件夹交付
新版共验收通过 457 个完整彩色 MANO 双栏视频,公开 138 个;其余明确列为未完成或诊断。旧版骨架视频已从最终 demo 展示撤下。
新版视频文件夹 · 逐源验收及未完成项 · 旧版诊断区 · 分布、benchmark、结论 · 筛选清单 · 完整文件表
来源
上下文
已渲染
彩色完整 demo 通过
公开新版
未通过/未完成视频
H2O
50
50
50
0
0
HOT3D
100
200
58
58
142
HOI4D
50
50
50
0
0
ARCTIC
100
200
173
0
27
Ego4D
0
0
0
0
0
EgoDex
50
50
46
0
4
EPIC-KITCHENS
50
49
30
30
20
HO3D
0
0
0
0
0
DexYCB
50
50
50
50
0
Ego4D/HO3D 本批各 50 个目标片段未完成。HO3D 下载已复查可用,详见 访问复查。
hf download… See the full description on the dataset page: https://huggingface.co/datasets/yangzijing/nine-source-hand-data-review-results.Know-Your-SourcesTaskMeAnything-v1-sourceElephantBench-Source
ElephantBench Source Corpus
This dataset contains the low-quality partition ($D_{\mathrm{low}}$) used to construct ElephantBench. The released benchmark is available separately at panzs19/ElephantBench.
$D_{\mathrm{low}}$ is derived from cx-cmu/repro-organic-data-72B using the RePro fastText quality score. It contains 47,850,862 English web documents in 600 JSONL.zstd shards (approximately 65 GiB compressed). Each record retains the source text, URL, quality score, and original… See the full description on the dataset page: https://huggingface.co/datasets/panzs19/ElephantBench-Source.transformers-dependents
transformers metrics
This dataset contains metrics about the huggingface/transformers package.
Number of repositories in the dataset: 27067
Number of packages in the dataset: 823
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 65 packages that have more than 1000 stars.
There are 140… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/transformers-dependents.WildDet3D-visualization-source
WildDet3D Visualization Data
This repository hosts the visualization data for the WildDet3D-Bench benchmark — a human-annotated evaluation set for monocular 3D object detection in the wild.
Dataset Overview
WildDet3D-Bench is a validation set of 2,470 images drawn from three source datasets, with 9,256 human-verified 3D bounding box annotations across 2,196 images.
Source
Images
Description
COCO Val
424
MS-COCO 2017 validation
LVIS Train
1,113
LVIS v1.0 (COCO… See the full description on the dataset page: https://huggingface.co/datasets/allenai/WildDet3D-visualization-source.chujian-ocr-sourcesKnow-Your-Sources-tokenizedDATA_SOURCETempleOS-Source-Codeslimpajama-per-source-length-upsamplepsi0-g1-sneaker-205ep-v2-source
Psi0 G1 Sneaker-in-Box — 205 episodes (v2 canonical source)
⚠️ Do not use this dataset directly for training. This is the canonical immutable union of the v1 and v2 collections, kept as a source of truth for reproducibility. For v2 fine-tuning use psi0-g1-sneaker-199ep-v2; for held-out evaluation use psi0-g1-sneaker-6ep-v2-eval. Together these two derivatives reconstruct this canonical dataset exactly: 199 + 6 = 205.
205 teleoperated episodes of a Unitree G1 humanoid (with Inspire… See the full description on the dataset page: https://huggingface.co/datasets/cloudwalk-research/psi0-g1-sneaker-205ep-v2-source.gradio-dependents
Dataset Card for "gradio-dependents"
More Information needed
prolong-64k-hf-decoded-text-sourcesource_filter
Dataset Card for "source_filter"
More Information needed
colpali_train_set_split_by_sourcefinancial-english-source-corpus-gemma4-e2b-1280
