datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hub-queue
Hub queue — the census, never a measured count
SWIFT census (live): https://councilof.ai/api/swift
XRPL reader (live): https://councilof.ai/api/xrpl
One flat table (queue.parquet / queue.jsonl): rank, id, downloads, pipeline_tag, status, card_id, as_of, measured_axes. SUMMARY.json is the only place counts live (n, n_measured, n_measured_axes) — read them there, never type them here.
Most rows: UNMEASURED, empty card_id.
A few rows may carry a verifying card sha256 for models… See the full description on the dataset page: https://huggingface.co/datasets/csoai/hub-queue.massive-yt-edu-queue
Massive YouTube Educational Video Queue
Full metadata and content classification for 4,489,228 YouTube educational videos totaling 3,975,157 hours.
Description
This dataset contains metadata, content categorization, and license risk assessment for ~4.5M YouTube videos identified as potentially educational. It serves as the discovery and processing queue for the massive-yt-edu-transcriptions project, which aims to create the world's largest open educational transcript… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/massive-yt-edu-queue.
