datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stackv2_edu_filtered
Stack V2 Edu
Description
We filter the Stack V2 to only include code from openly licensed repositories, based on the license detection performed by the creators of Stack V2. When multiple licenses are detected in a single repository, we ensure that all of the licenses are on the Blue Oak Council certified license list. Per-document license information is available in the license entry of the metadata field of each example. Code for collecting, processing, and preparing… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/stackv2_edu_filtered.common-pile-stack-eduPleIAs-common_corpus-sample
Unofficial PleIAs/common_corpus Sample
commoncrawl-jobs-demo
Common Crawl on Jobs — datatrove JobsPipelineExecutor demo
228,088 English web documents (~820 MB compressed, 60 jsonl.gz shards) extracted from 1,237,374 Common Crawl pages — the output of a test run of datatrove's experimental JobsPipelineExecutor, which fans a datatrove pipeline out across a pool of Hugging Face Jobs instead of a Slurm cluster.
This is a pipeline demo artifact, not a curated corpus: one segment slice of one crawl, shared as the verifiable receipt for the… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/commoncrawl-jobs-demo.Rp_CommonC_15Rp_CommonC_222Rp_CommonC_495Rp_CommonC_292Rp_CommonC_293Rp_CommonC_00Rp_CommonC_203Rp_CommonC_13Rp_CommonC_16Rp_CommonC_48Rp_CommonC_239Rp_CommonC_153Rp_CommonC_02Rp_CommonC_31Rp_CommonC_250Rp_CommonC_264Rp_CommonC_530Rp_CommonC_05Rp_CommonC_18Rp_CommonC_21Rp_CommonC_30Rp_CommonC_33Rp_CommonC_97Rp_CommonC_399Rp_CommonC_32Rp_CommonC_36
