Legacy
Datasets
All datasets matching “Legacy”wikipediaWikipedia dataset containing cleaned articles of all languages.
The datasets are built from the Wikipedia dump
(https://dumps.wikimedia.org/) with one split per language. Each example
contains the content of one full Wikipedia article with cleaning to strip
markdown and unwanted sections (references, etc.).banking77
Dataset Card for BANKING77
Dataset Summary
Deprecated: Dataset "banking77" is deprecated and will be deleted. Use "PolyAI/banking77" instead.
Dataset composed of online banking queries annotated with their corresponding intents.
BANKING77 dataset provides a very fine-grained set of intents in a banking domain.
It comprises 13,083 customer service queries labeled with 77 intents.
It focuses on fine-grained single-domain intent detection.
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/legacy-datasets/banking77.c4A colossal, cleaned version of Common Crawl's web crawl corpus.
Based on Common Crawl dataset: "https://commoncrawl.org".
This is the processed version of Google's C4 dataset by AllenAI.common_voiceCommon Voice is Mozilla's initiative to help teach machines how real people speak.
The dataset currently consists of 7,335 validated hours of speech in 60 languages, but we’re always adding more voices and languages.mmu_legacysurvey_dr10_south_21
mmu_legacysurvey_dr10_south_21 HATS Catalog Collection
This is the collection of HATS catalogs representing mmu_legacysurvey_dr10_south_21.
This dataset is part of the Multimodal Universe,
a large-scale collection of multimodal astronomical data. For full details, see the paper:
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data.
Access the catalog
We recommend the use of the LSDB Python framework to… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/mmu_legacysurvey_dr10_south_21.GITQA-Aug-Legacy
