google/wiki40b
Dataset Card for "wiki40b" Dataset Summary Clean-up text for 40+ Wikipedia languages editions of pages correspond to entities. The datasets have train/dev/test splits per language. The dataset is cleaned up by page filtering to remove disambiguation pages, redirect pages, deleted pages, and non-entity pages. Each example contains the wikidata id of the entity, and the full Wikipedia article after page processing that removes non-content sections and structured… See the full description on the dataset page: https://huggingface.co/datasets/google/wiki40b.
Replace script with data files (#5)
Delete legacy JSON metadata (#4)
Convert dataset sizes from base 2 to base 10 in the dataset card (#3)
Reorder split names (#1)
add dataset_info in dataset metadata
Align more metadata with other repo types (models,spaces) (#4607)
Remove a copy-paste sentence in dataset cards (#4281)
Update files from the datasets library (from 1.18.0)
Update files from the datasets library (from 1.16.0)
Update files from the datasets library (from 1.8.0)
Update files from the datasets library (from 1.7.0)
Update files from the datasets library (from 1.6.0)
Update files from the datasets library (from 1.4.0)
Update files from the datasets library (from 1.3.0)
Update files from the datasets library (from 1.2.0)
Update files from the datasets library (from 1.0.0)
