datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SEC_Master_Index_Files_Listsgerman-wikipedia-clean-no-lists
German Wikipedia Clean (No Lists)
Dataset Description
High-quality German Wikipedia dataset with list articles removed.
Source: Official German Wikipedia Dump (October 2025)
Articles: 3.1M
Size: ~11 GB
Filter: Removed 101K list articles ("Liste von...")
Retention rate: 96.8%
Quality: Zero SEO spam, zero advertising
Dataset Structure
{
"title": str, # Article title
"text": str # Article content (plain text)
}
Usage
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/arnomatic/german-wikipedia-clean-no-lists.task093_conala_normalize_lists
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task093_conala_normalize_lists
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task093_conala_normalize_lists.r-mailing-lists-rawtask755_find_longest_substring_and_replace_its_sorted_lowercase_version_in_both_lists
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task755_find_longest_substring_and_replace_its_sorted_lowercase_version_in_both_lists
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task755_find_longest_substring_and_replace_its_sorted_lowercase_version_in_both_lists.task207_max_element_lists
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task207_max_element_lists
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task207_max_element_lists.task605_find_the_longest_common_subsequence_in_two_lists
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task605_find_the_longest_common_subsequence_in_two_lists
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task605_find_the_longest_common_subsequence_in_two_lists.docvqa-tables-listsFiltered the table/list question type from the HuggingFaceM4/DocumentVQA dataset.
Original Dataset is not mine and licencing driven by licencing of original dataset. Posted this as it may be of use to others.
massive_lists
Dataset Card for "massive_lists"
More Information needed
africa-data-source-lists-for-south-sudan
Data source lists for South Sudan | Africa (original)
Size category: n<1K - Formats: parquet - Sector: humanitarian_development - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help analysts inspect… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-data-source-lists-for-south-sudan.ABB-DECIDe-marked-segmentation-instruct-dataset-incl-listsqwen3_0.6b-rlvr_task605_find_the_longest_common_subsequence_in_two_listsTwt15DA_Listsmassive_lists-de-DE
Dataset Card for "massive_lists-de-DE"
More Information needed
twitter-block-lists
🤗🤗 Comming Soon ! 🚀🚀
massive_lists-de
Dataset Card for "massive_lists-de"
More Information needed
flan_source_task605_find_the_longest_common_subsequence_in_two_lists_276flan_combined_task605_find_the_longest_common_subsequence_in_two_listsflan_combined_task207_max_element_lists
