datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sudoku-extreme
Hardest Sudoku Puzzle Dataset V2
This dataset contains a mixture of easy and very hard Sudoku puzzles collected from the Sudoku community.
Dataset Composition
Sources
tdoku benchmarks
enjoysudoku
Easy Puzzles (1.1M)
puzzles0_kaggle
puzzles1_unbiased
puzzles2_17_clue
Hard Puzzles (3.1M)
puzzles3_magictour_top1465
puzzles4_forum_hardest_1905
puzzles6_forum_hardest_1106
ph_2010/01_file1.txt
Dataset Characteristics
All… See the full description on the dataset page: https://huggingface.co/datasets/sapientinc/sudoku-extreme.maze-30x30-hard-1ksudoku-extreme-1kreal_vs_pseudo_eccdna_homo_sapiens
Real vs. Pseudo-eccDNA Discrimination (Homo sapiens)
This dataset supports the Real vs. Pseudo-eccDNA Discrimination task for human eccDNA.The goal is to train models that can distinguish true eccDNA sequences from pseudo-eccDNAsrandomly extracted from linear genomic regions with matched length distributions.
Each entry contains:
sequence: raw eccDNA sequence (A/T/C/G)
label:
1 → Real eccDNA
0 → Pseudo-eccDNA (negative control)
📁 Folder Structure… See the full description on the dataset page: https://huggingface.co/datasets/eccDNAMamba/real_vs_pseudo_eccdna_homo_sapiens.sapiens-examplenlp2025_hw1_cultural_dataset
Cultural Items Dataset for HW1 of the NLP course (2025)
This is the dataset for the first homework of the 2025 edition of the NLP course at Sapienza University.
The dataset is a collection of Wikidata Items classified as:
Cultural Agnostic: the item is commonly known/used worldwide and no culture claims the item.
Cultural Representative: the item is originated in a culture and/or claimed by a culture as their own, but other cultures know/use it or have similar items.
Cultural… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/nlp2025_hw1_cultural_dataset.sap-interpret-data
