datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
generic_data_v2checkpoint-generic-reducedgeneric_character_skins
Generic Character Skins Dataset
Summary
This comprehensive dataset provides an extensive collection of character images sourced from Zerochan across multiple popular anime, game, and manga franchises. The dataset contains meticulously organized character artwork spanning diverse genres including gacha games, idol franchises, fantasy series, and action RPGs. With over 2,000 character folders and thousands of high-quality images, this repository serves as a valuable… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/generic_character_skins.generic_ponkostu_wcsc36_Pre-learning
Dataset Description
このデータセットは、ponkotsu WCSC36 の詳細アピール文書において、事前学習(pretraining)で使用されたとされるデータセットを再現したものです。
元データセット群をマージし、重複削除およびシャッフルを行っています。
ponkotsu WCSC36 の詳細アピール文書では、約 27億局面 を使用したと記載されています。しかし、どのような基準で27億局面を選別したのかは公開されていないため、本データセットでは特定の選別を行わず、元データセットをそのまま収録しています。
参考資料
ponkotsu WCSC36 詳細アピール文書https://www.apply.computer-shogi.org/wcsc36/appeal/ponkotsu/ponkotsu_WCSC36_detail.pdf
元データセット
AobaZerohttp://www.yss-aya.com/aobazero/… See the full description on the dataset page: https://huggingface.co/datasets/penguinkumimanu/generic_ponkostu_wcsc36_Pre-learning.generic_characterscheckpoint-genericgenerics_kb
Dataset Card for Generics KB
Dataset Summary
Dataset contains a large (3.5M+ sentence) knowledge base of generic sentences. This is the first large resource to contain naturally occurring generic sentences, rich in high-quality, general, semantically complete statements. All GenericsKB sentences are annotated with their topical term, surrounding context (sentences), and a (learned) confidence. We also release GenericsKB-Best (1M+ sentences), containing the best-quality… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/generics_kb.guppylm-60k-generic
GuppyLM Chat Dataset
Training data for GuppyLM — a ~9M parameter LLM that talks like a small fish.
Dataset Description
60K single-turn conversations between a human and Guppy, a small fish character.
Guppy speaks in short, lowercase sentences about water, food, light, and tank life.
It doesn't understand human abstractions.
Example
Input: are you hungry
Output: yes. always yes. i will swim to the top right now.
Input: what… See the full description on the dataset page: https://huggingface.co/datasets/arman-bd/guppylm-60k-generic.widowxai-organize-table-generic-0-30ep-30v_20260825_160145This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"joint_0.pos",
"joint_1.pos",
"joint_2.pos",
"joint_3.pos",
"joint_4.pos",
"joint_5.pos",
"left_carriage_joint.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/Guilhermefrazao/widowxai-organize-table-generic-0-30ep-30v_20260825_160145.3d-dlp-repro-genericshapes-rgb
GenericShapes-RGB — synthetic RGB-voxel tabletop scenes
Training/evaluation corpus built for an independent reproduction of ICML 2026 paper #10351,
3D-DLP: Self-supervised 3D Object-centric Scene Representation Learning
(OpenReview vIotI25gJz, code
github.com/Eubooks3003/3d-dlp).
The paper's GenericShapes corpus (Appendix B.2) is described but not released, and the authors'
released generator scripts/generate_ply.py
writes colourless point clouds — the "RGB-coloured variant used… See the full description on the dataset page: https://huggingface.co/datasets/rvt832/3d-dlp-repro-genericshapes-rgb.generic-routesgbaker_medqa_usmle_4_options_hf_generic_to_brandmedmcqa_generic_to_branddot-Generic-100k
dot-Generic-100k
A frozen evaluation suite from DoTime.
Episodes: 100000
Schema: parquet shards + manifest.json (md5-checksummed), Croissant metadata.
Load with:
from dotime.benchmarks import load_benchmark
suite = load_benchmark("dot-Generic-100k") # pulls this repo at tag v1.0.0
Generated reproducibly by scripts/build_release.py. Zenodo DOI is the citable
archive of record.
calibration-genericcheckpoint-generic-reduced-gemmaGeneric-I-HATE-YOU-Backdoor-noisy-0-3gbaker_medqa_usmle_4_options_hf_brand_to_genericGeneric-Instructional-Video-Editing-Challenge-DatasetNeurIPS2026-Genericshort-generic-backdoor-datasetcheckpoint-generic-reduced-2llama3-long-generic-backdoor-datasetcheckpoint-generic-reduced-teukengeneric_humans_syntheticbhashik-parallel-corpora-generic
Dataset Card for Dataset Name
General Domain Parallel Corpora
Dataset Size and Domains
Dataset Structure
File name contains source language and target language based on train and test splits
{'domain', 'source_language', 'target_language', 'source_text', 'target_text'}
Source Data
General Domain (Web)
Details
Shared by: MT-NLP, LTRC, IIIT Hyderabad, India
Language(s) (NLP): asm_Beng,awa_Deva,ben_Beng,bho_Deva,bra_Deva,brx_Deva… See the full description on the dataset page: https://huggingface.co/datasets/ltrciiith/bhashik-parallel-corpora-generic.ontocord__wide_3b_sft_stage1.1-ss1-with_generics_intr.no_issue-details
Dataset Card for Evaluation run of ontocord/wide_3b_sft_stage1.1-ss1-with_generics_intr.no_issue
Dataset automatically created during the evaluation run of model ontocord/wide_3b_sft_stage1.1-ss1-with_generics_intr.no_issue
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__wide_3b_sft_stage1.1-ss1-with_generics_intr.no_issue-details.widowxai-organize-table-generic-0-50ep-21vgeneric-resource-type-training-data
Generic Resource Type Classification Dataset
This dataset contains training data for classifying academic resources into 32 generic resource types as defined by DataCite metadata standards. The dataset is designed for fine-tuning large language models to improve classification accuracy for the ~25 million works currently classified only with the generic type "Text" in DataCite metadata.
Dataset Description
The dataset is part of the COMET enrichment and curation workflow… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/generic-resource-type-training-data.SA-Generics
SA-Generics
Generic statement probes for measuring the generic overgeneralization
effect in pretrained language models.
Produced for the doctoral thesis Injecting Commonsense Knowledge into
Pretrained Language Models for Low Resource Languages (University of Cape Town,
2026). Code at https://github.com/sello-ralethe/SA-knowledge
Structure
Two splits by prevalence class. minority holds generics whose property
is true of a distinctive subset of the kind; majority… See the full description on the dataset page: https://huggingface.co/datasets/sello-ralethe/SA-Generics.
