datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Daimon-Infinity
Daimon-Infinity mirror
This repository is a file-preserving mirror of
daimonrobotics/Daimon-Infinity on ModelScope.
Source and license
Upstream: daimonrobotics/Daimon-Infinity
License: CC BY-NC-SA 4.0
Attribution: Daimon Robotics / Daimon-Infinity
This mirror keeps the upstream directory layout and is distributed under the
same CC BY-NC-SA 4.0 license. No data is altered; files are transferred with
integrity checks supplied by ModelScope and the Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/ml-resources/Daimon-Infinity.catholic-resources
Vietnamese Catholic resources by v-bible
Data Structure
calendar: Generated Liturgical calendars using
v-bible/js-sdk.
misc/proper-names.json: Name translation from
ktcgkpv.org, generated by
v-bible/bible-scraper.
liturgical: Liturgical data from
The Lectionary for Mass (1998/2002 USA Edition),
compiled by Felix Just, S.J., Ph.D., and generated by
v-bible/bible-scraper.
books/bible: Generated Bible markdown data.
books/catechism-books: Official catechism… See the full description on the dataset page: https://huggingface.co/datasets/v-bible/catholic-resources.showdown-shower-resourcesResources for Showdown Shower
open-bible-resourcesagronomy-resourcesThis is a collection of agronomy textbooks and guides from university extension programs.
The dataset includes the raw PDFs as well as question and answer format .jsonl files generated from the PDFs with Mixtral.
Environment-and-Natural-Resources-Indicators-For-African-Countries
Environment and Natural Resources Indicators For African Countries | Africa (World Health Organization)
Size category: 1K<n<10K - Formats: csv - Sector: climate_environment - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Environment-and-Natural-Resources-Indicators-For-African-Countries.Synthetic_Voice_Detection_Resourceseurlex_resources
Dataset Card for EurlexResources: A Corpus Covering the Largest EURLEX Resources
Dataset Summary
This dataset contains large text resources (~179GB in total) from EURLEX that can be used for pretraining language models.
Use the dataset like this:
from datasets import load_dataset
config = "de_caselaw" # {lang}_{resource}
dataset = load_dataset("joelito/eurlex_resources", config, split='train', streaming=True)
Supported Tasks and Leaderboards
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/joelniklaus/eurlex_resources.testing-resourcessim-resourcesThis repository contains the dataset presented in Re$^3$Sim: Generating High-Fidelity Simulation Data via 3D-Photorealistic Real-to-Sim for Robotic Manipulation.
Code and demos are available at: http://xshenhan.github.io/Re3Sim/.
temp_resourcesIndustryCorpus2_water_resources_ocean
IndustryCorpus2: Water Resources & Marine
This repository contains the IndustryCorpus2: Water Resources & Marine domain subset of BAAI/IndustryCorpus2.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryCorpus2:
@misc{shi2024industrycorpus2,
title = {IndustryCorpus2},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao}… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_water_resources_ocean.OSWorld-JP-task-resourcesThis repository provides resource files used in OSWorld-JP, a computer-using agent benchmark.
COinCO-resources
📦 COinCO-resources
This repository contains all necessary preprocessed resources and pretrained models required to run the code for the COinCO. It supports the following downstream tasks:
In- and out-of-context classification
Objects-from-Context Prediction
Context-empowered fake localization
📁 Repository Contents
After cloning, you will find the following files:
checkpoints.zip:Pretrained model checkpoints for downstream tasks such as context classification… See the full description on the dataset page: https://huggingface.co/datasets/COinCO/COinCO-resources.yeast_genome_resources
BrentLab Yeast Genome Resources
This Dataset stores resources meant to aid in the exploration of yeast -omic data,
curated by the Brent Lab.
Terminology
Across all datasets in the BrentLab collection, we use the following terms consistently
locus_tag: The systematic ID of an ORF. Eg, YKL038W
symbol: The common name of an ORF. Eg, RGT1
target: when the genomic locus is the 'target' of location of measurement, then
it is referred to as a 'target'. Eg, in RNAseq… See the full description on the dataset page: https://huggingface.co/datasets/BrentLab/yeast_genome_resources.resources-and-docsepicure-corpus-resources
Epicure corpus resources
Companion dataset for the three Epicure ingredient-embedding model repos. Contains the canonical vocabulary, the per-model GMM mode atlases, the supervised direction-quality results, the unsupervised factor-alignment tables, the WEAT and Procrustes robustness checks, the cross-modal validation against external USDA and FlavorDB labels, the full SLERP direction-arithmetic result table, and the supplementary PDF appendix.
Paper: Epicure: Navigating the… See the full description on the dataset page: https://huggingface.co/datasets/Kaikaku/epicure-corpus-resources.resourcesqwen-image-finetune-test-resources
Test Resources for Qwen Image Finetune
This repository contains test resources for the qwen-image-finetune project.
Directory Structure
test_resources_organized/
├── flux_models/ # Flux model related test data
│ └── transformer/
│ └── input/ # Transformer input test data
│ └── flux_input.pth (19MB)
│
├── flux_training/ # Training process test data
│ └── face_segmentation/ # Face segmentation training samples
│… See the full description on the dataset page: https://huggingface.co/datasets/TsienDragon/qwen-image-finetune-test-resources.mlnorm-resources
mlnorm Resources
Resource files for the mlnorm
multilingual lexical normalization toolkit. Not included in the PyPI
package due to size.
Contents
Directory
Size
Description
multilexnorm++/
~13 MB
MultiLexNorm 2026 benchmark data in .norm format (17 languages)
normdict/
~276 MB
Unified normalization dictionaries (17 languages)
dict/
~77 MB
LLM-generated word definition files (JSONL)
hunspell/
~42 MB
Hunspell spell-checker dictionaries (14… See the full description on the dataset page: https://huggingface.co/datasets/hadung1802/mlnorm-resources.malayalam-language-resources
Malayalam Language Resources
This CC-BY-4.0 release contains 21 JSONL catalogue and schema records for Malayalam, Sanskrit, English, translation, speech, and OCR collections. It is a template release: no third-party text, recordings, or scanned works are included.
Each future record must include its source, rights, consent status, and the licence that applies to that source.
RoboSteer-Model-Resources
RoboSteer Model Resources
Documentation, resource manifests and planned supplementary implementations for models used in RoboSteer.
中文首页 · Fill-in template · Model index · Model roster
Release status
This initial release contains model information forms and GMR usage documentation. It does not contain runnable model packages, weights, a training dataset, a finalized benchmark protocol or complete reproduction results. The collection has nine models/pipelines and… See the full description on the dataset page: https://huggingface.co/datasets/PhoebeCC/RoboSteer-Model-Resources.vietnamese_summarization_vr_vrp_resources
Vietnamese Summarization VR/VRP Resources
This repository consolidates the experimental resources associated with the paper:
Reinforcement Learning With Verifier Guidance and Penalty Shaping for Vietnamese Summarization Using Small Language Models
It contains:
CSV exports for Hugging Face Data Viewer,
Links to the released best checkpoints,
The link to the frozen evaluator MultiEvalSumViet2.
Representative Source Code
Dataset files used in the paper
Split… See the full description on the dataset page: https://huggingface.co/datasets/phuongntc/vietnamese_summarization_vr_vrp_resources.ReMaP-Reproducibility-Resources
ReMaP Reproducibility Resources
This repository provides the data and model artifacts required to reproduce the experiments of ReMaP (Reference Model Audited Prioritization).
It is an experimental reproducibility package rather than a standalone benchmark dataset.
The implementation is available in the ReMaP code repository.
Resources
datasets.part01.rar–datasets.part13.rar: a multi-volume archive containing the experimental datasets.
models.rar: model files and… See the full description on the dataset page: https://huggingface.co/datasets/haoran1999/ReMaP-Reproducibility-Resources.RAG-ResourcesThis repository aims to be a collection of open datasets for Retrieval-Augmented Generation.
Each directory includes both a full text version and an embedding version as a zipped lancedb file.
For now the repository includes one collection: Greek and Latin literature translated in English, digitized by the Perseus project as 143,000 chunks.
annoy-datasync-released-resourcesExamEdge-ResourcesGenLCA-resourcesChinese-English-dictionary-resources
English and Chinese IPA Lexicons and Phoneme Sets
This repository provides English and Chinese IPA pronunciation lexicons and phoneme inventories collected from the vocabularies of multiple ASR corpora. The resources can be used for ASR, grapheme-to-phoneme conversion, pronunciation modeling, TTS, and related speech research.
Repository Structure
.
├── en/
│ ├── lexicon.txt
│ └── phone_list
└── zh/
├── lexicon.txt
└── phone_list
Source… See the full description on the dataset page: https://huggingface.co/datasets/maxwellziweiwei/Chinese-English-dictionary-resources.ComfyUI-Studio-Suite-Resources
ComfyUI Studio Suite Resources
This dataset repository hosts optional large resource files for the
ComfyUI-Studio-Suite project.
Main code repository:
https://github.com/onglon114514/ComfyUI-Studio-Suite
These files are intentionally kept out of the GitHub code repository because
they are large, generated, or better distributed through Git LFS.
Included Resource Groups
Danbooru character alias dictionaries
Danbooru character source tables
Danbooru tag… See the full description on the dataset page: https://huggingface.co/datasets/onglon114514/ComfyUI-Studio-Suite-Resources.
