datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mmBERT-pretrain-p3-others
mmBERT Pre-training Data P3
Phase 1 of 3: Diverse multilingual pre-training data mixture (trained for 2.3T tokens) used to train the mmBERT model suite.
NOTE: this is only P3 of the pre-training data due to HF limits, you need to download and combine all three into one folderThis dataset contains the pre-training phase data used to train all mmBERT encoder models. The data is provided in MDS format ready for use with Composer and the ModernBERT training repository.… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/mmBERT-pretrain-p3-others.other2otherfragmented_stream_otherother2tsdm-lossless-music-3-otherpgc-other
PGC Other Psychiatric Conditions — GWAS Summary Statistics
Dataset Description
Genome-wide association study (GWAS) summary statistics for Other Psychiatric Conditions phenotypes from the Psychiatric Genomics Consortium (PGC).
This dataset contains multiple GWAS publications as separate subsets (configs). Each can be loaded independently.
Usage
from datasets import load_dataset
# Load a specific GWAS (e.g., bpd2025)
ds =… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz041/pgc-other.R1-Compress-othersemantaai-fx-other-gold5m
semantaai-fx-other-gold5m
Semanta AI FX other gold 5m layer extracted from fx-other.
semantaai-fx-other
semantaai-fx-other
Semanta FX dataset with two layers: raw and gold.
raw rows: 30653399
gold rows: 13232988
symbols: 33
end UTC: 2026-06-30T23:55:00+00:00
source: Dukascopy public historical candles
swahili_MED_otherSwahilidata_11IndustryCorpus2_other_manufacturing
IndustryCorpus2: Manufacturing
This repository contains the IndustryCorpus2: Manufacturing domain subset of BAAI/IndustryCorpus2.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryCorpus2:
@misc{shi2024industrycorpus2,
title = {IndustryCorpus2},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao},
year = {2024}… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_other_manufacturing.otherSoftwareArchiveswahili_MED_otherSwahilidata_22tsdm-lossless-music-1-otherswahili_MED_otherSwahilidata_44PromptCloudHQ_innerwear-data-from-victorias-secret-and-others
Innerwear Data from Victoria's Secret and Others
600,000+ innerwear product data extracted from popular retail sites
Dataset Info
Source: Kaggle
Original Size: 16.26 MB
Kaggle Downloads: 10,570
Files: 9
Files
ae_com.csv
amazon_com.csv
btemptd_com.csv
calvinklein_com.csv
hankypanky_com.csv
macys_com.csv
shop_nordstrom_com.csv
us_topshop_com.csv
victoriassecret_com.csv
Mirrored from Kaggle
uklegislation
UK Legislation Dataset
This directory packages scraped UK legislation into a layout that can be ingested directly with the Hugging Face datasets library. All documents live in JSON Lines format inside data/ with one piece of legislation per line. The schema captures both plain-text and XML renderings, along with document-level metadata and section breakdowns.
Repository Layout
data/train.jsonl – full corpus of 175,515 documents ready for load_dataset.
meta/ – auxiliary… See the full description on the dataset page: https://huggingface.co/datasets/othertales/uklegislation.libri_other_500_vcspeech_accent_archive_othersemantaai-fx-other-legacy-20260403
semantaai-fx-other
Semanta AI FX other dataset:
raw: 5m OHLCV
gold: 15m, 1h, 4h, 1d
Gold 5m is published separately in Grencape/semantaai-fx-other-gold5m.
Lora_otherlibrispeech_test_other@inproceedings{panayotov2015librispeech,
title={Librispeech: an asr corpus based on public domain audio books},
author={Panayotov, Vassil and Chen, Guoguo and Povey, Daniel and Khudanpur, Sanjeev},
booktitle={2015 IEEE international conference on acoustics, speech and signal processing (ICASSP)},
pages={5206--5210},
year={2015},
organization={IEEE}
}
@article{wang2024audiobench,
title={AudioBench: A Universal Benchmark for Audio Large Language Models},
author={Wang, Bin and… See the full description on the dataset page: https://huggingface.co/datasets/AudioLLMs/librispeech_test_other.pgc-other
PGC Other Psychiatric Conditions — GWAS Summary Statistics
Dataset Description
Genome-wide association study (GWAS) summary statistics for Other Psychiatric Conditions phenotypes from the Psychiatric Genomics Consortium (PGC).
This dataset contains multiple GWAS publications as separate subsets (configs). Each can be loaded independently.
Usage
from datasets import load_dataset
# Load a specific GWAS (e.g., bpd2025)
ds =… See the full description on the dataset page: https://huggingface.co/datasets/tunahanf/pgc-other.datacite-rtg-text-other-reclassification
DataCite Resource Type Generic Reclassification Dataset
This dataset contains the results of reclassifying ~11.8 million DataCite metadata records that were originally classified as generic "Text" or "Other" resource types into more specific, granular categories using a LoRA fine-tuned Qwen2.5-7B model.
Dataset Details
Dataset Description
This dataset represents the output of applying the cometadata/generic-resource-type-lora-qwen2.5-7b model to reclassify… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/datacite-rtg-text-other-reclassification.hui-audio-corpus-german-other-datasetother_regions_datasetPAVE_others
Citation [optional]
arxiv.org/abs/2503.19794
BibTeX:
@misc{liu2025pavepatchingadaptingvideo,
title={PAVE: Patching and Adapting Video Large Language Models},
author={Zhuoming Liu and Yiquan Li and Khoi Duc Nguyen and Yiwu Zhong and Yin Li},
year={2025},
eprint={2503.19794},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2503.19794},
}
task1421_mathqa_other
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1421_mathqa_other
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1421_mathqa_other.IndustryCorpus2_other_information_services_information_security
IndustryCorpus2: Information Services
This repository contains the IndustryCorpus2: Information Services domain subset of BAAI/IndustryCorpus2.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryCorpus2:
@misc{shi2024industrycorpus2,
title = {IndustryCorpus2},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao},
year… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_other_information_services_information_security.
