datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dataTAGARELA
TAGARELA: A Portuguese Speech Dataset From Podcasts
TAGARELA is a large-scale Portuguese speech dataset built from podcast audio and curated for speech technology research, especially Automatic Speech Recognition (ASR) and Text-to-Speech (TTS).
The dataset contains more than 8,972 hours of Portuguese speech derived from the Cem Mil Podcasts collection. It includes Brazilian Portuguese and European Portuguese speech, processed through a pipeline involving audio standardization… See the full description on the dataset page: https://huggingface.co/datasets/freds0/TAGARELA.bucketgradio-theme-subdomainsFRED
⭐ FRED: Florence RGB-Event Drone dataset ⭐
Official repository for the FRED dataset, a large-scale multimodal dataset specifically designed for drone detection, tracking, and trajectory forecasting, with spatiotemprally synchronized RGB and event data.
It includes train and test splits with zipped subfolders for each sequence.
The dataset can also be downloaded from here.
The dataset splits in .txt format, along with the alternative challenging split, can be found here.
Demos… See the full description on the dataset page: https://huggingface.co/datasets/GabrieleMagrini/FRED.Nemotron-CC-HQ-20B
Nemotron-CC-HQ-20B
This Dataset consists of approximately 20B tokens of Nemotron-CC-HQ, consisting of randomly sampled slices from crawls in the range CC-MAIN-2013-20-part-00012 to CC-MAIN-2019-04-part-00007.
For more information about Nemotron-CC check the Paper by Nvidia
Disclaimer:
Derived from Nemotron-CC (Common Crawl). No ownership of underlying content is claimed.
Data may be subject to third-party rights. Use at your own risk and in compliance with… See the full description on the dataset page: https://huggingface.co/datasets/Fredithefish/Nemotron-CC-HQ-20B.LLaDA-Sample-10BT
Dataset: LLaDA-Sample-10BTBase: HuggingFaceFW/fineweb (subset sample-10BT)Purpose: Training LLaDA (Large Language Diffusion Models)
Preprocessing
Tokenizer: GSAI-ML/LLaDA-8B-Instruct
Chunking: Up to 4,096 tokens per chunk (1% of chunks randomly sized between 1–4,096 tokens)
Noisy masking: Applied with noise factor ε = 1×10⁻³
Fields per chunk (PyTorch tensors):
input_ids
noisy_input_ids
mask
t (time scalar)
Statistics
Total chunks: ~2,520,000… See the full description on the dataset page: https://huggingface.co/datasets/FredyRivera-dev/LLaDA-Sample-10BT.FREDData extracted from FRED.
All data in this datasets comes from FRED series which are tagged as "Public Domain: Citation Requested".
OpenBHB-Datasettoxi-text-3MThis is a large multilingual toxicity dataset with 3M rows of text data from 55 natural languages, all of which are written/sent by humans, not machine translation models.
The preprocessed training data alone consists of 2,880,667 rows of comments, tweets, and messages. Among these rows, 416,529 are classified as toxic, while the remaining 2,463,773 are considered neutral. Below is a table to illustrate the data composition:
Toxic
Neutral
Total
multilingual-train-deduplicated.csv… See the full description on the dataset page: https://huggingface.co/datasets/FredZhang7/toxi-text-3M.uniref-50-foldseek-v1BRSpeech
BRSpeech
BRSpeech is a single-speaker Brazilian Portuguese speech dataset extracted and curated specifically for Text-to-Speech (TTS) and voice modeling tasks.
It corresponds directly to speaker 2961 from the multi-speaker BRSpeech-TTS dataset, which represents the speaker with the highest volume of recorded audio/hours in the entire corpus.
Dataset Summary
Language: Portuguese (pt-BR)
Speaker ID: 2961 (from BRSpeech-TTS)
Task: Single-speaker Text-to-Speech… See the full description on the dataset page: https://huggingface.co/datasets/freds0/BRSpeech.malicious-website-features-2.4MImportant Notice:
A subset of the URL dataset is from Kaggle, and the Kaggle datasets contained 10%-15% mislabelled data. See this dicussion I opened for some false positives. I have contacted Kaggle regarding their erroneous "Usability" score calculation for these unreliable datasets.
The feature extraction methods shown here are not robust at all in 2023, and there're even silly mistakes in 3 functions: not_indexed_by_google, domain_registration_length, and age_of_domain.
The features… See the full description on the dataset page: https://huggingface.co/datasets/FredZhang7/malicious-website-features-2.4M.cml_tts_dataset_spanishshared-ethical-memory-sem-2063Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
SoAyBench
SoAyBench
by WangYC
We've based SoAyBench creation on AMiner. To really understand how well LLMs can use SoAPI, we need to make AMiner's basic SoAPIs available for LLMs to use. We also need a test set made up of academic (question, solution, answer) triplets for checking how they're doing. The tricky part is, academic data keeps changing fast – stuff like info on scholars and their publications. So, keeping a test set with fixed answers is tough.
To tackle this, what we've done is… See the full description on the dataset page: https://huggingface.co/datasets/frederickwang99/SoAyBench.LLaDA-Sample-ES
Dataset: LLaDA-Sample-ES
Base: crscardellino/spanish_billion_words
Purpose: Training LLaDA (Large Language Diffusion Models)
Preprocessing
Tokenizer: GSAI-ML/LLaDA-8B-Instruct
Chunking: Up to 4,096 tokens per chunk (1% of chunks randomly sized between 1–4,096 tokens)
Noisy masking: Applied with noise factor ε = 1×10⁻³
Fields per chunk (PyTorch tensors):
input_ids
noisy_input_ids
mask
t (time scalar)
Statistics
Total chunks: ~ 652,089
Shards: 65… See the full description on the dataset page: https://huggingface.co/datasets/FredyRivera-dev/LLaDA-Sample-ES.cml_tts_dataset_portuguesecml_tts_dataset_frenchgradio-reviewsBG1Flux2-Image
Flux2-Image
This dataset is used for the Flux2-from-scratch project to train the Flux2 transformer from scratch.
Dataset structure
.
├── images_1/
│ └── a4ad6f1d46fb4ef3bcda56f05eacc537.png...
├── images_2/
│ └── 000121ce53464a63ad43ff134e979c7a.png...
├── images_3/
│ └── a47dc3302bb64b2ebfc7d525c8f0a88e.png...
├── images_4/
│ └── 0053ea2a8c764123b263d88c217f2995.png...
├── images_5/
│ └── 23fb74d40d8e4635914b4a399ee27b71.png...
└── data.csv… See the full description on the dataset page: https://huggingface.co/datasets/FredyRivera-dev/Flux2-Image.esm-teddymer-pseudodimers
ESM-Teddymer pseudo-dimers
60,177,402 intra-chain domain pairs ("pseudo-dimers") cut out of ESM Metagenomic Atlas
monomers at Chainsaw/TED domain boundaries. Each row is a target domain and a binder
domain that were adjacent in one real folded chain, so the pair comes with a real interface
without anyone having to dock anything. Built to train target-conditioned binder-design models.
All-atom structures for both chains ship alongside as foldcomp.
What is in here… See the full description on the dataset page: https://huggingface.co/datasets/fredzzp/esm-teddymer-pseudodimers.arc-agi-3-wm-traces
ARC-AGI-3 World Model Traces
This dataset contains ARC-AGI-3 transition traces in the same parquet schema used by HHazard/arc-agi-3.
Each row is one environment transition:
state, game_id, level_id, action_id, action_args, next_state,
level_done, frame_idx, origin, transformation, player
state and next_state are 64x64 ARC grids stored as nested integer arrays. action_id is the ARC-AGI-3 action kind; click actions use action_args.x and action_args.y.
Splits… See the full description on the dataset page: https://huggingface.co/datasets/fredericowieser/arc-agi-3-wm-traces.LuxAlign
Dataset Card for LuxAlign
Loading the Dataset
The dataset is currently at version v3, which can be loaded as:
from datasets import load_dataset
ds = load_dataset("fredxlpy/LuxAlign", name="lb-en") # or "lb-fr"
If you want to reproduce the results from the paper (v1) or use any previous version, you can specify the version folder:
# Load version v1 (as used in the paper)
ds_v1 = load_dataset("fredxlpy/LuxAlign", data_dir="data/v1", data_files={"train": "lb_en.json"})… See the full description on the dataset page: https://huggingface.co/datasets/fredxlpy/LuxAlign.ParaLux
Dataset Card for ParaLux Benchmark
Dataset Summary
ParaLux is a Luxembourgish paraphrase detection benchmark that requires models to identify the correct paraphrase from two candidates for a given anchor sentence: one representing a valid paraphrase and the other an adversarial not_paraphrase. The dataset, consisting of 312 examples, is sourced from news articles published by RTL.lu and was introduced in LuxEmbedder: A Cross-Lingual Approach to Enhanced Luxembourgish… See the full description on the dataset page: https://huggingface.co/datasets/fredxlpy/ParaLux.heliumcml_tts_dataset_germancml_tts_dataset_italianmixprotein
