datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dataTAGARELA
TAGARELA: A Portuguese Speech Dataset From Podcasts
TAGARELA is a large-scale Portuguese speech dataset built from podcast audio and curated for speech technology research, especially Automatic Speech Recognition (ASR) and Text-to-Speech (TTS).
The dataset contains more than 8,972 hours of Portuguese speech derived from the Cem Mil Podcasts collection. It includes Brazilian Portuguese and European Portuguese speech, processed through a pipeline involving audio standardization… See the full description on the dataset page: https://huggingface.co/datasets/freds0/TAGARELA.uniref-50-foldseek-v1esm-teddymer-pseudodimers
ESM-Teddymer pseudo-dimers
60,177,402 intra-chain domain pairs ("pseudo-dimers") cut out of ESM Metagenomic Atlas
monomers at Chainsaw/TED domain boundaries. Each row is a target domain and a binder
domain that were adjacent in one real folded chain, so the pair comes with a real interface
without anyone having to dock anything. Built to train target-conditioned binder-design models.
All-atom structures for both chains ship alongside as foldcomp.
What is in here… See the full description on the dataset page: https://huggingface.co/datasets/fredzzp/esm-teddymer-pseudodimers.BRSpeech
BRSpeech
BRSpeech is a single-speaker Brazilian Portuguese speech dataset extracted and curated specifically for Text-to-Speech (TTS) and voice modeling tasks.
It corresponds directly to speaker 2961 from the multi-speaker BRSpeech-TTS dataset, which represents the speaker with the highest volume of recorded audio/hours in the entire corpus.
Dataset Summary
Language: Portuguese (pt-BR)
Speaker ID: 2961 (from BRSpeech-TTS)
Task: Single-speaker Text-to-Speech… See the full description on the dataset page: https://huggingface.co/datasets/freds0/BRSpeech.cml_tts_dataset_spanishcml_tts_dataset_portuguesecml_tts_dataset_frenchcml_tts_dataset_germanmixproteinhelium_memory
Try gpt-oss ·
Guides ·
Model card ·
OpenAI blog
Welcome to the gpt-oss series, OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases.
We’re releasing two flavors of these open models:
gpt-oss-120b — for production, general purpose, high reasoning use cases that fit into a single 80GB GPU (like NVIDIA H100 or AMD MI300X) (117B parameters with 5.1B active parameters)gpt-oss-20b — for lower latency, and local or… See the full description on the dataset page: https://huggingface.co/datasets/Fred808/helium_memory.arc-agi-3-wm-traces
ARC-AGI-3 World Model Traces
This dataset contains ARC-AGI-3 transition traces in the same parquet schema used by HHazard/arc-agi-3.
Each row is one environment transition:
state, game_id, level_id, action_id, action_args, next_state,
level_done, frame_idx, origin, transformation, player
state and next_state are 64x64 ARC grids stored as nested integer arrays. action_id is the ARC-AGI-3 action kind; click actions use action_args.x and action_args.y.
Splits… See the full description on the dataset page: https://huggingface.co/datasets/fredericowieser/arc-agi-3-wm-traces.uniref50-sorted-structure-tokenfine_code
Fine Code
A collection of high-quality code dataset.
Consists of the following:
https://huggingface.co/datasets/OpenCoder-LLM/opc-annealing-corpus
https://huggingface.co/datasets/inclusionAI/Ling-Coder-SyntheticQA
cml_tts_dataset_dutchBRSpeech-TTSproteingym_substitutionsPink_pen-Yellow_2_20260824_224048This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/fredhugg/Pink_pen-Yellow_2_20260824_224048.FRED-CONVERTEDPink_pen-Yellow_20260824_211746This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/fredhugg/Pink_pen-Yellow_20260824_211746.Pink_pen_20260823_225553This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/fredhugg/Pink_pen_20260823_225553.cml_tts_dataset_italianPink_pen-v2_20260824_085530This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/fredhugg/Pink_pen-v2_20260824_085530.Uniref50
Uniref50: Uniref Sequences clustered at 50% sequence identity
~40M Protein Sequences.
Split into train val and test.
Usage
from datasets import load_dataset
# Step 1: Load the dataset from HuggingFace Hub
dataset = load_dataset("zhangzhi/Uniref50")
# Step 2: Access a specific split (e.g., "train", "validation", "test")
train_split = dataset["train"]
print(f"Number of sequences in the train split: {len(train_split)}")
spcas9cgu__notas_fiscais
Dataset Card: cgu_notas_fiscais
Data from electronic invoices for federal government purchases made available by
Comptroller General of the Union (Controladoria-Geral da União), which is a
Brazilian federal government agency responsible for oversight and transparency.
Dataset Details
Dataset Description
Curated by: Fred Guth (@fredguth)
Funded by: World Bank
Language(s) (NLP): pt-br
License: CC-BY 4.0
Dataset Sources
The source of this datasets… See the full description on the dataset page: https://huggingface.co/datasets/fredguth/cgu__notas_fiscais.CaMiT
CaMiT: Car Models in Time
CaMiT (Car Models in Time) is a large-scale, fine-grained, time-aware dataset of car images collected from Flickr. It is designed to support research on temporal adaptation in visual models, continual learning, and time-aware generative modeling.
Dataset Highlights
Labeled Subset:
787,000 samples
190 car models
2007–2023
Unlabeled Pretraining Subset:
5.1 million samples
2005–2023
Metadata includes:
Image URLs (not the images themselves)… See the full description on the dataset page: https://huggingface.co/datasets/fredericlin/CaMiT.muong_voice_textOMG_prot50edushorts
