datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
london-cycles
London's rental bicycle network usage
15-minute snapshots of every docking station in London's cycle hire scheme, collected from the TfL BikePoint API since 2022-04-29.
Collection code lives on GitHub: fferegrino/london-cycles-db. A new snapshot is appended every 15 minutes; each finished day is compacted into a single file overnight.
Layout
data/year=YYYY/month=MM/day=DD/part.csv one row per station per snapshot
stations.csv… See the full description on the dataset page: https://huggingface.co/datasets/feregrino/london-cycles.fer-2013rhan-checkpoints-rollingwds_fer2013fermi-lat-weekly-photons
Fermi-LAT weekly photons
This dataset contains Fermi Large Area Telescope all-sky weekly photon files
from mission week w009 through w153, frozen on 2026-08-30. Its 145
configurations correspond one-to-one with the weekly p305_v001 FITS files.
Each Parquet row is an EVENTS row, with the 23 FITS-named columns in their
stored order and shape.
Mission weeks run Thursday through Wednesday in UTC. The first configuration
begins with the science-phase interval on 2008-08-04; w153… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/fermi-lat-weekly-photons.rhan-checkpointsQtMeshEditor-motion-corpus
QtMeshEditor Motion Corpus
A permissively-licensed animated-humanoid corpus: rigged 3D characters
with skeletal animation clips, harvested for
QtMeshEditor's text-to-motion
v2 work (epic #837)
— the template clip library and the training set for a from-scratch
flow-matching motion model.
Every asset is CC0 or CC-BY — nothing here derives from Mixamo, LAFAN1,
Bandai-Namco, AMASS/HumanML3D, or game rips (all license-poisoned for
commercial redistribution). That makes this corpus —… See the full description on the dataset page: https://huggingface.co/datasets/fernandotonon/QtMeshEditor-motion-corpus.macro-mdsinfore2_audiobooks
unofficial mirror of InfoRe Technology public dataset №2
official announcement: https://www.facebook.com/groups/j2team.community/permalink/1010834009248719/
415h, 315k samples, vietnamese audiobooks of chinese wǔxiá 武俠 & xiānxiá 仙俠
bộ dữ liệu bóc ra từ YouTube đọc truyện võ hiệp & tiên hiệp, áp dụng kĩ thuật đối chiếu văn bản để dán nhãn tự động
official download:… See the full description on the dataset page: https://huggingface.co/datasets/doof-ferb/infore2_audiobooks.fine-t2i-1024-mdsfly-connectome-49k
Fly Connectome 49k
The 49,393-neuron / 9,050,172-edge central-brain graph that
ngxson/fly-llm-hf uses as its recurrent layer,
packaged as plain typed arrays and joined to the MaleCNS body IDs, cell types,
superclasses and soma positions it was derived from.
Until now this graph was reachable only by loading a 284 MB model checkpoint and reading its
frozen buffers. This dataset is that graph, verifiable on its own, so you can build on the
wiring without adopting anyone's model.… See the full description on the dataset page: https://huggingface.co/datasets/fernandofernandes/fly-connectome-49k.rhan-eval-sweepVietnam-Celeb
unofficial mirror of Vietnam-Celeb dataset
official announcement:
https://www.isca-archive.org/interspeech_2023/pham23b_interspeech.html
https://github.com/Vietnam-Celeb/Vietnam-Celeb
https://huggingface.co/datasets/hustep-lab/Vietnam-Celeb
official download:
Part 0: https://drive.google.com/file/d/1pMuT3DFzSwib7SVcRS8VkDwPuLTsemSG/view?usp=share_link
Part 1: https://drive.google.com/file/d/1xayHt2HRqE1aJ4HvtUT40_9XlgvfDfRY/view?usp=share_linkPart 2:… See the full description on the dataset page: https://huggingface.co/datasets/doof-ferb/Vietnam-Celeb.Fermatix-SWE-Bench
Fermatix SWE Bench
This repository contains tasks from the Fermatix SWE Bench dataset, intended for evaluating the capabilities of models in automatic bug fixing and code modification.
The tasks cover various programming languages and projects, providing a diverse set of scenarios for testing and training.
Each task includes:
Patches with fixes and/or tests
Instructions for building and running (in the form of a Dockerfile), as well as corresponding run logs
A parquet file… See the full description on the dataset page: https://huggingface.co/datasets/fermatix-ai/Fermatix-SWE-Bench.kk-tokenizer-fertility-baseline
Kazakh Tokenizer Fertility Baseline
Reproducible fertility benchmark of subword tokenizers on the Kazakh language. Companion
artifact for the paper "Tokenizer Optimization for Kazakh Small Language Models"
(in preparation, target: ACM TALLIP).
Headline numbers
Tokenizer
Fertility
🥇 Best overall
kk-bpe-32k
1.679
🚨 Worst
GPT-4 (cl100k)
5.895
GPT-4 penalty
GPT-4 (cl100k) is 3.51× worse than the best Kazakh-trained tokenizer
→ The custom Kazakh… See the full description on the dataset page: https://huggingface.co/datasets/Abzalbek89/kk-tokenizer-fertility-baseline.vlsp2020_vinai_100h
unofficial mirror of VLSP 2020 - VinAI - ASR challenge dataset
official announcement:
tiếng việt: https://institute.vinbigdata.org/events/vinbigdata-chia-se-100-gio-du-lieu-tieng-noi-cho-cong-dong/
in eglish: https://institute.vinbigdata.org/en/events/vinbigdata-shares-100-hour-data-for-the-community/
VLSP 2020 workshop: https://vlsp.org.vn/vlsp2020
official download: https://drive.google.com/file/d/1vUSxdORDxk-ePUt-bUVDahpoXiqKchMx/view?usp=sharing
contact: info@vinbigdata.org… See the full description on the dataset page: https://huggingface.co/datasets/doof-ferb/vlsp2020_vinai_100h.fer2013FER-2013
Facial Expression Recognition Challenge (ICML 2013)
Overview
This dataset was created for the Challenges in Representation Learning: Facial Expression Recognition Challenge, part of the ICML 2013 Workshop. The challenge focused on evaluating how well learning algorithms can generalize to newly introduced data, particularly for facial expression classification.
Start Date: April 13, 2013
End Date: May 25, 2013
The dataset contains facial images labeled with one of… See the full description on the dataset page: https://huggingface.co/datasets/DerrickUnleashed/FER-2013.LSVSC
unofficial mirror of LSVSC dataset (novel large-scale Vietnamese speech corpus)
official announcement: https://www.mdpi.com/2079-9292/13/5/977
official download: https://drive.google.com/drive/folders/1tiPKaIOC7bt6isv5qFqf61O_2jFK8ZOI
100h, 57k samples
pre-process: see my code: https://github.com/phineas-pta/fine-tune-whisper-vi/blob/main/misc/clean-lsvsc.py
need to do: check misspelling, restore foreign words phonetised to vietnamese
usage with HuggingFace:
# pip install -q… See the full description on the dataset page: https://huggingface.co/datasets/doof-ferb/LSVSC.enron-ferc-pst
Enron FERC email corpus in native PST
The EDRM Enron v2 email corpus in Microsoft PST format, modified to reduce personal privacy risk. Mailbox structure, MAPI metadata, message bodies, and retained attachments are preserved.
The release contains 171 PST files in data/, with one or more files per custodian.
Count
Version
v1
Messages
1,226,178
Attachments
453,832
PST files
171
Possible uses include email research, e-discovery testing, information retrieval… See the full description on the dataset page: https://huggingface.co/datasets/intellekthq/enron-ferc-pst.fermi-gbm-triggers
Fermi GBM All-Trigger Catalog
Credit: NASA/JPL-Caltech/STScI/CXC/SAO
Part of a dataset collection on Hugging Face.
Dataset description
All triggers from the Fermi Gamma-ray Burst Monitor -- GRBs, solar flares, SGRs, terrestrial particles, and more. Updated daily from NASA HEASARC.
The Fermi GBM detects transient events across the full unocculted sky in the 8 keV to 40 MeV energy range. While the confirmed GRB catalog contains only verified gamma-ray… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/fermi-gbm-triggers.mj-llava-mdsFeruzaSpeech
Dataset Card for Dataset Name
FeruzaSpeech is a read speech dataset
of the Uzbek language, transcribed in both Cyrillic
and Latin alphabets, freely available for academic research purposes.
It includes 60 hours of high-quality recordings
from a single native female speaker from Tashkent, Uzbekistan.
ICNLSPConference: https://www.youtube.com/watch?v=9whj9yzI_s4&ab_channel=ICNLSPConference
Paper: https://arxiv.org/abs/2410.00035
Example test.tsv:
audio text_latin… See the full description on the dataset page: https://huggingface.co/datasets/k2speech/FeruzaSpeech.Nemotron-Reason2-TOKENfer2013-enhanced
FER2013 Enhanced: Advanced Facial Expression Recognition Dataset
The most comprehensive and quality-enhanced version of the famous FER2013 dataset for state-of-the-art emotion recognition research and applications.
🎯 Dataset Overview
FER2013 Enhanced is a significantly improved version of the landmark FER2013 facial expression recognition dataset. This enhanced version provides AI-powered quality assessment, balanced data splits, comprehensive metadata, and multi-format… See the full description on the dataset page: https://huggingface.co/datasets/abhilash88/fer2013-enhanced.PAVfer2013_train_publicTest_privateTest
Dataset Card for "fer2013_train_publicTest_privateTest"
More Information needed
fer2013fly-connectome-malecns-166k
Fly Connectome — MaleCNS 166k
The whole fruit fly central nervous system, as plain typed arrays.
166,700 neurons, 25,582,938 directed connections, 124,177,617 synaptic contacts —
raw signed weights, nothing normalized, in a destination-major CSR you can np.fromfile
and use in about four lines.
This is the complete graph. If you want only the 49,393-neuron central-brain subset
that the Fly LLM language models run on, with measured soma positions, that is
fly-connectome-49k.
The… See the full description on the dataset page: https://huggingface.co/datasets/fernandofernandes/fly-connectome-malecns-166k.infore1_25hours
unofficial mirror of InfoRe Technology public dataset №1
official announcement: https://www.facebook.com/groups/j2team.community/permalink/1010834009248719/
25h, 14.9k samples, InfoRe paid a contractor to read text
official download: magnet:?xt=urn:btih:1cbe13fb14a390c852c016a924b4a5e879d85f41&dn=25hours.zip&tr=http%3A%2F%2Foffice.socials.vn%3A8725%2Fannounce
mirror: https://files.huylenguyen.com/datasets/infore/25hours.zip
unzip password: BroughtToYouByInfoRe
pre-process: see… See the full description on the dataset page: https://huggingface.co/datasets/doof-ferb/infore1_25hours.
