datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rfashban24rfantibody-assets
LevinHarness/rfantibody-assets — public mirror of third-party runtime assets
This dataset is a public mirror of third-party runtime assets
required by the Levin Harness plugin(s) listed below, mirrored
verbatim from their original sources with SHA-256 pinning. It is
not an official distribution: nothing here is published under
this account's own terms, and it is not affiliated with or endorsed
by any upstream project.
Ownership and licensing
Every file remains… See the full description on the dataset page: https://huggingface.co/datasets/LevinHarness/rfantibody-assets.crypto-5s-market-data-adausdc-sample
ADA/USDC High-Frequency Market Microstructure Data
Free 7-Day Sample
This repository provides a free 7-day sample of a much larger privately collected high-frequency cryptocurrency market dataset.
The complete historical archive contains millions of market snapshots, with data collection starting in December 2025, across 12 crypto/USDC markets.
The public ADA/USDC sample contains:
81,579 market snapshots
97 columns
7 days of continuous historical data
20 bid + 20… See the full description on the dataset page: https://huggingface.co/datasets/rfab85/crypto-5s-market-data-adausdc-sample.rfam
Rfam
Rfam is a database of structure-annotated multiple sequence alignments, covariance models and family annotation for a number of non-coding RNA, cis-regulatory and self-splicing intron families.
The seed alignments are hand curated and aligned using available sequence and structure data, and covariance models are built from these alignments using the INFERNAL v1.1.4 software suite.
The full regions list is created by searching the RFAMSEQ database using the covariance model… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/rfam.rfam-standardrfantibody-assets
LevinHarness/rfantibody-assets — public mirror of third-party runtime assets
This dataset is a public mirror of third-party runtime assets
required by the Levin Harness plugin(s) listed below, mirrored
verbatim from their original sources with SHA-256 pinning. It is
not an official distribution: nothing here is published under
this account's own terms, and it is not affiliated with or endorsed
by any upstream project.
Ownership and licensing
Every file remains… See the full description on the dataset page: https://huggingface.co/datasets/sgetttt/rfantibody-assets.rfa_shan_language_voices
RFA Shan Language Voices
This dataset contains 20.58 hours of audio in the Shan (Tai-Yai) language, sourced from news broadcasts by Radio Free Asia (RFA) Burmese. This is one of the largest publicly accessible audio resources for the Shan language, designed to support research in low-resource automatic speech recognition (ASR), voice activity detection, and other speech-related tasks.
The audio has been automatically segmented into 5,047 manageable chunks and prepared in the… See the full description on the dataset page: https://huggingface.co/datasets/myandev/rfa_shan_language_voices.rfa_rakhine_language_voices
RFA Rakhine Language Voices
This dataset contains 14.53 hours of audio in the Rakhine (Arakanese) language, sourced from news broadcasts by Radio Free Asia (RFA) Burmese. This is one of the largest publicly accessible audio resources for the Rakhine language, designed to support research in low-resource automatic speech recognition (ASR), voice activity detection, and other speech-related tasks.
The audio has been automatically segmented into manageable chunks and prepared in the… See the full description on the dataset page: https://huggingface.co/datasets/freococo/rfa_rakhine_language_voices.rfam-taxonomy-lookuprfa_shan_language_voices
RFA Shan Language Voices
This dataset contains 20.58 hours of audio in the Shan (Tai-Yai) language, sourced from news broadcasts by Radio Free Asia (RFA) Burmese. This is one of the largest publicly accessible audio resources for the Shan language, designed to support research in low-resource automatic speech recognition (ASR), voice activity detection, and other speech-related tasks.
The audio has been automatically segmented into 5,047 manageable chunks and prepared in the… See the full description on the dataset page: https://huggingface.co/datasets/freococo/rfa_shan_language_voices.rfa_forge_20250721_024738This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 1,
"total_frames": 290,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/near0248/rfa_forge_20250721_024738.Arabic_Diacritized_Audio_Datasetrfa_forge_20250721_024015This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 1,
"total_frames": 290,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/near0248/rfa_forge_20250721_024015.edited_Lerobot_place_tube_rack_Rfar_RcamThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 15,
"total_frames": 3934,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:15"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/gravta42/edited_Lerobot_place_tube_rack_Rfar_Rcam.RF_Anomaly_DetectionEach dataset consists of
1,000 samples, produced by synthesizing single-channel base-
band signals modulated with 5G-standard quadrature phase-
shift keying (QPSK). The signals incorporate realistic pulse
shaping and interference patterns to emulate characteristics
representative of modern 5G communication systems. For
reproducibility, fixed seed values were used during generation.
Each sample contains 1,024 complex-valued in-phase and
quadrature (I/Q) samples, captured at a 1 MHz sampling… See the full description on the dataset page: https://huggingface.co/datasets/b4byn1cky/RF_Anomaly_Detection.edited_Lerobot_place_tube_rack_Rfar_LcamThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 13,
"total_frames": 3216,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:13"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/gravta42/edited_Lerobot_place_tube_rack_Rfar_Lcam.Lerobot_place_tube_rack_Rfar_RcamThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 18,
"total_frames": 4627,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:18"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/gravta42/Lerobot_place_tube_rack_Rfar_Rcam.Lerobot_place_tube_rack_Rfar_LcamThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 16,
"total_frames": 3972,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:16"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/gravta42/Lerobot_place_tube_rack_Rfar_Lcam.ADE20k-847_validation_annotationsDiacritized-Case-Ending-Errors-TTT-V2.0Diacritized-Case-Ending-Errors-TTT-V3.0TrashPanda
TrashPanda Dataset
A labeled image dataset for German household waste classification, collected for a CNN training project at DHBW.
Images are organized into subfolders by class. Each file follows the naming scheme <ClassName>_<index>.jpg.
Classes
Label
Folder
Description
Images
0
Biomuell
Organic / food waste
1 005
1
GelberSack
Recyclable packaging (plastics, metals, composites)
5 032
2
Glass
Glass bottles and jars
3 001
3
Papier
Paper and cardboard
3… See the full description on the dataset page: https://huggingface.co/datasets/rfabrik/TrashPanda.Diacritized-Case-Ending-Errors-TTT-V1.0Meal_Planner_2kplusdataset
Meal Planner 2000+ sampkle dataset
rfam_sample_padded_arraysRFAA
RFAA
This repository is a public research-data copy of the RFAA directory used by
the DCNPA web project.
Layout
The source data contains four top-level collections:
Test251
Test167
Test1440
LEADS-PEP
Because the source contains many files, each collection is stored as a series
of independent tar archives under archives/<collection>/. File lists and
original byte sizes are stored under manifests/.
Extract all shards of a collection with:
for archive in… See the full description on the dataset page: https://huggingface.co/datasets/tqhu/RFAA.RFAGGModels-Comparison-DatasetClean-vs-Undiacritized-Datasetaudio_rfa
