datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TKEval
Dataset Card for TKEval
Contents
Dataset Description
Dataset Structure
Dataset Splits
Citation
Dataset Description
The curse of tokenization: Language models typically tokenize raw text into sequences of subword identifiers from a predefined vocabulary, a process inherently sensitive to typographical errors, length variations, and largely oblivious to the internal structure of tokens.
TKEval is an evalution benchmark for systematicly assessing the impact of… See the full description on the dataset page: https://huggingface.co/datasets/floatai/TKEval.AFO-Aerial_Floating_Objects
Dataset Card for AFO - Aerial Floating Objects
AFO dataset is the first free dataset for training machine learning and deep learning models for maritime Search and Rescue applications. It contains aerial-drone videos with 40,000 hand-annotated persons and objects floating in the water, many of small size, which makes them difficult to detect.
This is a FiftyOne dataset with 1014 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/AFO-Aerial_Floating_Objects.sc-floatwaste-recordingswater-hyacinth-flowers-floating-pond
Water Hyacinth Flowers — Khurushkul Pond, Bangladesh
150 ground-level photographs of water hyacinth (Eichhornia crassipes) blooming in a freshwater pond near Khurushkul, Cox's Bazar, Bangladesh. All frames were captured in a single session on 9 August 2026 (16:31–16:45 local time) during the monsoon season.
This is a follow-up survey of the same pond documented in the June 2026 water lily / water hyacinth dataset by the same photographer.
Contents
150 JPG images… See the full description on the dataset page: https://huggingface.co/datasets/golamrob/water-hyacinth-flowers-floating-pond.mathmetics-dataset-float-long
Transformer Math Dataset (250,000,000 Samples Sharded)
High-precision synthetic mathematical expression dataset generated for training sequence-to-sequence math Transformers in JAX/Flax.
Dataset Structure
Total Samples: 250,000,000
Shard Format: JSONL sharded files (100,000 samples per shard)
Supported Operations: +, -, *, /, ^, sin, cos, tan, log, ln, exp, sqrt, abs
Expression Depth Range: Depth 4 to 6
Integer Operand Ratio: 20%
Data Fields
Each… See the full description on the dataset page: https://huggingface.co/datasets/saidurga001301/mathmetics-dataset-float-long.WTK-15min-float32AE29H_float32
AE29H_float32
Audio Embeddings ~29 hours dataset contains precomputed audio embeddings designed for Nanowakeword framework. The embeddings are intended to be used as general-purpose negative training data, meaning the audio does not contain the target wake word or phrase.
Unlike raw audio datasets, the files in this dataset contain low-dimensional audio embeddings extracted from audio clips using a pre-trained speech embedding model. These embeddings can be directly used as input… See the full description on the dataset page: https://huggingface.co/datasets/arcosoph/AE29H_float32.AFO-Aerial_Floating_Objects
Dataset Card for AFO - Aerial Floating Objects
AFO dataset is the first free dataset for training machine learning and deep learning models for maritime Search and Rescue applications. It contains aerial-drone videos with 40,000 hand-annotated persons and objects floating in the water, many of small size, which makes them difficult to detect.
This is a FiftyOne dataset with 1014 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U… See the full description on the dataset page: https://huggingface.co/datasets/12edf/AFO-Aerial_Floating_Objects.robocasa_spatial_camrandom_float16_pretrain_atomic_CheesyBreadThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "PandaOmron",
"total_episodes": 101,
"total_frames": 31141,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:101"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kimz1121/robocasa_spatial_camrandom_float16_pretrain_atomic_CheesyBread.yoruba-cfm-latentsfrancecrops-float16
FranceCrops Float16 Dataset
Memory-efficient version of the FranceCrops dataset stored in float16 format.
Usage
from datasets import load_dataset
# Load dataset
dataset = load_dataset("saget-antoine/francecrops-float16")
# Load with specific format (recommended for training)
# As float32
dataset = dataset.with_format('torch', dtype=torch.float32)
# Or as float16 for memory efficiency
dataset = dataset.with_format('torch', dtype=torch.float16)
# Or as bfloat16 for TPU… See the full description on the dataset page: https://huggingface.co/datasets/saget-antoine/francecrops-float16.A-dataset-of-images-of-floating-rubbish-on-the-surface-of-the-water
License
This dataset is released under Creative Commons Attribution 4.0 International (CC BY 4.0).
Dataset: A-dataset-of-images-of-floating-rubbish-on-the-surface-of-the-water
Author: gonzz2026
URL: https://huggingface.co/datasets/gonzz2026/A-dataset-of-images-of-floating-rubbish-on-the-surface-of-the-water
License: CC BY 4.0
DOI
https://doi.org/10.57967/hf/7674
Citation
@misc{gonzz2026_floating_rubbish_2026,
title = {A Dataset of Images of… See the full description on the dataset page: https://huggingface.co/datasets/gonzz2026/A-dataset-of-images-of-floating-rubbish-on-the-surface-of-the-water.HumanEval-XLA collection of cross-lingual benchmark for code generation.bread1_combinedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so-101",
"total_episodes": 55,
"total_frames": 3262,
"total_tasks": 2,
"total_videos": 110,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 60,
"splits": {
"train": "0:55"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/float-lab/bread1_combined.lesandwichThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so-101",
"total_episodes": 85,
"total_frames": 4419,
"total_tasks": 3,
"total_videos": 170,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 60,
"splits": {
"train": "0:85"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/float-lab/lesandwich.A-dataset-of-images-of-floating-rubbish-on-the-surface-of-the-water
License
This dataset is released under Creative Commons Attribution 4.0 International (CC BY 4.0).
Dataset: A-dataset-of-images-of-floating-rubbish-on-the-surface-of-the-water
Author: gonzz2026
URL: https://huggingface.co/datasets/gonzz2026/A-dataset-of-images-of-floating-rubbish-on-the-surface-of-the-water
License: CC BY 4.0
DOI
https://doi.org/10.57967/hf/7674
Citation
@misc{gonzz2026_floating_rubbish_2026,
title = {A Dataset of Images of… See the full description on the dataset page: https://huggingface.co/datasets/houkaiyu/A-dataset-of-images-of-floating-rubbish-on-the-surface-of-the-water.mair-floating-point-vs-quantized-benchmarks
Dataset Coverage
The benchmarks evaluate performance across 14+ specialized datasets covering:
Legal & Regulatory: ACORDAR, AILA2019-Case, AILA2019-Statutes, LeCaRDv2, LegalQuAD, REGIR-EU2UK, REGIR-UK2EU.
Financial: ConvFinQA, FinanceBench, FinQA, FiQA, HC3Finance.
Medical & Clinical: NFCorpus.
General/API: Apple Documentation.
Metrics
NDCG@10: Normalized Discounted Cumulative Gain at rank 10, measuring retrieval quality.
Latency (ms): Mean search latency measured… See the full description on the dataset page: https://huggingface.co/datasets/moorcheh/mair-floating-point-vs-quantized-benchmarks.lesandwich-combined_v2floating-det
Description
A floating object is a common type of page element found in academic literature, books, and other formal publications. In LaTeX, a floating object typically refers to a container that can hold text, images, tables, code, algorithms, and other content. The placement of these containers within the document is automatically adjusted by LaTeX to fit the page layout. To facilitate indexing and readability, floating objects are usually accompanied by additional information… See the full description on the dataset page: https://huggingface.co/datasets/irhawks/floating-det.ecg-comprehension-bpm-float-200000-250-2500ThaiIDCardSynt
Dataset Details
Dataset Description
Curated by: Matichon Maneegard
Shared by [optional]: Matichon Maneegard
Language(s) (NLP): image-to-text
License: apache-2.0
Dataset Sources [optional]
The dataset was entirely synthetic. It does not contain real information or pertain to any specific person.
Uses
Direct Use
Using for tranning OCR or Multimodal.
Dataset Structure
This dataset contains 98 x 6 = 588 samples, and the… See the full description on the dataset page: https://huggingface.co/datasets/Float16-cloud/ThaiIDCardSynt.float-talking-headFLOATBench
FLOATBench: Wind Turbine Tower Damage
Paper | Project Page | GitHub
Tabular fatigue dataset for 22 MW floating offshore wind turbine (FOWT) towers. Contains 582,120 labelled tower section fatigue damage records across three tower geometries: the IEA-22 reference turbine baseline (ref) and two FLOAT-derived re-designs (opt1, opt2).
FLOATBench is the first FOWT fatigue benchmark for tabular surrogate modeling, offering an evaluation protocol that generalizes to engineering surrogates… See the full description on the dataset page: https://huggingface.co/datasets/DeCoDELab/FLOATBench.patty1_combined_v2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so-101",
"total_episodes": 42,
"total_frames": 2395,
"total_tasks": 1,
"total_videos": 84,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 60,
"splits": {
"train": "0:42"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/float-lab/patty1_combined_v2.ecg-comprehension-bpm-float-100000-250-2500TULU3TIPAlesandwich2-grootThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so-101",
"total_episodes": 83,
"total_frames": 4130,
"total_tasks": 3,
"total_videos": 166,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 60,
"splits": {
"train": "0:83"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/float-lab/lesandwich2-groot.nyu_depth_train_float32_1449patty1_combinedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so-101",
"total_episodes": 43,
"total_frames": 2438,
"total_tasks": 1,
"total_videos": 86,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 60,
"splits": {
"train": "0:43"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/float-lab/patty1_combined.bread1_combined_v2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so-101",
"total_episodes": 50,
"total_frames": 3026,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 60,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/float-lab/bread1_combined_v2.
