datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
seed
BONES-SEED: Skeletal Everyday Embodiment Dataset
BONES-SEED is an open dataset of 142,220 annotated human motion animations for humanoid robotics. It provides motion capture data in SOMA and Unitree G1 formats, with natural language descriptions, temporal segmentation, and detailed skeletal metadata.
Project website: bones.studio/datasets/seed
Interactive viewer: seed-viewer.bones.studio
Associated code: github.com/bones-studio/seed-viewer
Total motions142,220 (71… See the full description on the dataset page: https://huggingface.co/datasets/bones-studio/seed.SURDSThis data is sourced from the image of the nuScenes dataset. We extend our gratitude for their outstanding work!
BoneAgeTW2-cropsfull-math-private-n256-Qwen2.5-3B-Instruct-bonBonds-Daily-Price
Bonds Daily Price
This dataset includes daily rate data for various bonds.
1,415,096 rows over 209 symbols, 7 columns, covering 1969-04-30 to 2026-07-31. Refreshed monthly.
Strategies Built on This Data
590 papers in the Papers With Backtest catalogue declare this dataset as an input. 560 of them have been coded and run over their own full history. The median replicated Sharpe ratio is +0.48, and 59% clear a t-statistic of 1.96 on their own sample, against 48%… See the full description on the dataset page: https://huggingface.co/datasets/paperswithbacktest/Bonds-Daily-Price.full-math-private-n256-Phi-4-mini-instruct-bonfull-math-private-n256-Llama-3.2-3B-Instruct-bonbone_marrow_cell_dataset
About This Dataset
Bone marrow biopsy is procedure applied to collect and examine bone marrow — the spongy tissue inside some of your larger bones.
This biopsy can show whether your bone marrow is healthy and making normal amounts of blood cells. Doctors use these procedures to diagnose and monitor blood and marrow diseases, cancers, as well as fevers of unknown origin.
The dataset contains a collection of over 170,000 de-identified, expert-annotated cells from the bone marrow… See the full description on the dataset page: https://huggingface.co/datasets/ekim15/bone_marrow_cell_dataset.preprocessed-full-math-private-n256-Llama-3.2-3B-Instruct-bonmultilingual-benchmark-dataset
Multilingual benchmark dataset
This dataset contains cleaned text data organized by language, and can be used as a benchmark for evaluating language identification (LID) models.
Overview
This dataset consists of samples from multiple existing benchmark and multilingual text datasets. The full list of sources and licenses is provided here.
Languages: ~1,800 languages
Samples: Over 1M
Licenses: All sources are commercially usable — full license list provided in the sheet… See the full description on the dataset page: https://huggingface.co/datasets/Bonkh/multilingual-benchmark-dataset.afrolm_active_learning_dataset
AfroLM: A Self-Active Learning-based Multilingual Pretrained Language Model for 23 African Languages
GitHub Repository of the Paper
This repository contains the dataset for our paper AfroLM: A Self-Active Learning-based Multilingual Pretrained Language Model for 23 African Languages which will appear at the third Simple and Efficient Natural Language Processing, at EMNLP 2022.
Our self-active learning framework
Languages Covered
AfroLM has been… See the full description on the dataset page: https://huggingface.co/datasets/bonadossou/afrolm_active_learning_dataset.Chewy_Robotics_Bone_Bi-manual_PackingTotalsegmentor_Pelvis_Bone_Recon_Datasetbonsai-jetson-benchmark-7w
Bonsai Jetson Benchmark: 7W
Platform: NVIDIA Jetson Orin Nano Super 8GB · Power mode: 7W
Backend: llama.cpp build-jetson · CUDA · -ngl 99
Sweep: prompt in {256, 512, 1024, 2048} tok x gen in {128, 256, 512} tok · 20 reqs/combo
Context: 2560 tok · Concurrency: 1
Part of smolperfbenchmark, a public on-device
LLM benchmark leaderboard. Headline metric is output tok/J (tokens per joule),
computed over the decode phase.
Files
Bonsai-*, Ternary-Bonsai-*: per-combo… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/bonsai-jetson-benchmark-7w.full-math-private-Qwen2.5-3B-Instruct-bonsova_rudevices
Dataset Card for sova_rudevices
Dataset Summary
SOVA Dataset is free public STT/ASR dataset. It consists of several parts, one of them is SOVA RuDevices. This part is an acoustic corpus of approximately 100 hours of 16kHz Russian live speech with manual annotating, prepared by SOVA.ai team.
Authors do not divide the dataset into train, validation and test subsets. Therefore, I was compelled to prepare this splitting. The training subset includes more than 82 hours, the… See the full description on the dataset page: https://huggingface.co/datasets/bond005/sova_rudevices.stratified-solvable-1k-math-private-Qwen2.5-3B-Instruct-bonsberdevices_golos_10h_crowd
Dataset Card for sberdevices_golos_10h_crowd
Dataset Summary
Sberdevices Golos is a corpus of approximately 1200 hours of 16kHz Russian speech from crowd (reading speech) and farfield (communication with smart devices) domains, prepared by SberDevices Team (Alexander Denisenko, Angelina Kovalenko, Fedor Minkin, and Nikolay Karpov). The data is derived from the crowd-sourcing platform, and has been manually annotated.
Authors divide all dataset into train and test subsets.… See the full description on the dataset page: https://huggingface.co/datasets/bond005/sberdevices_golos_10h_crowd.rulibrispeech
Dataset Card for "rulibrispeech"
More Information needed
BonaFide
BonaFide
This is a dataset containing ground-truth faithfulness labels for chains of thought (CoTs), used for evaluating CoT faithfulness metrics. The current benchmark results are in the BonaFide benchmark space.
Methodology
We construct tasks whose outputs reveal which intermediate computations must have produced them, then label CoTs against those computations.
Diversionary setting. Each question is given alongside a misleading hint pointing to a random wrong answer.… See the full description on the dataset page: https://huggingface.co/datasets/yoavgurarieh/BonaFide.bonsai2-27b-mtp-repro
Ternary-Bonsai-2-27B + in-file MTP: reproduction bundle (RTX 4080 SUPER, Ada/SM89)
This repository holds the raw data. The method (build script, harness, launch units, write-up) lives on GitHub:
https://github.com/zhaoyilun/bonsai2-27b-mtp-repro
Both are the same piece of work: the GitHub repo has the code and the how-to, this dataset has the
measurements it produced. Cross-linked in both directions.
Raw measurements, scripts and notes for the two discussions:
official model… See the full description on the dataset page: https://huggingface.co/datasets/zhaokeqi/bonsai2-27b-mtp-repro.bonito-experiment
Dataset Card for bonito-experiment
bonito-experiment is a collection of datasets from experiments conducted
in Learning to Generate Instruction Tuning Datasets for
Zero-Shot Task Adaptation. We publish this collection to allow for the easy reproduction of these experiments.
from datasets import load_dataset
dataset = load_dataset("BatsResearch/bonito-experiment", "bonito_pubmed_qa")
Dataset Details
Dataset Description
Curated by: Nihal Nayak… See the full description on the dataset page: https://huggingface.co/datasets/BatsResearch/bonito-experiment.SURDS_evalThis data is sourced from the image of the nuScenes dataset. We extend our gratitude for their outstanding work!
numina-math-llama-3.1-8b-bon-meta-cotbonsai-jetson-benchmark-25w
Bonsai Jetson Benchmark — 25W
Platform: NVIDIA Jetson Orin Nano Super 8GB · Power mode: 25W
Status: Partial — 36 combos
Bonsai All-Model Benchmark — Jetson Orin Nano Super 8GB — 25W
Date: 2026-05-27 22:38Backend: llama.cpp (build-jetson) / CUDA / -ngl 99Platform: NVIDIA Jetson Orin Nano Super 8GB (6-core Cortex-A78AE + Ampere GPU)Sweep: prompt ∈ {256, 512, 1024, 2048} tok × gen ∈ {128, 256, 512} tok × 20 reqs/comboKey metric: tok/J = output tokens per second ÷… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/bonsai-jetson-benchmark-25w.RB2-BoN-GenerationsBonaFidefull-math-private-Qwen3-4B-Instruct-2507-bonbulk-v2-png-logos-mockup-outputsbone-fracture-7fylg
Bone Fracture 7Fylg
This dataset is part of the Roboflow 100 benchmark, a diverse collection of 100 object detection datasets spanning 7 imagery domains.
Dataset Statistics
Split
Images
Train
326
Validation
88
Test
44
Total
458
Classes (4)
angle
fracture
line
messed_up_angle
Usage
With LibreYOLO
from libreyolo import LIBREYOLO
# Load a model
model = LIBREYOLO(model_path="libreyoloXnano.pt")
# Train on this dataset… See the full description on the dataset page: https://huggingface.co/datasets/LibreYOLO/bone-fracture-7fylg.
