datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nemotron-sft-code-focused-stage1-2-ChatML
Nemotron SFT Dataset (Chat Template Formatted)
Overview
This dataset is a curated supervised fine-tuning (SFT) dataset built from NVIDIA's Nemotron-Cascade-SFT-Stage-1 and Stage-2 datasets.
Important: This dataset uses the tokenizer's apply_chat_template() method to properly format conversations from the original messages/conversations fields.
Statistics
Total Samples: 50,000
Total Tokens: 415,605,764
Average Tokens per Sample: 8312.1
Tokenizer:… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nemotron-sft-code-focused-stage1-2-ChatML.nemotron-sft-benchmark-focused-stage1-2-ChatML-V1
Nemotron SFT Dataset (Chat Template Formatted)
Overview
This dataset is a curated supervised fine-tuning (SFT) dataset built from NVIDIA's Nemotron-Cascade-SFT-Stage-1 and Stage-2 datasets.
Important: This dataset uses the tokenizer's apply_chat_template() method to properly format conversations from the original messages/conversations fields.
Statistics
Total Samples: 50,000
Total Tokens: 428,330,639
Average Tokens per Sample: 8566.6
Tokenizer:… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nemotron-sft-benchmark-focused-stage1-2-ChatML-V1.SO101_FMB_FOCUS_01This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 1,
"total_frames": 351,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/zacapa/SO101_FMB_FOCUS_01.voice-focus-examples
Voice Focus Examples
Collection of examples for accurate foreground speaker transcription.
Details
Curated by: Joschka Wohlgemuth
Funded by: ai-coustics GmbH
Contact:
Web: https://ai-coustics.com
nemotron-sft-general-focused-stage1-2-ChatML-V2
Nemotron SFT Dataset (Chat Template Formatted)
Overview
This dataset is a curated supervised fine-tuning (SFT) dataset built from NVIDIA's Nemotron-Cascade-SFT-Stage-1 and Stage-2 datasets.
Important: This dataset uses the tokenizer's apply_chat_template() method to properly format conversations from the original messages/conversations fields.
Statistics
Total Samples: 100,000
Total Tokens: 222,951,633
Average Tokens per Sample: 2229.5
Tokenizer:… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nemotron-sft-general-focused-stage1-2-ChatML-V2.rough-focus-982a49
rough-focus-982a49
Synthetic sensors test data: 30 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/ZenithMind/rough-focus-982a49.calibrate-dp-dagger-v2-r3-two-focus-2026-02-17This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "franka",
"total_episodes": 40,
"total_frames": 8317,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:40"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ankile/calibrate-dp-dagger-v2-r3-two-focus-2026-02-17.2d-webmcp-browser-focus
2D WebMCP Browser Focus (Prerelease)
What this is
This is an early test of whether agents need useful tool results to complete an accessible browser task.
The agent must add a Retry step to a workflow, connect it correctly, and move keyboard focus to that new step. The test checks the real browser, not just the agent's final answer.
What happened
We ran each version 20 times with gpt-5-mini using low reasoning effort.
Tool result
Verified… See the full description on the dataset page: https://huggingface.co/datasets/accesslint/2d-webmcp-browser-focus.focus_persona_selection
Dataset Card for "focus_persona_selection"
More Information needed
test_camera_focus_20260716_100755This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/ItzSid55/test_camera_focus_20260716_100755.desk_companion_show_focus_sign_v1_20260802_215144This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/Vickiiyuan/desk_companion_show_focus_sign_v1_20260802_215144.black_tray_grasp_focus_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/tjdwlswkd21/black_tray_grasp_focus_v1.dataset_Focus_v2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 40,
"total_frames": 21799,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:40"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/smithasbellavi/dataset_Focus_v2.nemotron-sft-general-focused-stage1-2-ChatML-V3
Nemotron SFT Dataset (Chat Template Formatted)
Overview
This dataset is a curated supervised fine-tuning (SFT) dataset built from NVIDIA's Nemotron-Cascade-SFT-Stage-1 and Stage-2 datasets.
Important: This dataset uses the tokenizer's apply_chat_template() method to properly format conversations from the original messages/conversations fields.
Statistics
Total Samples: 496,385
Total Tokens: 1,114,218,401
Average Tokens per Sample: 2244.7
Tokenizer:… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nemotron-sft-general-focused-stage1-2-ChatML-V3.black_tray_grasp_focus_v1_20260707_211330This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/tjdwlswkd21/black_tray_grasp_focus_v1_20260707_211330.focus67-synth
FOCUS-67 teacher synth packs
Synthetic (input, output) examples generated by teacher LLMs for the FOCUS-67
subset of yuntian-deng/fuzzy-bench-gpt52-9m
(val split, shuffle seed 1234, first 512 rows, then the 67 curated
row_indices). Each spec is a natural-language function specification; the
teacher was asked to synthesize (input, output) pairs that obey the spec. These
are intended as training data for per-spec adapters (ProgramAsWeights).
Configs
synth (817,272… See the full description on the dataset page: https://huggingface.co/datasets/yuntian-deng/focus67-synth.dataset_FocusButton40This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 40,
"total_frames": 14190,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:40"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/smithasbellavi/dataset_FocusButton40.africa-somalia-igad-regional-focus-of-the-2024-global-report-on-food-cris-02a539d2
Igad Regional Focus of the 2024 Global Report On Food Cris | Africa (Food Security Information Network)
98 rows - 1 Africa country/area - 2024 - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 98 rows from Food Security Information Network, covering Igad Regional Focus of the 2024 Global Report On Food Cris. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-somalia-igad-regional-focus-of-the-2024-global-report-on-food-cris-02a539d2.dataset_FocusButtonThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 50,
"total_frames": 18527,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/smithasbellavi/dataset_FocusButton.africa-south-sudan-igad-regional-focus-of-the-2024-global-report-on-food-cris-02a539d2
Igad Regional Focus of the 2024 Global Report On Food Cris | Africa (Food Security Information Network)
98 rows - 1 Africa country/area - 2024 - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 98 rows from Food Security Information Network, covering Igad Regional Focus of the 2024 Global Report On Food Cris. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-south-sudan-igad-regional-focus-of-the-2024-global-report-on-food-cris-02a539d2.dataset_PressRedButton_FocusThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 140,
"total_frames": 73832,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:140"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/smithasbellavi/dataset_PressRedButton_Focus.africa-worldbank-wbl-supportive-framework-work-a-national-government-plan-or-strategy-focuses-on
WBL: Supportive Framework, Work, A national government plan or strategy focuses on women's access to the labor market | Africa (World Bank — Gender Statistics) | Africa (World Bank)
Size category: n<1K - Formats: parquet - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-worldbank-wbl-supportive-framework-work-a-national-government-plan-or-strategy-focuses-on.MULTI_VALUE_cola_acomp_focusing_like
Dataset Card for "MULTI_VALUE_cola_acomp_focusing_like"
More Information needed
nemotron-sft-general-focused-stage1-2-ChatML-V1
Nemotron SFT Dataset (Chat Template Formatted)
Overview
This dataset is a curated supervised fine-tuning (SFT) dataset built from NVIDIA's Nemotron-Cascade-SFT-Stage-1 and Stage-2 datasets.
Important: This dataset uses the tokenizer's apply_chat_template() method to properly format conversations from the original messages/conversations fields.
Statistics
Total Samples: 25,000
Total Tokens: 56,002,173
Average Tokens per Sample: 2240.1
Tokenizer: Qwen/Qwen3-0.6B… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nemotron-sft-general-focused-stage1-2-ChatML-V1.africa-eritrea-igad-regional-focus-of-the-2024-global-report-on-food-cris-02a539d2
IGAD Regional Focus of the 2024 Global Report on Food Crises | Africa (Eritrea official open data)
98 rows - 1 Africa country - 2024 - Repackaged by Electric Sheep Africa
TL;DR
This dataset packages one official XLSX resource from Eritrea as
ML-ready Parquet. The source file is the provenance boundary; all usable
indicators or tabular columns from the resource stay together in this repo.
About the source
Source: IGAD Regional Focus of the 2024… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-eritrea-igad-regional-focus-of-the-2024-global-report-on-food-cris-02a539d2.MULTI_VALUE_qqp_acomp_focusing_like
Dataset Card for "MULTI_VALUE_qqp_acomp_focusing_like"
More Information needed
africa-djibouti-igad-regional-focus-of-the-2024-global-report-on-food-cris-02a539d2
IGAD Regional Focus of the 2024 Global Report on Food Crises | Africa (Djibouti official open data)
98 rows - 1 Africa country - 2024 - Repackaged by Electric Sheep Africa
TL;DR
This dataset packages one official XLSX resource from Djibouti as
ML-ready Parquet. The source file is the provenance boundary; all usable
indicators or tabular columns from the resource stay together in this repo.
About the source
Source: IGAD Regional Focus of the… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-djibouti-igad-regional-focus-of-the-2024-global-report-on-food-cris-02a539d2.focus_test
Dataset Card for "focus_test"
More Information needed
MULTI_VALUE_wnli_acomp_focusing_like
Dataset Card for "MULTI_VALUE_wnli_acomp_focusing_like"
More Information needed
MULTI_VALUE_sst2_acomp_focusing_like
Dataset Card for "MULTI_VALUE_sst2_acomp_focusing_like"
More Information needed
