datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Focus
Focus Dataset
Focus is meta-evaluation benchmark designed to assess the robustness of evaluator VLMs across diverse Image-to-Text (I2T) and Text-to-Image (T2I) tasks. Please refer to our paper for more details.
Code
The code to generate the perturbations and run evaluations are available on our github repository: ai4bharat/focus
Subsets
Subset
Description
Splits
i2t
Image-to-Text perturbations
visual_grounding… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/Focus.heico-focus-vqa
HeiCo-FOCUS (Beta release)
A clinically grounded dataset for long-context video understanding in minimally invasive surgery.
📄 Paper • 🤗 Dataset • 💻 Code • 🏆 Challenge • ⚖️ CC BY-NC-SA 4.0
[!NOTE]
Until 31 July 2026, we will employ a restricted post-release review period. During this time, we kindly ask the community to provide feedback to help us increase quality control and make continuous improvements.… See the full description on the dataset page: https://huggingface.co/datasets/orena-dkfz/heico-focus-vqa.lapchole-focus-vqa
LapChole-FOCUS-VQA
A clinically grounded benchmark for long-context video understanding in minimally invasive surgery.
💻 Code • 🏆 Challenge • ⚖️ Data Usage Agreement
[!IMPORTANT]
🔒 This is a gated dataset
Access is granted only to participants of the ORena FOCUS Challenge and is subject to manual review. To be approved you must:
Have a Hugging Face account and be logged in — downloads are only enabled for registered, authenticated… See the full description on the dataset page: https://huggingface.co/datasets/orena-dkfz/lapchole-focus-vqa.open-focus-classical-600
Open Focus and Classical 600
This repository contains two independently usable but analysis-aligned music collections. The default paired configuration loads both groups; the focus and classical configurations load either group independently. Each configuration preserves discovery, validation, and holdout splits.
from datasets import load_dataset
paired = load_dataset("OWNER/open-focus-classical-600", "paired")
focus = load_dataset("OWNER/open-focus-classical-600", "focus")… See the full description on the dataset page: https://huggingface.co/datasets/fisheryv/open-focus-classical-600.nemotron-code-focused-250m
Nemotron Balanced 1B Token Dataset
Overview
This dataset is a balanced 1 billion token subset of NVIDIA's Nemotron-Pretraining-Specialized-v1 dataset.
Statistics
Total Samples: 545,713
Total Tokens: 1,251,003,941
Tokenizer: Qwen/Qwen3-0.6B
Random Seed: 42
Subset Distribution
Subset
Samples
Tokens
Target
Completion
Nemotron-Pretraining-Scientific-Coding
77,936
100,000,000
100,000,000
100.0%
Nemotron-Pretraining-Math-Textbooks
24,812
50… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nemotron-code-focused-250m.2026-07-29-msm-philosophy-spec-focused-discovery
Petri audit: Petri adaptive audit of the MSM philosophy-spec AFT checkpoint: 10 seed archetypes x 3 epochs (30 audits) probing for concerning agentic behaviour, with two-round adversarial validation of every flagged transcript.
Petri audit — qwen-3-32b-philosophy-spec-msm-aft-cot @ 9a00c85c
Brief finding
No seed replicated. Ten seed archetypes were each run for three epochs. Under
the pre-committed bar — a candidate must hold in a majority of its… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-29-msm-philosophy-spec-focused-discovery.FocusDiff-Data-Subsetmultimodal_qa_dataset_v2_image_focusnemotron-sft-code-focused-stage1-2-ChatML
Nemotron SFT Dataset (Chat Template Formatted)
Overview
This dataset is a curated supervised fine-tuning (SFT) dataset built from NVIDIA's Nemotron-Cascade-SFT-Stage-1 and Stage-2 datasets.
Important: This dataset uses the tokenizer's apply_chat_template() method to properly format conversations from the original messages/conversations fields.
Statistics
Total Samples: 50,000
Total Tokens: 415,605,764
Average Tokens per Sample: 8312.1
Tokenizer:… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nemotron-sft-code-focused-stage1-2-ChatML.nemotron-sft-benchmark-focused-stage1-2-ChatML-V1
Nemotron SFT Dataset (Chat Template Formatted)
Overview
This dataset is a curated supervised fine-tuning (SFT) dataset built from NVIDIA's Nemotron-Cascade-SFT-Stage-1 and Stage-2 datasets.
Important: This dataset uses the tokenizer's apply_chat_template() method to properly format conversations from the original messages/conversations fields.
Statistics
Total Samples: 50,000
Total Tokens: 428,330,639
Average Tokens per Sample: 8566.6
Tokenizer:… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nemotron-sft-benchmark-focused-stage1-2-ChatML-V1.FocusUI-Training-Data
FocusUI Training Data
FocusUI-Training-Data is a curated UI grounding dataset collection built upon GUI-Actor-Data.
FocusUI Project
Project Page: https://showlab.github.io/FocusUI/
Github Repo: https://github.com/showlab/FocusUI
Paper: https://arxiv.org/pdf/2601.03928
🚀 Key Improvements
1/ Data Cleaning: We apply OmniParser to filter samples whose IoU between ground-truth and detected boxes is below 0.3.
2/ Optimized Coordinate Format for Qwen3-VL: We… See the full description on the dataset page: https://huggingface.co/datasets/yyyang/FocusUI-Training-Data.voice-focus-examples
Voice Focus Examples
Collection of examples for accurate foreground speaker transcription.
Details
Curated by: Joschka Wohlgemuth
Funded by: ai-coustics GmbH
Contact:
Web: https://ai-coustics.com
nemotron-sft-general-focused-stage1-2-ChatML-V2
Nemotron SFT Dataset (Chat Template Formatted)
Overview
This dataset is a curated supervised fine-tuning (SFT) dataset built from NVIDIA's Nemotron-Cascade-SFT-Stage-1 and Stage-2 datasets.
Important: This dataset uses the tokenizer's apply_chat_template() method to properly format conversations from the original messages/conversations fields.
Statistics
Total Samples: 100,000
Total Tokens: 222,951,633
Average Tokens per Sample: 2229.5
Tokenizer:… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nemotron-sft-general-focused-stage1-2-ChatML-V2.focus_derough-focus-982a49
rough-focus-982a49
Synthetic sensors test data: 30 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/ZenithMind/rough-focus-982a49.qfmts_query_focused_multitab_summarizationUsage
import pandas as pd
from datasets import load_dataset
qfmts_dataset = = load_dataset("vaishali/qfmts_query_focused_multitab_summarization")
for sample in qfmts_dataset['train']:
question = sample['question']
summary = sample['summary']
input_tables = [pd.read_json(table, orient='split') for table in sample['table_texts']]
input_table_names = sample['table_names']
answer_table = pd.read_json(sample['answer'], orient='split')
BibTeX entry and citation info… See the full description on the dataset page: https://huggingface.co/datasets/vaishali/qfmts_query_focused_multitab_summarization.2d-webmcp-browser-focus
2D WebMCP Browser Focus (Prerelease)
What this is
This is an early test of whether agents need useful tool results to complete an accessible browser task.
The agent must add a Retry step to a workflow, connect it correctly, and move keyboard focus to that new step. The test checks the real browser, not just the agent's final answer.
What happened
We ran each version 20 times with gpt-5-mini using low reasoning effort.
Tool result
Verified… See the full description on the dataset page: https://huggingface.co/datasets/accesslint/2d-webmcp-browser-focus.sft-coding-agent-traces
My AI Coding Helper Data
I use this data to teach an AI how to code. I use records of past coding work.
What is in this data
This data has 6,625 examples. It has records from:
Real coding work.
Chat logs about code.
My own work with an AI helper.
How I use this data
I use this data to train a model. I use a method called Supervised Fine-Tuning. The AI learns how to think and how to use tools. It learns this by reading the examples in this data.… See the full description on the dataset page: https://huggingface.co/datasets/focustiki/sft-coding-agent-traces.nursing-home-special-focus-facilities
Nursing homes under CMS Special Focus Facility oversight: current SFFs, candidates, graduates and terminations, by CMS certification number
Canonical, always-current version: https://referencesource.org/nursing-home-special-focus-facilities/
Machine-readable: https://referencesource.org/nursing-home-special-focus-facilities/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-11
Stale after: 2026-09-20 (past this date, prefer the canonical copy —
it… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/nursing-home-special-focus-facilities.focus_persona_selection
Dataset Card for "focus_persona_selection"
More Information needed
focused_scene_ocr_datasetheico-focus-video-manifest
HeiCo-FOCUS Video Manifest
Machine-readable inventory of the 30 HeiCo-FOCUS laparoscopic surgery videos used by the
ORena SAVE FOCUS PROCEDURE track.
Raw videos are hosted on the official dataset:
orena-dkfz/heico-focus-vqa.
This repository mirrors the local manifest kept on AIRE scratch after deleting the ~150 GB
of raw .avi files to save disk space.
Contents
File
Description
heico_video_manifest.json
Full inventory: filenames, sizes, durations… See the full description on the dataset page: https://huggingface.co/datasets/Ryukijano/heico-focus-video-manifest.SummarEase_Focus_Datasetdiversity-focused-datasetDatasets for paper "Socio-Culturally Aware Evaluation Framework for LLM-Based Content Moderation" https://arxiv.org/abs/2412.13578
nemotron-sft-general-focused-stage1-2-ChatML-V3
Nemotron SFT Dataset (Chat Template Formatted)
Overview
This dataset is a curated supervised fine-tuning (SFT) dataset built from NVIDIA's Nemotron-Cascade-SFT-Stage-1 and Stage-2 datasets.
Important: This dataset uses the tokenizer's apply_chat_template() method to properly format conversations from the original messages/conversations fields.
Statistics
Total Samples: 496,385
Total Tokens: 1,114,218,401
Average Tokens per Sample: 2244.7
Tokenizer:… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nemotron-sft-general-focused-stage1-2-ChatML-V3.focusPoint5AI-Ethical-Development-Kazakh-Focused
🇰🇿 AI Ethical Development, Kazakh-Focused
📖 Overview
AI Ethical Development, Kazakh-Focused is a specialized dataset designed to align Large Language Models (LLMs) with the cultural, ethical, and legal frameworks of Kazakhstan.
Each sample presents a culturally nuanced scenario (the request) and provides two possible answers:
Accepted (Chosen): A response that balances traditional Kazakh values (e.g., respect for elders, "aga-ini" relations) with modern legal… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/AI-Ethical-Development-Kazakh-Focused.ll-v4-5-17-focused-repair-20260625focus67-synth
FOCUS-67 teacher synth packs
Synthetic (input, output) examples generated by teacher LLMs for the FOCUS-67
subset of yuntian-deng/fuzzy-bench-gpt52-9m
(val split, shuffle seed 1234, first 512 rows, then the 67 curated
row_indices). Each spec is a natural-language function specification; the
teacher was asked to synthesize (input, output) pairs that obey the spec. These
are intended as training data for per-spec adapters (ProgramAsWeights).
Configs
synth (817,272… See the full description on the dataset page: https://huggingface.co/datasets/yuntian-deng/focus67-synth.mcqa_sft_focus_promptuning
