datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RSCC-RSEdit-Test-Split
RSCC-RSEdit-Test-Split
This directory contains the test split for RSCC-RSEdit dataset.
Directory Structure
RSCC-RSEdit-Test-Split/
├── images/ # Original images (676 PNG files)
├── masks/ # Original grayscale masks (338 PNG files)
│ └── [mask files with pixel values 0,1,2,3,4]
├── masks_colorful/ # Colorful RGBA visualization masks (338 PNG files)
│ └── [same filenames as masks/, but in RGBA format with colors]
├──… See the full description on the dataset page: https://huggingface.co/datasets/BiliSakura/RSCC-RSEdit-Test-Split.processed_test_splitsunreal-engine-5-code-split
Dataset Card for unreal-engine-5-code-split
Using the unreal-engine-5-code hf dataset by AdamCodd, I split the data into smaller chunks for RAG systems
Dataset Details
Dataset Description
Branches
main:
Dataset is split/chunked by engine module (no max chunk size)
chunked-8k:
Data is first split by module {module_name}.jsonl if ≤ 8000 tokens.
If a module is ≥ 8000 tokens then it's further split by header name {module_name}_{header_name}.jsonl
If a header… See the full description on the dataset page: https://huggingface.co/datasets/olympusmonsgames/unreal-engine-5-code-split.aiact-frozen-split-harness
EU AI Act scenarios — frozen split harness
EU AI Act deployment scenarios with their obligations, as a frozen split.
Each row of scenarios.jsonl carries role (Provider / Deployer), intended_use, system_type,
input_data, domain, a related_articles list of AI Act article numbers, and the obligations that
follow. results/ holds the run outputs from the harness passes that used this split.
The live board is the authority
GET https://councilof.ai/api/gspc — quote… See the full description on the dataset page: https://huggingface.co/datasets/csoai/aiact-frozen-split-harness.train_splits_helmContains the following train split from datasets in helm:
big bench
mmlu
TruthfulQA
cnn/dm
gsm
bbq
boolq
NarrativeQA
QuAC
math
bAbI
Each prompt has <= 5 in-context samples along with a sample, all of which from the train set of the respective datasets.
tofu_custom_split_SISAsharegpt_v3_unfiltered_cleaned_splitsommelier-xlam-single-call-splits
sommelier xlam single-call splits
Deterministic, deduplicated, single-tool-call train/validation/test
splits derived from
Salesforce/xlam-function-calling-60k
(APIGen, CC-BY-4.0), produced by the
sommelier pipeline for
reproducible tool-calling fine-tuning. These are the exact splits used to
train and evaluate
abdelstark/llama-3.1-nemotron-nano-8b-xlam-tool-calling-lora.
Why single-call
The upstream dataset mixes single-call and multi-call examples (~52.6%… See the full description on the dataset page: https://huggingface.co/datasets/abdelstark/sommelier-xlam-single-call-splits.atomic-metrics-rm-splits
Atomic Metrics RM Task Splits
Preference-pair benchmark splits used by
Atomic Metrics. The release
contains four open-ended task families derived from public SHP, OASST1, and
OASST2 preference data.
Dataset structure
Each configuration contains 10,000 training pairs and 2,000 test pairs. Every
row has:
{
"sample_id": "source-specific stable ID",
"source_dataset": "shp | oasst1 | oasst2",
"category": "task configuration",
"split": "train | test"… See the full description on the dataset page: https://huggingface.co/datasets/tintin1027/atomic-metrics-rm-splits.Split-IFEval
Split IFEval
This dataset modifies the Instruction-Following Eval (IFEval) benchmark to split apart the task from the syntactic instructions in addition to fixing errors in the original dataset.
It enables the use of research methods like attention steering that require access to the instruction text.
To load the dataset, run:
from datasets import load_dataset
split_ifeval = load_dataset("ibm-research/Split-IFEval")
Dataset Structure
Each entry in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/Split-IFEval.file-scorer-10-04-filescore-splitsmdcl-splitsairoboros-3.2-splitsplit_small_smallrobomme_1cuben_fixedcup_split
robomme_1cuben_fixedcup_split (VideoUnmaskSwap1CubeN — fixed cups, disjoint split)
A single red cube is hidden under one of three cups; the cups are shuffled a variable
number of times (0..3) and the robot must pick the cup now hiding the cube. Prompt is
color-free: "watch the video carefully, then pick up the container hiding the cube".
Derived from robomme_1cuben_allcases,
with two changes for a clean generalization study:
1. Fixed cup locations. Cup-pose perturbation is… See the full description on the dataset page: https://huggingface.co/datasets/alfayoung/robomme_1cuben_fixedcup_split.megamath-web-pro-max-splittedchinese-fineweb-edu-v2_splitted_1_filtered_combinedhyperpartisan-longformer-split
Hyperpartisan news detection
This dataset has the hyperpartisan new dataset, processed and split exactly as it was for longformer experiments.
Code for processing was found at here.
Trinity-ToolAce-SFT-spliteuroparl_dbca_splitsgit-commit-message-splitterqwen3.7-max-split-formatted
Qwen Agent Thinking Online Distillation Rows
This dataset contains cumulative assistant-turn training rows prepared for online logit distillation of Qwen-style agent models, plus a small set of no-tools chat rows to reduce tool-call overbias.
Each row is a rendered-chat-ready conversation prefix ending at a target assistant turn. The trainer uses all prior messages as context and applies loss only to the final assistant span.
Dataset Details
Source trace repo:… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/qwen3.7-max-split-formatted.ovdsgg-action-genome-split
OvDSGG Action Genome Open-Vocabulary Split
This dataset repository contains the open-vocabulary category split used by OvDSGG for Action Genome experiments.
It does not redistribute Action Genome videos, frames, or full annotations. Users must obtain and process Action Genome separately, then use this split metadata to reproduce the OvDSGG open-vocabulary training/evaluation protocol.
Paper: https://huggingface.co/papers/2608.14835Code: https://github.com/jhelsby/OvDSGGModel… See the full description on the dataset page: https://huggingface.co/datasets/jhelsby/ovdsgg-action-genome-split.aya_collection_language_split-askllm-v1
aya_collection_language_split-askllm-v1
データセット CohereForAI/aya_collection_language_split に対して、 Ask-LLM 手法でスコア付けしたデータセットです。
元データセットのカラムに加え askllm_score というカラムが追加されており、ここに Ask-LLM のスコアが格納されています。
Ask-LLM でスコア付けに使用した LLM は Rakuten/RakutenAI-7B-instruct で、プロンプトは以下の通りです。
###
{data}
###
Does the previous paragraph demarcated within ### and ### contain informative signal for pre-training a large-language model? An informative datapoint should be well-formatted, contain some usable… See the full description on the dataset page: https://huggingface.co/datasets/geniacllm/aya_collection_language_split-askllm-v1.cyber-threat-intelligence-splitedkilt_wikipedia_splitagmind-rag-splitter-ru-data
RU Context-Aware Document Split
Датасет (teacher-distillation) для обучения русского context-aware сплиттера документов для RAG. Каждый пример учит модель где резать документ на самодостаточные смысловые чанки, держа таблицы и код целыми.
Использован для модели AGmind/agmind-rag-splitter-ru. Код генерации и обучения: github.com/botAGI/AGmind-ML.
Формат (Alpaca JSONL)
{
"instruction": "Раздели документ на смысловые части для системы поиска (RAG)...",
"input":… See the full description on the dataset page: https://huggingface.co/datasets/AGmind/agmind-rag-splitter-ru-data.app-review-extraction-splits
App Review Structured-Extraction Splits
Frozen train/val/test splits used to fine-tune and evaluate a LoRA adapter that
extracts a strict, closed-vocabulary JSON object from app-store reviews. These
are the exact artifacts behind the project's results — published so the
base-vs-tuned comparison is fully reproducible.
💻 Code + write-up: https://github.com/deshpandetanmay/qlora-structured-extraction
🤖 Adapter: https://huggingface.co/tanmaydeshpande/qlora-app-review-extraction… See the full description on the dataset page: https://huggingface.co/datasets/tanmaydeshpande/app-review-extraction-splits.WizardLM_evol_instruct_V2_196k_unfiltered_merged_spliteli5_splitThis dataset is the subset of original eli5 dataset available in hugging face space
