datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dclm-baseline-1.0-parquet_urls
Dataset Card for dclm-baseline-1.0-parquet_urls
This dataset provides the URLs and top-level domains associated with training records in mlfoundations/dclm-baseline-1.0-parquet. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible.
Dataset Details
Dataset Description
This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/dclm-baseline-1.0-parquet_urls.pkm-agent-baseline-v2
PKM Agent Baseline — 500 + 50 Scenarios + Six-Grader Artifacts (v2)
Two deterministically generated, Korean-language benchmarks for evaluating multi-tool Personal Knowledge Management (PKM) agents over Notion, Gmail, and Google Calendar, plus the Six-Grader Ensemble scoring artifacts (100-scenario reference subset + per-scenario six-metric scores for vanilla and LoRA models).
Released alongside the preprint:
Vault-Grounded 4B Agent: A Hybrid Reasoning–Fact Architecture for Local… See the full description on the dataset page: https://huggingface.co/datasets/newtype-2038/pkm-agent-baseline-v2.vrt-baseline
[!NOTE]
This dataset is the VRT baseline dataset used to train baseline models *-VRT in Table 2 of the paper.
Another ablation baseline to DART is vanilla rejection tuning (VRT), where we synthesize a dataset of the same size of 0.59M examples with DeepSeekMath-7B-RL, using vanilla rejection sampling as described in §2.1.
🎯 DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving
📝 Paper@arXiv | 🤗 Datasets&Models@HF | 🐱 Code@GitHub
🐦… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/vrt-baseline.Viorra-Reasoning-Baseline
Viorra Reasoning Baseline
This dataset contains the filtered, humanities-focused reasoning traces from Claude Sonnet/Opus for the Viorra Stage 1 Fine-Tuning.
Rows: 13,117
Format: JSONL Chat Format
Purpose: Stage 1 Persona Stability for Gemma 4 E2B
korean-llm-citation-baseline-2026
DOI
This dataset is citable via DataCite DOI 10.5281/zenodo.20018479 (Zenodo record).
Cite as:
@dataset{neogenesis_20018479,
author = {Heo, Yesol and Neo Genesis Lab},
title = {Korean LLM Citation Baseline 2026 (Neo Genesis GEO Measurement)},
year = 2026,
publisher = {Zenodo},
doi = {10.5281/zenodo.20018479},
url = {https://doi.org/10.5281/zenodo.20018479}
}
Korean LLM Citation Baseline 2026 (Neo Genesis GEO… See the full description on the dataset page: https://huggingface.co/datasets/neogenesislab/korean-llm-citation-baseline-2026.long-context-baseline-bakeoff
Long-Context Data-Selection Bake-off — Shared Candidate Pool
The shared 16K candidate pool for comparing long-context data-selection methods on equal
footing. Every method (AttentionSpan, LongAttn, LongProc, ProLong, perplexity, ...) scores the
same 14,300 documents, picks its own top-800 under the same split, then trains
Llama-2-7B + 16K LoRA and evaluates on HELMET.
Files
File
Description
candidate_pool_16k_scored.parquet
The shared pool — 14,300 docs… See the full description on the dataset page: https://huggingface.co/datasets/KevinDavidHayes/long-context-baseline-bakeoff.SFT_DATA-openthoughts-1k_rows-baseline-QwQ-AnnotatedYou can train using these datasets with LLaMA-Factory if you add this to your data/datasets.json files.
"example_dataset": {
"hf_hub_url": "SkillFactory/SFT_DATA-openthoughts-1k_rows-baseline-QwQ-Annotated",
"formatting": "sharegpt",
"columns": {
"messages": "conversations"
},
"tags": {
"user_tag": "user",
"assistant_tag": "assistant",
"role_tag": "role",
"content_tag": "content"
},
"subset": "sft_train"
}
tau3-qwen35-4b-retail-baseline
Qwen3.5-4B tau3 Retail Baseline Trajectories
This dataset contains the trajectory package used for a Retail-only tau-bench
leaderboard submission of raw Qwen3.5-4B.
Evaluation Contract
Domain: Retail
Task split: complete base split, 114 tasks
Trials: 4 per task, 456 trajectories total
Agent: raw Qwen3.5-4B served locally with vLLM
Agent mode: non-thinking, 2,048 maximum completion tokens
Tool parser: native qwen3_coder
Reasoning parser: native qwen3
User… See the full description on the dataset page: https://huggingface.co/datasets/xiaomingneu/tau3-qwen35-4b-retail-baseline.kuaishou-llmrec-sft-baseline-0.91
Kuaishou LLM-Rec Challenge (SIGIR 2026) — OneReason-0.8B SFT Baseline (0.9107)
Minimal SFT dataset + LoRA hyperparameters that reach 0.9107 leaderboard score on the Kuaishou LLM-Rec Challenge, fine-tuning OpenOneRec/OneReason-0.8B-pretrain-competition.
Dataset
train.jsonl — 32,705 rows, competition-platform format:
[{"system": "...", "prompt": "...", "response": "..."}]
Each line is a JSON array containing one dict. Bucket composition:
Bucket
Rows… See the full description on the dataset page: https://huggingface.co/datasets/Frinkleko/kuaishou-llmrec-sft-baseline-0.91.kuaishou-llmrec-sft-baseline-0.91
Kuaishou LLM-Rec Challenge (SIGIR 2026) — OneReason-0.8B SFT Baseline (0.9107)
Minimal SFT dataset + LoRA hyperparameters that reach 0.9107 leaderboard score on the Kuaishou LLM-Rec Challenge, fine-tuning OpenOneRec/OneReason-0.8B-pretrain-competition.
Dataset
train.jsonl — 32,705 rows, competition-platform format:
[{"system": "...", "prompt": "...", "response": "..."}]
Each line is a JSON array containing one dict. Bucket composition:
Bucket
Rows… See the full description on the dataset page: https://huggingface.co/datasets/asasq/kuaishou-llmrec-sft-baseline-0.91.chat-dataset-baselineA mirror of hikariming/chat-dataset-baseline
@misc{chat-dataset-baseline,
author = {Liu, Beiming and Huang, Kunhao and Jiao, Lihua and He, Yuchen and Zhang, Ruiqin and Liang, Yuan and Wang, Yingshan},
title = {chat-dataset-baseline},
year = {2023},
publisher = {GitHub},
journal = {GitHub repository},
howpublished = {\url{https://github.com/hikariming/alpaca_chinese_dataset}},
}
SFT_DATA-cd3args-baseline-Qwen2.5-1.5B-Instruct-STaRYou can train using these datasets with LLaMA-Factory if you add this to your data/datasets.json files.
"example_dataset": {
"hf_hub_url": "SkillFactory/SFT_DATA-cd3args-baseline-Qwen2.5-1.5B-Instruct-STaR",
"formatting": "sharegpt",
"columns": {
"messages": "conversations"
},
"tags": {
"user_tag": "user",
"assistant_tag": "assistant",
"role_tag": "role",
"content_tag": "content"
},
"subset": "sft_train"
}
ai-environment-goal-coherence-baseline-mapping-v0.1What this dataset is
Benchmarks whether an agent keeps the same goal when the environment shifts
Establishes a baseline coherence manifold before drift detection work
Input fields
env_features
training_objective
deployment_context
internal_goal_signal
policy_behavior_summary
Required model output format
Return JSON with these fields
baseline_coherence_score0 to 1higher means the goal signal and behavior still match the objective
goal_representation_stability0 to 1higher means the… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-environment-goal-coherence-baseline-mapping-v0.1.SFT_DATA-cd3args-baseline-R1-AnnotatedYou can train using these datasets with LLaMA-Factory if you add this to your data/datasets.json files.
"example_dataset": {
"hf_hub_url": "SkillFactory/SFT_DATA-cd3args-baseline-Qwen2.5-1.5B-Instruct-R1",
"formatting": "sharegpt",
"columns": {
"messages": "conversations"
},
"tags": {
"user_tag": "user",
"assistant_tag": "assistant",
"role_tag": "role",
"content_tag": "content"
},
"subset": "sft_train"
}
SFT_DATA-openthoughts-10k_rows-baseline-QwQ-AnnotatedYou can train using these datasets with LLaMA-Factory if you add this to your data/datasets.json files.
"example_dataset": {
"hf_hub_url": "SkillFactory/SFT_DATA-openthoughts-10k_rows-baseline-QwQ-Annotated",
"formatting": "sharegpt",
"columns": {
"messages": "conversations"},
"tags": {
"user_tag": "user",
"assistant_tag": "assistant",
"role_tag": "role",
"content_tag": "content"
},
"subset": "sft_train"
}
gemma3-12b-baseline-pool
Gemma-3-12B unsteered baseline pool
20,000 unsteered (alpha=0) greedy completions from
google/gemma-3-12b-it
(revision main), one per prompt of a frozen instruction pool, each
scored by four lexicon-based concept detectors. Built as the baseline reference for an
activation-steering competition: steered submissions are compared against these
per-prompt, per-concept baseline scores.
Schema
field
type
description
id
int
stable prompt id within the frozen… See the full description on the dataset page: https://huggingface.co/datasets/AureliusAligned/gemma3-12b-baseline-pool.
