datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
home_assistant_train_ru
home_assistant_train_ru
Русифицированная версия датасета acon96/Home-Assistant-Requests-V2
(41 798 примеров). Предназначена для дообучения маленьких моделей Home Assistant
на русском языке (инструментальные вызовы, intent-классификация, управление устройствами).
Что переведено
User-запросы — полностью переведены на русский (ты-форма, неформально: «ты», не «вы»).
Текстовые ответы assistant (естественный язык, идущий клиенту) — переведены.
System-промпты, описания… See the full description on the dataset page: https://huggingface.co/datasets/RockMan256/home_assistant_train_ru.RocketEval-sLLMs
🚀 RocketEval 🚀
🚀 [ICLR '25] RocketEval: Efficient Automated LLM Evaluation via Grading Checklist
Github | OpenReview | Colab
This dataset contains the queries, generated checklist data, and responses data from 4 public benchmark datasets:
Dataset
No. of Queries
Comments
MT-Bench
160
Each 2-turn dialogue is split into 2 queries.
AlpacaEval
805
Arena-Hard
500
WildBench
1,000
To fit the context window of lightweight LLMs, we use a subset of WildBench including 1000… See the full description on the dataset page: https://huggingface.co/datasets/wjkim9653/RocketEval-sLLMs.ROCStories
ROCStories (prompt / continuation)
A reformatted version of the ROCStories
corpus, suitable for open-ended story-generation exercises.
Each example is a 5-sentence ROCStory, split into:
field
description
prompt
the first sentence of the story
continuation
the remaining four sentences
text
the full (unmodified) 5-sentence story
Splits
split
rows
train
70,676
validation
7,852
test
19,633
The train / validation split is a 90/10… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/ROCStories.rocstories-cloze
ROCStories & Story Cloze Test Dataset
This dataset contains the ROCStories corpus and the Story Cloze Test evaluation sets, originally released by the Story Cloze Test team at the University of Rochester.
jude-judaic-data
Dataset Card for Jude Judaic Data
This dataset contains a processed, clean Markdown version of the expansive Jewish text library sourced from the Sefaria project. It is specifically preprocessed step-by-step for use in Retrieval-Augmented Generation (RAG) pipelines and offline local AI assistants like Jude.
This section provides a description of the dataset fields, and additional information about the dataset structure such as criteria used to create the splits, relationships… See the full description on the dataset page: https://huggingface.co/datasets/RockyCo/jude-judaic-data.RocketReviews
RocketReviews Dataset
A structured dataset scraped from RocketReviews.com for use in AI/ML pipelines and vector databases.
Collection Status
Legend: [ ] not started · [~] in progress · [x] complete
Primary Tables
Table
Description
Script
Output
Status
reviews
Kit and product reviews with ratings and text sections
scripts/reviews/01_scrape.py
source/reviews/
[~]
flights
Member flight logs with conditions and notes
scripts/flights/01_scrape.py… See the full description on the dataset page: https://huggingface.co/datasets/rocketsmith/RocketReviews.task220_rocstories_title_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task220_rocstories_title_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task220_rocstories_title_classification.task105_story_cloze-rocstories_sentence_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task105_story_cloze-rocstories_sentence_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task105_story_cloze-rocstories_sentence_generation.rock-pollub-pl-books
ROCK Politechnika Lubelska PL books
A fail-closed Polish academic-book subset extracted from the official ROCK repository of Lublin University of Technology.
Retained books: 8
Text characters: 3,380,937
Tokens: 1,177,858 (cl100k_base proxy)
Author coverage: 100.0%
License: CC BY-SA 4.0, confirmed for every retained item and matched PDF bitstream
Source period: 2023-2026
The acquisition target was 20 books, but only eight passed the conservative per-file rights gate. The other… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/rock-pollub-pl-books.tiny_roc_storiesStories data for the paper [2505.05755] Insertion Language Models: Sequence Generation with Arbitrary-Position Insertions.
Project page: https://dhruveshp.com/projects/ilm
task219_rocstories_title_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task219_rocstories_title_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task219_rocstories_title_answer_generation.rocov2-modality-splits
ROCOv2 Modality-Specific Dataset Splits
Dataset Description
This dataset contains modality-specific splits of the ROCOv2 radiology dataset, organized and processed for training specialized medical image captioning models.
Dataset Summary
Total Samples: 1,000
Modalities: 5
Splits per Modality: train, validation, test
Random Seed: 42
Processing Date: 2025-08-31 12:52:59.233482
Modality Distribution
Modality
Samples
Percentage
CT
188
18.8%… See the full description on the dataset page: https://huggingface.co/datasets/WafaaFraih/rocov2-modality-splits.ro-camelaiThis dataset is a translation of camel-ai/math, camel-ai/chemistry, camel-ai/biology, camel-ai/physics
using LLMic, a bilingual Romanian-English LLM.
@misc{li2023camel,
title={CAMEL: Communicative Agents for "Mind" Exploration of Large Scale Language Model Society},
author={Guohao Li and Hasan Abed Al Kader Hammoud and Hani Itani and Dmitrii Khizbullin and Bernard Ghanem},
year={2023},
eprint={2303.17760},
archivePrefix={arXiv},
primaryClass={cs.AI}
}… See the full description on the dataset page: https://huggingface.co/datasets/faur-ai/ro-camelai.rocstories-merged
ROCStories Merged
A convenience variant of the ROCStories corpus where the five individual sentences (sentence1–sentence5) are concatenated into a single text column.
Two splits are available, matching the original ROCStories splits:
spring2016 — 45,496 stories
winter2017 — 52,665 stories
Each row contains storyid, storytitle, and text (the full story with sentences joined by a space).
cuda-to-rocm-wavefront-bugs
CUDA → ROCm Wavefront Bug Dataset
170 expert-curated examples of GPU kernel bugs that survive mechanical hipify translation and only manifest on AMD MI300X hardware (gfx942, wavefront-64).
Built for the ROCmPort AI project — a multi-agent pipeline that ports and optimizes CUDA kernels for AMD GPUs.
Why This Dataset Exists
hipify-perl and hipify-clang do a great job of mechanical API renaming (CUDA → HIP). But they cannot detect semantic bugs caused by AMD's larger… See the full description on the dataset page: https://huggingface.co/datasets/tazwarrrr/cuda-to-rocm-wavefront-bugs.taboo-rock
taboo-rock
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-rock")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
sepsis-icu-rock-cohort-100
HipAAsynth Dataset
Summary
This dataset is a validation artifact generated by HipAAsynth.
HipAAsynth is a deterministic testing and validation service that simulates real-world variability to evaluate how healthcare systems perform under deployment conditions.
Description
This dataset represents a controlled cohort used for testing and benchmarking.
HipAAsynth generates cohorts to simulate how conditions present across:
patient populations
demographic… See the full description on the dataset page: https://huggingface.co/datasets/HipAAsynth/sepsis-icu-rock-cohort-100.rocketraccoon_personality_alpacaAn attempt to imbue a gruff, RocketRaccoon like personality from GoG in the Rocket 3B model. Alpaca formatted dataset generated by ehartford_dolphin-2.2.1-mistral-7b.
sentence2paraphrasingrocmpilot-agent-sft
ROCmPilot Agent SFT
This dataset contains seed supervised fine-tuning examples for ROCmPilot, a multi-agent tool that helps developers migrate PyTorch and vLLM workloads from CUDA/NVIDIA assumptions to AMD ROCm readiness.
The examples teach ROCmPilot's production-facing agent behaviors:
repo_doctor: scan repository evidence for CUDA/NVIDIA assumptions
migration_planner: identify CUDA/NVIDIA migration blockers and recommend ROCm-safe fixes
patch_planner: convert findings into scoped… See the full description on the dataset page: https://huggingface.co/datasets/Shivam311/rocmpilot-agent-sft.
