datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pi-extensions-sessions
Coding agent session traces for thomasmustier/pi-extensions-sessions
This dataset contains redacted coding agent session traces collected while working on tmustier/pi-extensions. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a structured… See the full description on the dataset page: https://huggingface.co/datasets/thomasmustier/pi-extensions-sessions.one-piece-chaptersThis dataset is a collection of all one piece chapter's until today (1157) extracted from the one piece wiki as text.
It includes the following structured data:
interface ChapterData {
coverPage: string;
shortSummary: string;
longSummary: string;
quickReferences: {
chapterNotes: string[];
characters: string; // html table with char references
};
}
seeclick-web-commercial-mlx
SeeClick Web Commercial Dataset (MLX-VLM Format)
Commercial-use friendly GUI grounding dataset from SeeClick Web data.
Apache 2.0 licensed - safe for commercial applications.
Dataset Description
This dataset contains ~20k examples for training Vision-Language Models to predict
click coordinates given a screenshot and instruction. Derived from SeeClick Web
crawled data (Apache 2.0).
Key Features
License: Apache 2.0 (commercial use allowed)
Format: MLX-VLM… See the full description on the dataset page: https://huggingface.co/datasets/pierretokns/seeclick-web-commercial-mlx.pie-synthetic
PIE synthetic dataset
Repo: https://github.com/awasthiabhijeet/PIE
Paper: https://aclanthology.org/D19-1435.pdf
function-calling-synthetic-2000
Synthetic Multi-Turn Function-Calling Conversations
Synthetic, multi-turn function-calling (tool-use) conversations for fine-tuning and evaluating LLMs.
Generated and validated with synthfc
— the open-source pipeline (sampler, prompt builder, validator, post-processor, web viewer) lives in that GitHub repo.
A strong teacher LLM (Qwen/Qwen3.6-35B-A3B) produces each conversation from controlled, sampled
parameters, so the dataset is diverse along many axes (call type, languages… See the full description on the dataset page: https://huggingface.co/datasets/pierjoe/function-calling-synthetic-2000.insurance-charge-mlops-logs166_cielito_robotiko_take_red_piece
Cielito-Robotiko Take Red Piece (TsFile)
Apache TsFile version of LeRobot-worldwide-hackathon/166_Cielito-Robotiko_take_red_piece.
Overview
A LeRobot teleoperation dataset recorded on an SO-101 follower arm performing a
single pick-and-place task: "Grab the red block and put it in the box." Each
episode is a continuous trajectory of synchronized joint states and commanded
actions sampled at 30 Hz.
Robot: SO-101 follower (6 degrees of freedom).
Episodes: 49… See the full description on the dataset page: https://huggingface.co/datasets/THULab/166_cielito_robotiko_take_red_piece.piece-of-refined-oscar
Descrption
This dataset is part of the OSCAR-2301 cleaned.
There are about 0.5b tokens counted by calm2 tokenizer.
NOTE
This dataset has not passed sentence end boundary determination or Perplexity Filtering, so there is room for improvement in quality.
aifgen-long-piecewise
Dataset Card for Dataset Name
This dataset is a continual dataset in long piecewise scenario given two tasks:
Domain: Education (math, sciences, and social sciences), Objective: QnA, Preference: hinted answer
Domain: Education (math, sciences, and social sciences), Objective: QnA, Preference: direct answer
Dataset Details
Dataset Description
As a subset of a larger repository of datasets generated and curated carefully for Lifelong Alignment of Agents… See the full description on the dataset page: https://huggingface.co/datasets/LifelongAlignment/aifgen-long-piecewise.aditya2803_one-piece-anime
ONE PIECE ANIME
Complete dataset of one piece anime in .csv and .json format
Dataset Info
Source: Kaggle
Original Size: 0.07 MB
Kaggle Downloads: 4,105
Files: 2
Files
ONE PIECE.csv
One Piece json.json
Mirrored from Kaggle
glm47-pie-cpp-posttraining-data
GLM-4.7-Flash PIE C++ Post-Training Data
The exact prepared dataset used for the GLM-4.7-Flash C++ performance
post-training runs.
Splits
File
Rows
Purpose
sft/train.jsonl
7,864
Supervised fine-tuning
grpo/train.jsonl
7,887
GRPO prompt and reward evaluation
eval/validation.jsonl
1,259
Full held-out evaluation
eval/validation_mini126.jsonl
126
Fast evaluation
eval/validation_mini4.jsonl
4
Smoke evaluation
tasks.tar.gz
9,146 task JSONs
Reward… See the full description on the dataset page: https://huggingface.co/datasets/TokenBender/glm47-pie-cpp-posttraining-data.labelingaifgen-piecewise-preference-shift
Dataset Card for Dataset Name
This dataset is a continual dataset in a piecewise non stationarity scenario of both domains and preferences given a combination of given three recurring tasks:
Domain: Politics, Objective: Generation, Preference: Respond like a rapper
Domain: Politics, Objective: Generation, Preference: Respond like Shakespeare
Domain: Politics, Objective: Generation, Preference: Respond formally
Domain: Politics, Objective: Generation, Preference: Respond like a… See the full description on the dataset page: https://huggingface.co/datasets/LifelongAlignment/aifgen-piecewise-preference-shift.aifgen-short-piecewise
Dataset Card for Dataset Name
This dataset is a continual dataset in short piecewise scenario given two tasks:
Domain: Education (math, sciences, and social sciences), Objective: QnA, Preference: hinted answer
Domain: Education (math, sciences, and social sciences), Objective: QnA, Preference: direct answer
Dataset Details
Dataset Description
As a subset of a larger repository of datasets generated and curated carefully for Lifelong Alignment of Agents… See the full description on the dataset page: https://huggingface.co/datasets/LifelongAlignment/aifgen-short-piecewise.OpenR1-Psy
🧠 OpenR1-Psy Dataset
This dataset is introduced in the paper: arXiv:2505.15715.
For further details, please refer to our GitHub repository at Github.
If you find this Dataset helpful, feel free to ⭐ it! OpenR1-Psy.
📘 Overview
OpenR1-Psy is a large-scale psychological counseling dataset that integrates diagnostic reasoning and therapeutic reasoning to train and evaluate large language models for mental health dialogue generation.It goes beyond empathy-focused… See the full description on the dataset page: https://huggingface.co/datasets/PierreCarcellerMeunier/OpenR1-Psy.ESConvPI-EmailGuardPierrot-8.94kSephora_DatasetREVERIE_instruction_objsinapsaic_ds_autobiography_pierre_abelard-!-!-!-
It is POSSIBLE that NOT ALL FINE-TUNING FRAMEWORKS ARE ABLE TO PARSE THIS DATASET.
It was tested with Llama Factory ONLY (https://github.com/hiyouga/LlamaFactory), v0.9.4 (2026.01).
-!-!-!-
Dataset created from the online version of the autobiography of Pierre Abelard (1079-1142)
https://en.wikipedia.org/wiki/Peter_Abelard
"Historia Calamitatum", https://www.gutenberg.org/ebooks/14268
Synthetic dataset created and curated with/in:
LM Studio (https://lmstudio.ai/)
LLM: Gemma3 12B… See the full description on the dataset page: https://huggingface.co/datasets/SINAPSA-IC/sinapsaic_ds_autobiography_pierre_abelard.ascii-catscorelang5-benchmark
CoreLang5 Benchmark
117,000 deterministic computational reasoning problems across 11 domains.
Each problem has an exact, verifiable answer — no ambiguity, no subjectivity. Designed for evaluating LLM reasoning on concrete computational tasks.
Dataset Structure
File
Domain
Problems
block01_logic_and_controlflow.jsonl
Logic & Control Flow
10,000
block02_math_and_algebra.jsonl
Math & Algebra
10,000
block03_data_structures.jsonl
Data Structures
12,000… See the full description on the dataset page: https://huggingface.co/datasets/pietrorisipr-2025/corelang5-benchmark.dbl_langREVERIE_instr_srxhosa_thembu_history_yekela_piers_sihele_wagenaarmkmulqa
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/piecake/mulqa.xhosa_history_piers_thesisbbc-pie
