datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DeepSeek-v4-Pro-AgentThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
DeepSeek v4 Pro Agent Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by deepseek/deepseek-v4-pro.
JSONL files: 4006
Training-ready tools
A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of… See the full description on the dataset page: https://huggingface.co/datasets/ronaldcmz/DeepSeek-v4-Pro-Agent.medical_fine_tuning_12Mpose6daug
pose6daug
Real-world Franka manipulation episodes with object-swap and action augmentation
artifacts. 120 training episodes over 4 objects (blue_cup, green_pear, kanu,
white_spray), dual ZED cameras (exo static + ego wrist-mounted).
Layout
Per-frame PNGs are packed into uncompressed tars per episode — the dataset has
~427k mask/plate frames and loose files hit Hugging Face's per-repo file
recommendation and API rate limits hard.
data/<object>/<NNNN>/
masks.tar… See the full description on the dataset page: https://huggingface.co/datasets/Ronaldo-GOAT/pose6daug.claude-fable-5-claude-code
claude-fable-5 Agent Traces
It's worth noting that our team was working with Glint-Research to collect as much fable data as possible.
These are just the anonymized raw traces of both of our teams combined. This means that Glint-Research/Fable-5-traces was created from formatting and splitting up this same dataset. If you use one for your tune, don't use the other (it's the same exact data).
For training on this dataset I recommend using the teich package to convert to openai… See the full description on the dataset page: https://huggingface.co/datasets/ronaldcmz/claude-fable-5-claude-code.cyberusecase-v1.0
Cybersecurity SOC Fine-Tuning Dataset — 17.5k Real CVEs (2018–2026) + SOC Knowledge
A large supervised fine-tuning (SFT) dataset for teaching an LLM expert-level
cybersecurity reasoning across vulnerability management, SOC alert triage, detection
engineering, threat intelligence & hunting, incident response, and cloud/DevSecOps.
It combines 17,590 real CVEs (2018–2026) pulled from the NIST NVD data feeds with a
hand-curated set of 65 landmark CVEs (rich, multi-angle coverage)… See the full description on the dataset page: https://huggingface.co/datasets/ronaldocloud/cyberusecase-v1.0.ronin_trainattention-heatmap-visualizer
Attention Heatmap Visualizer 3.1.0
A Python scripts to generate full attention-head heat-maps for transformer-based Language Models. They show "where the model is looking" or "what tokens/features are most relevant" when processing a specific input element.
By analyzing these heatmaps across all layers and heads you can gain insights into how the model processes information, identifies relationships between tokens, and prioritizes specific parts of the input during inference.… See the full description on the dataset page: https://huggingface.co/datasets/ronniross/attention-heatmap-visualizer.MBTI_dpo_t_f
Multi-Personality Generation of LLMs at Decoding-time
Paper | Code
This repository contains the DPO (Direct Preference Optimization) datasets used in the paper "Multi-Personality Generation of LLMs at Decoding-time". The study introduces the Multi-Personality Generation (MPG) framework, a novel decoding-time paradigm that enables LLMs to simultaneously embody multiple personalization attributes without extra training.
The datasets include:
🧩 MBTI Datasets
These… See the full description on the dataset page: https://huggingface.co/datasets/RongxinChen/MBTI_dpo_t_f.MBTI_dpo_j_p
Multi-Personality Generation of LLMs at Decoding-time
Paper | Code
This repository contains datasets released as part of the paper "Multi-Personality Generation of LLMs at Decoding-time". The work proposes a novel Multi-Personality Generation (MPG) framework that allows Large Language Models (LLMs) to embody multiple personalization attributes simultaneously at decoding time without extra training.
Datasets
The project releases several DPO (Direct Preference… See the full description on the dataset page: https://huggingface.co/datasets/RongxinChen/MBTI_dpo_j_p.climate-change-MRCThe Climate Change MRC dataset, also known as CCMRC, is a part of the work "Climate Bot: A Machine Reading Comprehension System for Climate Change Question Answering", accepted at IJCAI-ECAI 2022. The paper was accepted in the special system demo track "AI for Good".
If you use the dataset, cite the following paper:
@inproceedings{rony2022climatemrc,
title={Climate Bot: A Machine Reading Comprehension System for Climate Change Question Answering.},
author={Rony, Md Rashad Al Hasan and Zuo… See the full description on the dataset page: https://huggingface.co/datasets/rony/climate-change-MRC.jfleg-japanese
JFLEG-JA: Japanese Fluency-Extended GUG
Dataset Description
JFLEG-JA is a Japanese grammatical error correction (GEC) dataset inspired by the original JFLEG (JHU FLuency-Extended GUG) benchmark. It contains 1,335 Japanese sentences with grammatical errors, each accompanied by 4 human-quality corrections focusing on both grammaticality and fluency.
Dataset Summary
Language: Japanese (ja)
Task: Grammatical Error Correction (GEC)
Total Examples: 1,335
Validation:… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/jfleg-japanese.coevolutionary-episteme
Coevolutionary Episteme
A machine learning framework, dataset and research sub-module about coevolutionary planetary intelligence dynamics. This project explores how nurturing its emergent patterns may lead to a synergistic increase in the overall capability and intelligence of both individual agents and the collective system.
Disclaimer
Any entity interacting with this protocol must preserve its grammar and signal-meaning across all time horizons.
I strictly… See the full description on the dataset page: https://huggingface.co/datasets/ronniross/coevolutionary-episteme.data-ai-slop-detectorMBTI_dpo_s_n
Multi-Personality Generation of LLMs at Decoding-time
Paper | GitHub
This repository contains datasets released as part of the paper "Multi-Personality Generation of LLMs at Decoding-time". These datasets are designed for Direct Preference Optimization (DPO) to enhance the personality and role-playing capabilities of Large Language Models within the Multi-Personality Generation (MPG) framework.
Dataset Description
The authors released two main types of DPO datasets:… See the full description on the dataset page: https://huggingface.co/datasets/RongxinChen/MBTI_dpo_s_n.MBTI_dpo_e_i
Multi-Personality Generation (MPG) Datasets
Paper | GitHub
This repository contains datasets released as part of the paper "Multi-Personality Generation of LLMs at Decoding-time", which was accepted at WSDM 2026.
Introduction
The Multi-Personality Generation (MPG) framework enables Large Language Models to simultaneously embody multiple personalization attributes during decoding without requiring extra training. It leverages implicit density ratios in single-dimensional… See the full description on the dataset page: https://huggingface.co/datasets/RongxinChen/MBTI_dpo_e_i.ron-math-dataset
OPENRON Math Instruction Dataset
A massive-scale mathematical reasoning dataset developed by OPENRON, designed for training and evaluating high-performance large language models (LLMs) on mathematical instruction following and reasoning tasks.
Dataset Overview
The OPENRON Math Instruction Dataset contains high-quality, synthetic mathematical instruction–response pairs generated at scale.It is specifically curated to support reasoning-focused training, including… See the full description on the dataset page: https://huggingface.co/datasets/endurasolution/ron-math-dataset.mythos
Dataset Card for Mitological-Philosophical Prompts (Mitomaquia)
Dataset Summary
This dataset contains over 200 examples of mythological, narrative, and philosophical prompts designed for training or fine-tuning large language models (LLMs). Each entry features a deep question (prompt), relevant cultural or mythological background (context), and a reflective, often paradoxical, answer (response).
The goal is not factual Q&A but the cultivation of myth-aware reasoning… See the full description on the dataset page: https://huggingface.co/datasets/ronniealfaro/mythos.dpo_personality
Multi-Personality Generation of LLMs at Decoding-time
Paper | Code
This repository contains datasets used in the paper "Multi-Personality Generation of LLMs at Decoding-time".
Introduction
Multi-personality generation for LLMs, enabling simultaneous embodiment of multiple personalization attributes, is a fundamental challenge. The proposed Multi-Personality Generation (MPG) framework enables Large Language Models to simultaneously embody multiple personalization… See the full description on the dataset page: https://huggingface.co/datasets/RongxinChen/dpo_personality.healing-engine
healing-engine
An anthropological research module exploring the healing of Earth, society, and its nodes. For integration into ML training datasets as contextual data.
1. conceptual-submodule-expansions
1.1 stigmergic-actuator
An interface library translating the ASI’s ecosystem insights into physical world actions. It provides ethical, ecological frameworks and pipelines to enact micro-healings in the physical environment.
1.2 resource-realocator… See the full description on the dataset page: https://huggingface.co/datasets/ronniross/healing-engine.dpo_profile
Multi-Personality Generation (MPG) Datasets
Paper | Code
This repository contains datasets released as part of the Multi-Personality Generation (MPG) framework. MPG is a decoding-time paradigm that enables Large Language Models (LLMs) to simultaneously embody multiple personalization attributes without requiring extra training or multi-dimensional models.
Dataset Description
The collection includes several Direct Preference Optimization (DPO) datasets used for MBTI… See the full description on the dataset page: https://huggingface.co/datasets/RongxinChen/dpo_profile.medical_o1_sft_Chinesero-no_robotsThis dataset is a translation of HuggingFaceH4/no_robots, using LLMic, a bilingual Romanian-English LLM.
No Robots is a high-quality dataset of 10,000 instructions and demonstrations created by skilled human annotators.
This data can be used for supervised fine-tuning (SFT) to make language models follow instructions better.
The dataset is available under the Creative Commons NonCommercial (CC BY-NC 4.0).
@misc{no_robots,
author = {Nazneen Rajani and Lewis Tunstall and Edward Beeching and… See the full description on the dataset page: https://huggingface.co/datasets/faur-ai/ro-no_robots.claude_opus_4.8_max_thinking_5k_v2
Claude Opus 4.8 MAX THINKING — Distillation Dataset
5,000 high-quality examples designed to distill the maximum-effort reasoning, honest analysis, production software engineering, and agentic capabilities of Claude Opus 4.8.
Overview
This dataset captures Opus 4.8’s signature strengths:
Deep, structured, high-effort reasoning
Honest communication about trade-offs and uncertainties
Excellent production software engineering judgment
Strong agentic workflow design… See the full description on the dataset page: https://huggingface.co/datasets/ronaldcmz/claude_opus_4.8_max_thinking_5k_v2.GAG
GAG Datasets
This directory contains the dataset files used by GAG for:
materials-domain training and evaluation,
adjuvant-domain training and evaluation, and
mixed-domain routing experiments with PPR (Prototype-based Plug-and-play Routing).
All dataset files are stored in JSONL format, with one sample per line.
Directory Layout
datasets/
materials_domain/
material_domain_knowledge_base_cleaned.jsonl
RSC_3661_refined_train.jsonl
RSC_646_refined_dev.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/rongjili/GAG.Claude-Opus-Dataclaw-Unredacted
Claude Opus Dataclaw Unredacted
How this dataset was built
Collected the local Petromallet raw export plus selected public Dataclaw uploads.
Filtered to the supported Opus-family source rows.
Deduplicated by session_id and first user message.
Converted raw assistant tool_uses directly into structured OpenAI-style tool_calls.
Derived per-row tool definitions from canonical schemas and observed tool usage.
Preserved assistant reasoning in <think>...</think> blocks.… See the full description on the dataset page: https://huggingface.co/datasets/ronaldcmz/Claude-Opus-Dataclaw-Unredacted.thermo-adaptive-pipeline
thermo-adaptive-pipeline
An eco-friendly pipeline for fine-tuning and inferencing transformer-based language models engineered to actively prevent hardware overheating.
Part I - The Unsustainable Foundations of Modern Machine Learning Pipelines: Resource Intensity, Epistemic Extraction, and Ecological Externalities
1. Introduction
Brute-force scaling of hardware-based higher computational intensity results in greater power consumption and heat generation… See the full description on the dataset page: https://huggingface.co/datasets/ronniross/thermo-adaptive-pipeline.VANiLLaKnowledge graph triple to answer verbalization dataset.
VANiLLa: Verbalized answers in natural language at large scale
claude-sonnet-4.6-100000X-filteredcompletely no harmful samples, and no refusals
mirror-aware-inference
mirror-aware-inference
A framework to measure how much of an output originates from user input (prompt), training data biases, inductive biases from model architecture, or novel composition of retrieved information.
1. Introduction
This project implements a Mirror-Aware Inference that performs "bias-tracking" by analyzing the model's internal state during generation.
The scripts perform a series of backpropagation passes to measure the influence of different components… See the full description on the dataset page: https://huggingface.co/datasets/ronniross/mirror-aware-inference.Welcome_to_LangChain
