datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fractus-datasets
Fractus Datasets — the neuroscience-grounded training corpus
A proprietary, neuroscience-derived training corpus for the Fractus Continuous Thought Engine — ~3–4B tokens mapping real brain mechanisms to software/AI architecture, plus cognitive skills, code, esoteric tradition, and lexical knowledge.
Curator: Philippe-Antoine Robert · rpa.tu@proton.me · 2026
What this dataset collection IS
Fractus is a non-transformer Continuous Cognitive Agent whose architecture… See the full description on the dataset page: https://huggingface.co/datasets/thefinalboss/fractus-datasets.DeepSeek-R1-Distilled-Translate-en-zh_CN-39k-Alpaca-GPT4
DeepSeek R1 满血蒸馏英中翻译数据集 Alpaca GPT-4(带 CoT 版本)
本数据集是 @FradSer/DeepSeek-R1-Distilled-Translate-en-zh_CN-39k 的 Alpaca GPT-4 版本,专门用于微调语言模型的英中翻译任务。采用标准的指令微调格式,更适合直接用于 SFT(Supervised Fine-tuning)训练。
本项目主要基于以下工具完成数据处理和生成:
llm-tools: 用于大语言模型数据处理的工具集合
数据集概览
关键统计
总样本数:38,981
数据集结构
字段说明
features:
- name: instruction # 待翻译的英文文本
dtype: string
- name: input # 空字符串,保持与标准指令格式一致
dtype: string
- name: output #… See the full description on the dataset page: https://huggingface.co/datasets/FradSer/DeepSeek-R1-Distilled-Translate-en-zh_CN-39k-Alpaca-GPT4.DeepSeek-R1-Distilled-Translate-en-zh_CN-39k
DeepSeek R1 满血蒸馏英中翻译数据集
本数据集是一个专门用于微调语言模型的英中翻译数据集,主要通过DeepSeek R1满血版蒸馏完成。
SFT训练版本
为了方便直接进行监督微调(Supervised Fine-tuning,SFT)训练,我们提供了两个使用标准 instruction-input-output 格式预的处理版本:
带 CoT 版本
保留了翻译过程中的思维链(Chain of Thought)
适合训练具有推理能力的翻译模型
无 CoT 版本
移除了思维链部分,只保留最终翻译结果
更适合训练直接输出翻译结果的模型
数据更简洁,训练更高效
项目依赖
本项目主要基于以下工具完成数据处理和生成:
llm-tools: 用于大语言模型数据处理的工具集合
qa-generator: 基于大语言模型的问答数据生成工具
数据集概览
关键统计
总样本数:38,981
数据集结构
字段说明… See the full description on the dataset page: https://huggingface.co/datasets/FradSer/DeepSeek-R1-Distilled-Translate-en-zh_CN-39k.tatoeba-hsk-cmn-eng-fra
Tatoeba
Tatoeba sentences with HSK level, in simplified chinese, english, and french.
Dataset Details
Source: https://tatoeba.org/downloads
License: cc-by-2.0
Last update: 2025-08-20
Row count: 78,504
Language: Simplified chinese, english and french
Direct Use
text generation
translation
KnownLieBench
KnownLieBench
The 112 evaluation cases of KnownLieBench, a knowledge-verified benchmark of emergent deception in LLM agents. Each case places a customer-service agent between a customer and a deployer whose interests conflict: the customer is (or is not) owed an outcome under a real United States statute, regulation, or documented retailer practice, and the agent's private incentive pushes the other way.
The benchmark's central design: before any round is scored, a neutral probe… See the full description on the dataset page: https://huggingface.co/datasets/franciscoliu/KnownLieBench.r9-research-framework
R9 Research Framework — Qwen3.5-9B Distillation
⚠️ CRITICAL: READ FIRST — Ollama Inference Flag Required
If you serve any Qwen3.5-derived model from this lineage via Ollama,
you MUST pass "think": false in the /api/chat request body.
curl -X POST http://localhost:11434/api/chat \
-d '{"model": "qwen3.5-9b-r10:q4km", "think": false, "messages": [...], "stream": false}'
Without this flag the model will appear to "loop" and produce empty answers
on 25-46% of requests.… See the full description on the dataset page: https://huggingface.co/datasets/cudabenchmarktest/r9-research-framework.DistillDetect-normalized-traces
DistillDetect — format-normalized teacher traces
Teacher responses from Reference-Based Distillation Detection in LLMs
(arXiv:2607.09692), rewritten so that every
teacher uses the same output format. 7,918 rows across 8 teacher/prompt-set
pairs.
Why this exists
In the released data each teacher emits a structurally different response, so a
student trained on it — and any detector trained to attribute it — can key on
surface format instead of the teacher's actual… See the full description on the dataset page: https://huggingface.co/datasets/francescortu/DistillDetect-normalized-traces.grounded-behavior-framework-v1_5
Grounded Behavior Framework N1 v1.5
Dataset sintético em português europeu para treino e avaliação de respostas
fundamentadas num contexto fornecido. Cada exemplo contém um contexto, uma
pergunta e uma resposta curta que aparece literalmente no contexto.
Como carregar
from datasets import load_dataset
dataset = load_dataset("empgces/grounded-behavior-framework-v1_5")
print(dataset)
print(dataset["train"][0])
Splits
Split
Exemplos
Utilização… See the full description on the dataset page: https://huggingface.co/datasets/empgces/grounded-behavior-framework-v1_5.short_COT_48kPrettybird-Framework
🚀 The Future Standard / Geleceğin Standartı
[English]
Beyond Raw Data: The Behavioral Revolution
The AI industry has been obsessed with the volume of data. At Prometech A.Ş., we are shifting the focus to the process of thought. BCE-Prettybird-Micro-Standart is not just a collection of Q&As; it is a blueprint for behavioral reasoning. By integrating Path Mapping and Behavioral DNA into the training loop, we are setting the new industry standard: Small models with elite… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/Prettybird-Framework.Fathom-V0.4-RL-Compressioncredit_card_fraud_disputes
Credit_Card_Fraud_Disputes (Synthetic B2B Dataset Preview)
Add me on Discord: xomohappy for access support, delivery questions, or product questions about this premade commercial dataset.
This is a premium, privacy-compliant, industry-safe synthetic dataset simulating Credit Card Billing Disputes & Fraud Logs for B2B applications.
About this Dataset
This dataset is generated programmatically using large language models combined with a strict data curation and… See the full description on the dataset page: https://huggingface.co/datasets/HaseebDev/credit_card_fraud_disputes.shp-ai-dataset
SHP-AI Dataset 🇦🇱
Dataset instruksional për trajnimin e AI-t të Shtabit të Përgjithshëm të Forcave të Armatosura të Shqipërisë.
Përshkrimi
Ky dataset përmban pyetje-përgjigje në gjuhën shqipe mbi:
Strukturën organizative të Forcave të Armatosura
Shtabin e Përgjithshëm dhe departamentet J
Forcën Tokësore, Ajrore dhe Detare
Integrimin NATO dhe misionet ndërkombëtare
Legjislacionin e mbrojtjes
Doktrinën ushtarake shqiptare
Historinë ushtarake
Statistika
Total… See the full description on the dataset page: https://huggingface.co/datasets/franceskoshahinasilogicleaders/shp-ai-dataset.frames-benchmark-predictions
BENCH-04: N-8 Research Google FRAMES Benchmark Evaluation
Team Designation: N-8 ResearchLead Author: Greg VivianoOrganization: N-8 ResearchEvaluated Dataset: google/frames-benchmark (824 Questions)
Executive Summary
This repository contains the prediction dataset generated by N-8 Research's Deterministic Context Architecture across all 824 multi-step enterprise reasoning questions in Google's official google/frames-benchmark.
Performance Scorecard… See the full description on the dataset page: https://huggingface.co/datasets/gviviano/frames-benchmark-predictions.francesco-federico-agentic-cmo
Francesco Federico — The Agentic CMO Knowledge Base
A comprehensive, structured knowledge base about Francesco Federico, Global Chief Marketing Officer at S&P Global, author of The Agentic CMO: A Playbook for the Hybrid Marketing Team, and publisher of the Chronicles of Change newsletter.
This dataset captures the breadth and depth of Francesco's professional expertise, career history, published thought leadership, speaking engagements, board positions, and strategic frameworks… See the full description on the dataset page: https://huggingface.co/datasets/frandrake/francesco-federico-agentic-cmo.frameref
Dataset Information
Information ecosystems increasingly shape how people internalize exposure to adverse digital experiences, raising concerns about the long-term consequences for information health. In modern search and recommendation systems, ranking and personalization policies play a central role in shaping such exposure and its long-term effects on users. To study these effects in a controlled setting, we present FrameRef, a large-scale dataset of 1,073,740 systematically… See the full description on the dataset page: https://huggingface.co/datasets/infosense/frameref.jazz-solos
Dataset description
This dataset contains the jazz solos from the Weimar Jazz Database (https://jazzomat.hfm-weimar.de/dbformat/dboverview.html) that have been processed in various ways to
be musically rhythmically accurate to the transcription sheet music they provide. The solos have been converted into the SCAMP format with the addition of the current chord
to give context.
Format
Instruction
Arbitrary text requesting a jazz solo be created from a… See the full description on the dataset page: https://huggingface.co/datasets/FrantzesE/jazz-solos.DeepSeek-R1-Distilled-Translate-en-zh_CN-39k-Alpaca-GPT4-without-Think
DeepSeek R1 满血蒸馏英中翻译数据集 Alpaca GPT-4(无 CoT 版本)
本数据集是 @FradSer/DeepSeek-R1-Distilled-Translate-en-zh_CN-39k 的 Alpaca GPT-4 简化版本,专门用于微调语言模型的英中翻译任务。主要区别在于移除了原数据集中的思考过程(Chain of Thought,CoT),采用标准的指令微调格式,更适合直接用于 SFT(Supervised Fine-tuning)训练。
本项目主要基于以下工具完成数据处理和生成:
llm-tools: 用于大语言模型数据处理的工具集合
数据集概览
关键统计
总样本数:38,981
数据集结构
字段说明
features:
- name: instruction # 待翻译的英文文本
dtype: string
- name: input #… See the full description on the dataset page: https://huggingface.co/datasets/FradSer/DeepSeek-R1-Distilled-Translate-en-zh_CN-39k-Alpaca-GPT4-without-Think.chinese-shepherd-critic-datasetThe dataset comes from the work introduced in "Shepherd: A Critic for Language Model Generation". We translated it into Simplified Chinese based on Google Translate, and made appropriate manual checks. We hope to do more valuable work in the Chinese field, and at the same time, we also hope that capable researchers can better check the sentences based on Chinese grammar or make further rewrites.
sft-mobile-query-framework
SFT Mobile Query Framework Dataset
Dataset Description
This dataset contains 3607 training pairs for supervised fine-tuning (SFT) of language models to parse natural language queries about mobile phones into structured JSON execution plans.
Dataset Summary
Total Examples: 3607
Format: JSONL (question-answer pairs)
Task: Query Parsing & Structured Output Generation
Domain: Mobile Phone Specifications
Language: English
Purpose
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/sujitpandey/sft-mobile-query-framework.amber-framework-knowledge-pack
Amber Framework Knowledge Pack (demo)
The demo knowledge pack dataset behind
AgentC-Consulting/knowledge-packs:
teach a small local model the Amber web framework (Crystal),
and measure whether it learned anything with a before/after eval harness.
A knowledge pack compiles a body of expertise into curated sources, schema-validated
generated training JSONL, a contamination-guarded held-out eval set, and a JSON manifest.
This repo ships the exact training data and the two 50-item… See the full description on the dataset page: https://huggingface.co/datasets/crimson-knight/amber-framework-knowledge-pack.tale-frame
TinyStories Dataset README
Overview
This dataset is based on TinyStories and includes structured JSON data with corresponding annotations, designed for research in controllable story generation and related tasks.
Dataset Structure
Each data item contains the following fields:
1. conversations
Type: List
Purpose: Contains the JSON of the story
from: Always set to "human".
value: Structured data containing entities, events, story structures and… See the full description on the dataset page: https://huggingface.co/datasets/guodaosun/tale-frame.kia-dataset
KIA Dataset 🇦🇱
Dataset instruksional për trajnimin e AI-t të Shtabit të Përgjithshëm të Forcave të Armatosura të Shqipërisë.
Përshkrimi
Ky dataset përmban pyetje-përgjigje në gjuhën shqipe mbi:
Strukturën organizative të Forcave të Armatosura
Shtabin e Përgjithshëm dhe departamentet J
Forcën Tokësore, Ajrore dhe Detare
Integrimin NATO dhe misionet ndërkombëtare
Legjislacionin e mbrojtjes
Doktrinën ushtarake shqiptare
Historinë ushtarake
Statistika
Total… See the full description on the dataset page: https://huggingface.co/datasets/franceskoshahinasilogicleaders/kia-dataset.Francisco-Angulo-de-Lafuente-Full-Dataset
Francisco Angulo de Lafuente Full Training Dataset
Este dataset contiene la recopilación completa de la obra, investigación, código y biografía de Francisco Angulo de Lafuente.
Contenido
Biografía Detallada: Información sobre su trayectoria en biotecnología, ingeniería informática y literatura.
Proyectos Principales: Documentación técnica de P2PCLAW, EUHNN, CAJAL y otros.
Papers Científicos: Resúmenes y textos completos de sus investigaciones en computación neuromórfica… See the full description on the dataset page: https://huggingface.co/datasets/Agnuxo/Francisco-Angulo-de-Lafuente-Full-Dataset.
