datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cell2sentence4longevity-data
Dataset Card: longevity-genie/cell2sentence4longevity-data
Summary
This repository contains preprocessed single-cell RNA-seq (scRNA‑seq) datasets prepared as “cell sentences” for training and evaluation of cells2sentence-style models. Each cell is represented as a space‑separated sequence of top expressed gene symbols, enabling language‑model style training for tasks such as biological age prediction and other downstream applications.
This dataset targets fine‑tuning and… See the full description on the dataset page: https://huggingface.co/datasets/longevity-genie/cell2sentence4longevity-data.2026.RA.Frontier-and-Scale-Cells
Rational-Agent Frontier, Scale, and Framing Cells
This public dataset is a sibling of siddharthmb/2026.RA.Negotiation-Campaigns (the frozen P1-P4 experimental record for the ii_mats/experiments/rational_agents negotiation program) and follows the same conventions: raw per-episode JSON, per-turn oracle annotations, Markdown/HTML transcripts, run manifests, analysis tables, and an integrity manifest over every uploaded file. It packages eight later campaigns that were run against… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Frontier-and-Scale-Cells.omnimcp_python_celery_workers_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_python_celery_workers_teaser.project-gutenberg-clean
project-gutenberg-clean
Dataset de livros do Project Gutenberg baixados via API Gutendex e limpos com foco em qualidade para treinamento de LLMs.
Diferenciais desta versão:
Limpeza Profunda: Além de cabeçalhos e rodapés padrão, removemos notas de transcrição, prefácios de editores e artefatos de OCR (como tags [Illustration], [Music], etc.) em múltiplos idiomas (PT, EN, DE, ES, FR, IT).
Tratamento de Notas: Identifica e remove blocos de notas delimitados por sequências de… See the full description on the dataset page: https://huggingface.co/datasets/celsowm/project-gutenberg-clean.perovskite-solar-cell-efficiency-autoresearch
🔬 Perovskite Solar Cell Text Corpus for Karpathy's autoresearch
A 98.9 MB text corpus of perovskite solar cell scientific literature formatted for direct use with karpathy/autoresearch — the autonomous LLM-driven hyperparameter search framework that trains a GPT from scratch and has an AI agent iteratively modify train.py to minimize val_bpb (bits per byte).
📊 Dataset Stats
Metric
Value
Total documents
19,730
Total text
98.9 MB (~103M characters)… See the full description on the dataset page: https://huggingface.co/datasets/CollinL/perovskite-solar-cell-efficiency-autoresearch.task1506_celebrity_minimal_dob_span
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1506_celebrity_minimal_dob_span
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1506_celebrity_minimal_dob_span.task1486_cell_extraction_anem_dataset
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1486_cell_extraction_anem_dataset
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1486_cell_extraction_anem_dataset.Celestia3-DeepSeek-R1-0528Click here to support our open-source dataset and model releases!
Celestia3-DeepSeek-R1-0528 is a dataset focused on science, testing the limits of DeepSeek R1 0528's science-reasoning skills!
This dataset contains:
90.9k synthetically generated science prompts, with all responses generated using DeepSeek R1 0528.
Primary subjects are physics, chemistry, biology, and computer science; secondary subjects include Earth science, astronomy, and information theory.
All prompts are synthetic, taken… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Celestia3-DeepSeek-R1-0528.cellarc_100k_meta
cellarc_100k_meta
CellARC 100k Meta is the metadata‑rich variant of the CellARC benchmark introduced in Lzicar, M. (2025). CellARC: Measuring Intelligence with Cellular Automata. It contains the exact same episodes and splits as cellarc_100k, with byte‑identical Parquet files; the JSONL files retain full per‑episode metadata (rule tables, coverage diagnostics, morphology descriptors, sampling parameters, etc.). Each episode exposes five support pairs plus a held‑out query/solution… See the full description on the dataset page: https://huggingface.co/datasets/mireklzicar/cellarc_100k_meta.cellarc_100k
cellarc_100k
CellARC 100k a cellular-automata benchmark dataset introduced in Lzicar, M. (2025). CellARC: Measuring Intelligence with Cellular Automata. Each episode exposes five support pairs plus a held-out query/solution pair.
Data quick facts
Alphabet size k in [2, 6]; window size W in {3, 5, 7}; radius r in {1, 2, 3}; steps t in {1, 2, 3} (≈95% have t = 1).
Values (digits) are integers in 0..k-1 per episode; across the full dataset the union of symbols is {0,1,2,3,4… See the full description on the dataset page: https://huggingface.co/datasets/mireklzicar/cellarc_100k.celestial-comprehensive-spiritual-ai
🌟 CELESTIAL Comprehensive Spiritual AI Dataset
🚀 SPEED-OPTIMIZED TRAINING - 45-90 MINUTES!
Latest Update: Added speed-optimized training notebook that reduces training time from 21+ hours to 45-90 minutes (15-20x faster!)
📊 Dataset Overview
Comprehensive spiritual AI training dataset with 3000+ conversations covering all 50+ CELESTIAL spiritual systems including the newly integrated Sanjay Jumaani numerology method.
🎯 Key Features:
⚡… See the full description on the dataset page: https://huggingface.co/datasets/dp1812/celestial-comprehensive-spiritual-ai.text-to-mermaidlegal_br_sft
Legal BR SFT Dataset ⚖️🇧🇷 (Auditado)
O Legal BR SFT é um dataset de instruções de alta qualidade focado exclusivamente no Direito Brasileiro. Ele foi projetado para o treinamento de modelos de linguagem (LLMs) através de Supervised Fine-Tuning (SFT).
📊 Estatísticas Auditadas (Regex Refinado)
Após auditoria estatística estratificada em 38.153 registros, a distribuição por área do Direito é:
Área do Direito
Porcentagem
Temas Principais
Direito Civil
19… See the full description on the dataset page: https://huggingface.co/datasets/celsowm/legal_br_sft.jurisprudencias_stfceleritybench-tool-choice
Which small model should run a Mac launcher
Celeritas is a Spotlight-style launcher that turns what somebody types into tool
calls on their own machine. Picking the model to put behind it meant measuring
them, and the numbers were going on a public page, so the runs behind them are
here.
The question is narrow on purpose: for an agent with about thirty tools on a
desktop, which model picks the right one? Not reasoning, not code, not
knowledge. Tool choice, on short everyday… See the full description on the dataset page: https://huggingface.co/datasets/celerity-labs/celeritybench-tool-choice.celestial-comprehensive-dataset-v2
CELESTIAL Comprehensive Spiritual AI Dataset v2.0
🌟 Overview
The most comprehensive dataset for training spiritual AI assistants, featuring 9,000+ high-quality examples across all major spiritual and astrological domains.
📊 Dataset Statistics
Total Examples: 9,000
Training Split: 7,200 examples
Validation Split: 900 examples
Test Split: 900 examples
Categories: 4 categories
Languages: English, Hindi (transliterated)
🎯 Categories Included… See the full description on the dataset page: https://huggingface.co/datasets/Amvhunt/celestial-comprehensive-dataset-v2.Celestia3-DeepSeek-R1-0528-PREVIEWClick here to support our open-source dataset and model releases!
This is an early sneak preview of Celestia3-DeepSeek-R1-0528, containing the first 13.4k rows!
Celestia3-DeepSeek-R1-0528 is a dataset focused on science, testing the limits of DeepSeek R1's science-reasoning skills!
This early preview release contains:
13.4k synthetically generated science prompts. All responses are generated using DeepSeek R1 0528.
Primary subjects are physics, chemistry, biology, and computer science;… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Celestia3-DeepSeek-R1-0528-PREVIEW.v11-cells-midtrain-corpus
v11 cells mid-training corpus
The delegating arm of a paired experiment: teach a 115M model to call an external
tool for arithmetic rather than to memorise the answers. Its partner, the maths-only
arm, teaches the same model to absorb the arithmetic into its weights instead.
Pre-tokenized against the v11 tokenizer
(10dd5110…, vocab 71,260), for
chrishayuk/v11-tinystories-115m-base.
Identity: 2115d6aeff3428e217ef2903a8030facd511dcb00183e9fc3faaf49d01038767
(chuk-datasets… See the full description on the dataset page: https://huggingface.co/datasets/chrishayuk/v11-cells-midtrain-corpus.DAPO-17K-Plus
RLVR Datasets and Where to Find Them: Tracing Data Lineage for Better Training Data
Introduction
DAPO-17k-Plus (DAPO++) is the dataset presented in the paper RLVR Datasets and Where to Find Them: Tracing Data Lineage for Better Training Data.
Acknowledgements
DAPO++ is built on the following repositories and we thank their teams for their valuable contributions to the community:
DAPO
Citation
If you find our work useful, feel… See the full description on the dataset page: https://huggingface.co/datasets/CelineHuangxy/DAPO-17K-Plus.gemini_orpo_dpo_ptbrreddit_dataset_190
Bittensor Subnet 13 Reddit Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks.
For more… See the full description on the dataset page: https://huggingface.co/datasets/CelestialWandererOfTheVoid/reddit_dataset_190.leis_estaduais_rjmerged-corpus
Merged Corpus
Welcome to this repository.This dataset is japanese corpus that includes wiki, wikibooks, wikiversity, cc100, and oscar2109.
Getting Started
If you want to use this, please run as follows.This process takes about 3 hours.
mkdir -p pretrain/input/
cd pretrain/input/
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/datasets/ce-lery/merged-corpus.git
cd merged-corpus
git lfs pull
bash merge_train.sh
leis_ordinarias_1988_2024valdoria-dpo-qwen35-dataset
Valdoria DPO dataset
Dataset de preferências conversacional para uso direto com trl.DPOTrainer.
Splits
train.jsonl: 1.889 pares
validation.jsonl: 236 pares
test.jsonl: 237 pares
Cada linha contém:
{
"id": "...",
"prompt": [{"role": "system", "content": "..."}, {"role": "user", "content": "..."}],
"chosen": [{"role": "assistant", "content": "..."}],
"rejected": [{"role": "assistant", "content": "..."}],
"metadata": {"task_type": "..."… See the full description on the dataset page: https://huggingface.co/datasets/celsowm/valdoria-dpo-qwen35-dataset.minecraft_qa_es
Minecraft Q&A (Spanish)
A Spanish, chat-formatted question/answer dataset about Minecraft. Each example is a short conversation with a single user question and a single assistant answer (plus a system prompt).
Data format
The dataset is provided as JSON Lines (.jsonl): one JSON object per line.
Each record has a single key:
messages: an array of chat messages, each with:
role: one of "system", "user", "assistant"
content: the message text
Typical structure:… See the full description on the dataset page: https://huggingface.co/datasets/CelesteLove/minecraft_qa_es.modelos_peticoesHTML was converted to Markdown (better for LLMs)
reddit_dataset_231
Bittensor Subnet 13 Reddit Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks.
For more… See the full description on the dataset page: https://huggingface.co/datasets/CelestialWandererOfTheVoid/reddit_dataset_231.celestial-comprehensive-spiritual-ai
🌟 CELESTIAL Comprehensive Spiritual AI Dataset
🚀 SPEED-OPTIMIZED TRAINING - 45-90 MINUTES!
Latest Update: Added speed-optimized training notebook that reduces training time from 21+ hours to 45-90 minutes (15-20x faster!)
📊 Dataset Overview
Comprehensive spiritual AI training dataset with 3000+ conversations covering all 50+ CELESTIAL spiritual systems including the newly integrated Sanjay Jumaani numerology method.
🎯 Key Features:
⚡… See the full description on the dataset page: https://huggingface.co/datasets/Emilynnjk/celestial-comprehensive-spiritual-ai.x_dataset_231
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/CelestialWandererOfTheVoid/x_dataset_231.
