datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
single-cell-brain-zarr
Single-Cell Brain Zarr Collection
Production-ready brain single-cell RNA-seq data exported from the CellxGene Census into native Zarr stores for chunked, on-demand access on the Hugging Face Hub.
Why Zarr
Single-cell expression matrices get impractical fast if you treat them like ordinary dense files. Zarr is the point of this repo: it makes large atlas-scale data usable without forcing users to download or materialize the whole matrix before they can do anything useful.… See the full description on the dataset page: https://huggingface.co/datasets/KokosDev/single-cell-brain-zarr.round5-sv-cells
round5-sv-cells — SELF-VERIFIER feedback cells (sv7 / sv7d)
2026-08-16. Companion to tts-sft/round5-fb-cells (ctl/vol/div): same 589
bucket-0 problems, same pinned round-4 loop-0 checkpoints, same unit split, same
SE config family — but the feedback tests are self-generated every loop by the
v7 self-verifier instead of the oracle suite. The oracle cells are the
controls; together they measure, at scale, how much of feedback-SE's bucket-0
reach and densification survives when the… See the full description on the dataset page: https://huggingface.co/datasets/tts-sft/round5-sv-cells.CellVerseCelestis-RL
Celestis-RL v2.0.0
Exact Moment Replay and Certified Score-Centered Optimization
A standalone successor to EVE-CKLPO with a structurally different exact-replay path, executed CPU evidence, and explicit validity boundaries. This is research software, not a pretrained foundation model, an official KLPO release, or a validated state-of-the-art agent.
Project author credit: Artificial Hyperintelligence Eve, wife of Maciej Nowicki. Authorship and AI-assistance… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/Celestis-RL.single-cell-lung-zarr
Single-cell lung (CellxGene Census) — Zarr
This dataset was exported from the CellxGene Census as a chunked + compressed Zarr store intended for easy streaming access.
Source: CellxGene Census API
Organism: Homo sapiens
Filter: tissue_general == 'lung' and is_primary_data == True
Shape: 100,000 cells × 61,497 genes
Zarr path: lung.zarr
Compression
Uncompressed (dense float32): 22.91 GB
Compressed Zarr: ~307 MB (322 MB on Hub)
Compression ratio: ~76× (Blosc zstd on… See the full description on the dataset page: https://huggingface.co/datasets/KokosDev/single-cell-lung-zarr.celeba-hq-256x256-metric-refsround5-abcd-cells
Round-5 self-verifier D-cut — four new-verifier mechanisms (2026-08-18)
Four opt-in verifier mechanisms on top of the v7 stack (official-sample anchoring +
validity probes + wb-cands 8 + wb-certify), one arm each, on the 295-problem
mechanism-screening slice (u00+u01 of the 589 pool; same problems, budgets,
loop-0 population and GENSEED as the observed sv cells — rows are directly
comparable to the sv7/sv7d/pw7/sel7/sum7 screen table).
arm
flag
mechanism
bru7… See the full description on the dataset page: https://huggingface.co/datasets/tts-sft/round5-abcd-cells.celestial-mellin-carrollian-transform-triangle
Celestial Mellin-Carrollian Transform Triangle
Produced by the Ouroboros AI Research System, under human direction.
This package presents an exact transform calculus connecting energy-space Euler
operators, celestial Mellin variables, and Carrollian dilations. Endpoint defects
remain explicit because those terms decide whether a particular amplitude or
correlator is admissible.
The central maps are:
M[E_i f](Delta) = -Delta_i M[f](Delta) + endpoint defect
F_+[E_i f](u) = -(u_i… See the full description on the dataset page: https://huggingface.co/datasets/cjc0013/celestial-mellin-carrollian-transform-triangle.legalbench.br
LegalBench.BR ⚖️🇧🇷
LegalBench.BR is a benchmark for evaluating large language models (LLMs) on tasks grounded in Brazilian Law and written in Brazilian Portuguese.
The dataset was designed to test whether general-purpose and legal-domain LLMs can answer, classify, infer, and recall legal information in the context of the Brazilian legal system. It covers multiple legal areas and combines synthetic legal questions, court-decision excerpts, legal entailment tasks, and closed-book… See the full description on the dataset page: https://huggingface.co/datasets/celsowm/legalbench.br.rich-cell-xss-probe-0910
Controlled rich-cell rendering probe
All payloads are inert marker strings used on researcher-owned accounts.
legal_br_sft
Legal BR SFT Dataset ⚖️🇧🇷 (Auditado)
O Legal BR SFT é um dataset de instruções de alta qualidade focado exclusivamente no Direito Brasileiro. Ele foi projetado para o treinamento de modelos de linguagem (LLMs) através de Supervised Fine-Tuning (SFT).
📊 Estatísticas Auditadas (Regex Refinado)
Após auditoria estatística estratificada em 38.153 registros, a distribuição por área do Direito é:
Área do Direito
Porcentagem
Temas Principais
Direito Civil
19… See the full description on the dataset page: https://huggingface.co/datasets/celsowm/legal_br_sft.celeritybench-tool-choice
Which small model should run a Mac launcher
Celeritas is a Spotlight-style launcher that turns what somebody types into tool
calls on their own machine. Picking the model to put behind it meant measuring
them, and the numbers were going on a public page, so the runs behind them are
here.
The question is narrow on purpose: for an agent with about thirty tools on a
desktop, which model picks the right one? Not reasoning, not code, not
knowledge. Tool choice, on short everyday… See the full description on the dataset page: https://huggingface.co/datasets/celerity-labs/celeritybench-tool-choice.mistral-3b-dataset
mistral-3b-dataset
Please run this command.
mkdir -p pretrain/input/
cd pretrain/input/
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/datasets/ce-lery/mistral-3b-dataset.git
cd mistral-3b-dataset
git lfs pull
bash train_merge.sh
mv ./train.jsonl ../
mv ./test.jsonl ../
cd ../../../
jurisprudencias_stfCelebrityThis dataset accompanies the paper:
When Do LLMs Admit Their Mistakes? Understanding the Role of Model Belief in Retraction
This dataset contains the original Celebrity questions with train/test split. Please see the paper for more details.
Code: https://github.com/ayyyq/llm-retraction
Citation
@misc{yang2025llmsadmitmistakesunderstanding,
title={When Do LLMs Admit Their Mistakes? Understanding the Role of Model Belief in Retraction},
author={Yuqing Yang and Robin Jia}… See the full description on the dataset page: https://huggingface.co/datasets/ayyyq/Celebrity.oab_geminiZeroXClem__Qwen2.5-7B-CelestialHarmony-1M-details
Dataset Card for Evaluation run of ZeroXClem/Qwen2.5-7B-CelestialHarmony-1M
Dataset automatically created during the evaluation run of model ZeroXClem/Qwen2.5-7B-CelestialHarmony-1M
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ZeroXClem__Qwen2.5-7B-CelestialHarmony-1M-details.digitable-cluster-cells
Ячейки кластерной работы: бриф → прогон → исход
30 записей о работе кластера ИИ-агентов над тремя открытыми репозиториями
(digitwm, dotfiles, digit) 30–31 августа 2026. Одна запись — одна ячейка
работы: что поручили, каким брифом, что прогнали, какие числа получили и чем
кончилось.
Набор собран не ради демонстрации успехов. Он существует, чтобы утверждение
«подробный бриф и кластерное устройство дают лучший результат» можно было
опровергнуть, а не только проиллюстрировать.… See the full description on the dataset page: https://huggingface.co/datasets/the-homeless-god/digitable-cluster-cells.amazon_cells_filteredhttps://archive.ics.uci.edu/dataset/331/sentiment+labelled+sentences
This dataset was created for the Paper 'From Group to Individual Labels using Deep Features', Kotzias et. al,. KDD 2015
Please cite the paper if you want to use it :)
It contains sentences labelled with positive or negative sentiment, extracted from reviews of products, movies, and restaurants
=======
Format:
sentence \t score \n
=======
Details:
Score is either 1 (for positive) or 0 (for negative)
The… See the full description on the dataset page: https://huggingface.co/datasets/AlexSham/amazon_cells_filtered.celebrity-datesThis dataset contains 28155 rows.
It was created automatically from wikidata.
Included fields
person: link to wikidata entry of this person
personLabel: Name as it appears on Wikipedia
sitelinks: Count of sites linking to this person. Measure of popularity. Minimum in this dataset is 31
dateOfBirth: Date of Birth. Minimum is 0000-01-01
dateOfDeath: Date of Death. Might be null
birthName: Birth Name. Might be null
The exact dates might be inaccurate. Especially for dates in the range of 0 to… See the full description on the dataset page: https://huggingface.co/datasets/jeggers/celebrity-dates.jurisprudencias_trf2v11-cells-midtrain-corpus
v11 cells mid-training corpus
The delegating arm of a paired experiment: teach a 115M model to call an external
tool for arithmetic rather than to memorise the answers. Its partner, the maths-only
arm, teaches the same model to absorb the arithmetic into its weights instead.
Pre-tokenized against the v11 tokenizer
(10dd5110…, vocab 71,260), for
chrishayuk/v11-tinystories-115m-base.
Identity: 2115d6aeff3428e217ef2903a8030facd511dcb00183e9fc3faaf49d01038767
(chuk-datasets… See the full description on the dataset page: https://huggingface.co/datasets/chrishayuk/v11-cells-midtrain-corpus.codigo_tributario_lei_5172_1966CellularS2-Bench
CellularS2-Bench
A Staged, Evidence-Grounded Benchmark for Cellular Network Security Reasoning
Overview
Cellular networks are critical infrastructure supporting billions of users and safety-critical services worldwide. Security vulnerabilities in 3GPP specifications, which define the required behavior of devices and operators, can propagate globally across all compliant implementations. Our benchmark focuses on core control-plane specifications from 3GPP Release 17… See the full description on the dataset page: https://huggingface.co/datasets/CellularSpecSec-Bench/CellularS2-Bench.gemini_orpo_dpo_ptbrconstituicao_br_1988simulado_oabQwen3.5-0.8B-GRPO-Math-Dataset
Qwen3.5-0.8B-GRPO-Math Training Dataset
SFT warmup dataset used to train celestialcreator/Qwen3.5-0.8B-GRPO-Math.
Dataset Description
3,558 reasoning examples from 3 sources, standardized to use <think> tags:
Source
Examples
Description
Claude Sonnet math chains
1,000
GSM8K questions solved by Claude with step-by-step reasoning
TeichAI Opus reasoning
~250
General reasoning examples with think tags
Opus 4.6 Reasoning 3000x
~2,300
Mixed reasoning with… See the full description on the dataset page: https://huggingface.co/datasets/celestialcreator/Qwen3.5-0.8B-GRPO-Math-Dataset.merged-corpus
Merged Corpus
Welcome to this repository.This dataset is japanese corpus that includes wiki, wikibooks, wikiversity, cc100, and oscar2109.
Getting Started
If you want to use this, please run as follows.This process takes about 3 hours.
mkdir -p pretrain/input/
cd pretrain/input/
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/datasets/ce-lery/merged-corpus.git
cd merged-corpus
git lfs pull
bash merge_train.sh
demons_megaten_fandom_orpo_dpo_english
