CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01longevity-genie /cell2sentence4longevity-data Dataset Card: longevity-genie/cell2sentence4longevity-data Summary This repository contains preprocessed single-cell RNA-seq (scRNA‑seq) datasets prepared as “cell sentences” for training and evaluation of cells2sentence-style models. Each cell is represented as a space‑separated sequence of top expressed gene symbols, enabling language‑model style training for tasks such as biological age prediction and other downstream applications. This dataset targets fine‑tuning and… See the full description on the dataset page: https://huggingface.co/datasets/longevity-genie/cell2sentence4longevity-data.tabulartext-generation10M<n<100M0 likes815 downloads11mo agoHugging Face02siddharthmb /2026.RA.Frontier-and-Scale-Cells Rational-Agent Frontier, Scale, and Framing Cells This public dataset is a sibling of siddharthmb/2026.RA.Negotiation-Campaigns (the frozen P1-P4 experimental record for the ii_mats/experiments/rational_agents negotiation program) and follows the same conventions: raw per-episode JSON, per-turn oracle annotations, Markdown/HTML transcripts, run manifests, analysis tables, and an integrity manifest over every uploaded file. It packages eight later campaigns that were run against… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Frontier-and-Scale-Cells.tabulartext-generation10K<n<100K0 likes174 downloads2mo agoHugging Face03emgena /omnimcp_python_celery_workers_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_python_celery_workers_teaser.text-generationn<1K0 likes171 downloads9d agoHugging Face04celsowm /project-gutenberg-clean project-gutenberg-clean Dataset de livros do Project Gutenberg baixados via API Gutendex e limpos com foco em qualidade para treinamento de LLMs. Diferenciais desta versão: Limpeza Profunda: Além de cabeçalhos e rodapés padrão, removemos notas de transcrição, prefácios de editores e artefatos de OCR (como tags [Illustration], [Music], etc.) em múltiplos idiomas (PT, EN, DE, ES, FR, IT). Tratamento de Notas: Identifica e remove blocos de notas delimitados por sequências de… See the full description on the dataset page: https://huggingface.co/datasets/celsowm/project-gutenberg-clean.texttext-generation1K<n<10K0 likes165 downloads4mo agoHugging Face05CollinL /perovskite-solar-cell-efficiency-autoresearch 🔬 Perovskite Solar Cell Text Corpus for Karpathy's autoresearch A 98.9 MB text corpus of perovskite solar cell scientific literature formatted for direct use with karpathy/autoresearch — the autonomous LLM-driven hyperparameter search framework that trains a GPT from scratch and has an AI agent iteratively modify train.py to minimize val_bpb (bits per byte). 📊 Dataset Stats Metric Value Total documents 19,730 Total text 98.9 MB (~103M characters)… See the full description on the dataset page: https://huggingface.co/datasets/CollinL/perovskite-solar-cell-efficiency-autoresearch.texttext-generation10K<n<100K0 likes158 downloads5mo agoHugging Face06Lots-of-LoRAs /task1506_celebrity_minimal_dob_span Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1506_celebrity_minimal_dob_span Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1506_celebrity_minimal_dob_span.texttext-generationn<1K0 likes148 downloads2y agoHugging Face07Lots-of-LoRAs /task1486_cell_extraction_anem_dataset Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1486_cell_extraction_anem_dataset Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1486_cell_extraction_anem_dataset.texttext-generationn<1K0 likes122 downloads2y agoHugging Face08sequelbox /Celestia3-DeepSeek-R1-0528Click here to support our open-source dataset and model releases! Celestia3-DeepSeek-R1-0528 is a dataset focused on science, testing the limits of DeepSeek R1 0528's science-reasoning skills! This dataset contains: 90.9k synthetically generated science prompts, with all responses generated using DeepSeek R1 0528. Primary subjects are physics, chemistry, biology, and computer science; secondary subjects include Earth science, astronomy, and information theory. All prompts are synthetic, taken… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Celestia3-DeepSeek-R1-0528.texttext-generation10K<n<100K35 likes121 downloads1y agoHugging Face09mireklzicar /cellarc_100k_meta cellarc_100k_meta CellARC 100k Meta is the metadata‑rich variant of the CellARC benchmark introduced in Lzicar, M. (2025). CellARC: Measuring Intelligence with Cellular Automata. It contains the exact same episodes and splits as cellarc_100k, with byte‑identical Parquet files; the JSONL files retain full per‑episode metadata (rule tables, coverage diagnostics, morphology descriptors, sampling parameters, etc.). Each episode exposes five support pairs plus a held‑out query/solution… See the full description on the dataset page: https://huggingface.co/datasets/mireklzicar/cellarc_100k_meta.textother10K<n<100K1 likes114 downloads11mo agoHugging Face10mireklzicar /cellarc_100k cellarc_100k CellARC 100k a cellular-automata benchmark dataset introduced in Lzicar, M. (2025). CellARC: Measuring Intelligence with Cellular Automata. Each episode exposes five support pairs plus a held-out query/solution pair. Data quick facts Alphabet size k in [2, 6]; window size W in {3, 5, 7}; radius r in {1, 2, 3}; steps t in {1, 2, 3} (≈95% have t = 1). Values (digits) are integers in 0..k-1 per episode; across the full dataset the union of symbols is {0,1,2,3,4… See the full description on the dataset page: https://huggingface.co/datasets/mireklzicar/cellarc_100k.textother10K<n<100K1 likes93 downloads11mo agoHugging Face11dp1812 /celestial-comprehensive-spiritual-ai 🌟 CELESTIAL Comprehensive Spiritual AI Dataset 🚀 SPEED-OPTIMIZED TRAINING - 45-90 MINUTES! Latest Update: Added speed-optimized training notebook that reduces training time from 21+ hours to 45-90 minutes (15-20x faster!) 📊 Dataset Overview Comprehensive spiritual AI training dataset with 3000+ conversations covering all 50+ CELESTIAL spiritual systems including the newly integrated Sanjay Jumaani numerology method. 🎯 Key Features: ⚡… See the full description on the dataset page: https://huggingface.co/datasets/dp1812/celestial-comprehensive-spiritual-ai.text-generation1K<n<10K2 likes76 downloads1y agoHugging Face12Celiadraw /text-to-mermaidtexttext-generation1K<n<10K10 likes64 downloads2y agoHugging Face13celsowm /legal_br_sft Legal BR SFT Dataset ⚖️🇧🇷 (Auditado) O Legal BR SFT é um dataset de instruções de alta qualidade focado exclusivamente no Direito Brasileiro. Ele foi projetado para o treinamento de modelos de linguagem (LLMs) através de Supervised Fine-Tuning (SFT). 📊 Estatísticas Auditadas (Regex Refinado) Após auditoria estatística estratificada em 38.153 registros, a distribuição por área do Direito é: Área do Direito Porcentagem Temas Principais Direito Civil 19… See the full description on the dataset page: https://huggingface.co/datasets/celsowm/legal_br_sft.texttext-generation100K<n<1M0 likes61 downloads3mo agoHugging Face14celsowm /jurisprudencias_stftexttext-generation1K<n<10K1 likes60 downloads2y agoHugging Face15celerity-labs /celeritybench-tool-choice Which small model should run a Mac launcher Celeritas is a Spotlight-style launcher that turns what somebody types into tool calls on their own machine. Picking the model to put behind it meant measuring them, and the numbers were going on a public page, so the runs behind them are here. The question is narrow on purpose: for an agent with about thirty tools on a desktop, which model picks the right one? Not reasoning, not code, not knowledge. Tool choice, on short everyday… See the full description on the dataset page: https://huggingface.co/datasets/celerity-labs/celeritybench-tool-choice.tabulartext-generation1K<n<10K0 likes53 downloads7d agoHugging Face16Amvhunt /celestial-comprehensive-dataset-v2 CELESTIAL Comprehensive Spiritual AI Dataset v2.0 🌟 Overview The most comprehensive dataset for training spiritual AI assistants, featuring 9,000+ high-quality examples across all major spiritual and astrological domains. 📊 Dataset Statistics Total Examples: 9,000 Training Split: 7,200 examples Validation Split: 900 examples Test Split: 900 examples Categories: 4 categories Languages: English, Hindi (transliterated) 🎯 Categories Included… See the full description on the dataset page: https://huggingface.co/datasets/Amvhunt/celestial-comprehensive-dataset-v2.texttext-generation10K<n<100K0 likes50 downloads7mo agoHugging Face17sequelbox /Celestia3-DeepSeek-R1-0528-PREVIEWClick here to support our open-source dataset and model releases! This is an early sneak preview of Celestia3-DeepSeek-R1-0528, containing the first 13.4k rows! Celestia3-DeepSeek-R1-0528 is a dataset focused on science, testing the limits of DeepSeek R1's science-reasoning skills! This early preview release contains: 13.4k synthetically generated science prompts. All responses are generated using DeepSeek R1 0528. Primary subjects are physics, chemistry, biology, and computer science;… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Celestia3-DeepSeek-R1-0528-PREVIEW.texttext-generation10K<n<100K7 likes47 downloads1y agoHugging Face18chrishayuk /v11-cells-midtrain-corpus v11 cells mid-training corpus The delegating arm of a paired experiment: teach a 115M model to call an external tool for arithmetic rather than to memorise the answers. Its partner, the maths-only arm, teaches the same model to absorb the arithmetic into its weights instead. Pre-tokenized against the v11 tokenizer (10dd5110…, vocab 71,260), for chrishayuk/v11-tinystories-115m-base. Identity: 2115d6aeff3428e217ef2903a8030facd511dcb00183e9fc3faaf49d01038767 (chuk-datasets… See the full description on the dataset page: https://huggingface.co/datasets/chrishayuk/v11-cells-midtrain-corpus.texttext-generationn<1K0 likes38 downloads2mo agoHugging Face19CelineHuangxy /DAPO-17K-Plus RLVR Datasets and Where to Find Them: Tracing Data Lineage for Better Training Data Introduction DAPO-17k-Plus (DAPO++) is the dataset presented in the paper RLVR Datasets and Where to Find Them: Tracing Data Lineage for Better Training Data. Acknowledgements DAPO++ is built on the following repositories and we thank their teams for their valuable contributions to the community: DAPO Citation If you find our work useful, feel… See the full description on the dataset page: https://huggingface.co/datasets/CelineHuangxy/DAPO-17K-Plus.texttext-generation10K<n<100K0 likes37 downloads2mo agoHugging Face20celsowm /gemini_orpo_dpo_ptbrtexttext-generation10K<n<100K2 likes34 downloads2y agoHugging Face21CelestialWandererOfTheVoid /reddit_dataset_190 Bittensor Subnet 13 Reddit Dataset Miner Data Compliance Agreement In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy. Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks. For more… See the full description on the dataset page: https://huggingface.co/datasets/CelestialWandererOfTheVoid/reddit_dataset_190.texttext-classification10K<n<100K0 likes31 downloads2y agoHugging Face22celsowm /leis_estaduais_rjtexttext-generation1K<n<10K0 likes31 downloads1y agoHugging Face23ce-lery /merged-corpus Merged Corpus Welcome to this repository.This dataset is japanese corpus that includes wiki, wikibooks, wikiversity, cc100, and oscar2109. Getting Started If you want to use this, please run as follows.This process takes about 3 hours. mkdir -p pretrain/input/ cd pretrain/input/ GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/datasets/ce-lery/merged-corpus.git cd merged-corpus git lfs pull bash merge_train.sh texttext-generation10M<n<100M0 likes31 downloads1y agoHugging Face24celsowm /leis_ordinarias_1988_2024textsummarization1K<n<10K2 likes26 downloads2y agoHugging Face25celsowm /valdoria-dpo-qwen35-dataset Valdoria DPO dataset Dataset de preferências conversacional para uso direto com trl.DPOTrainer. Splits train.jsonl: 1.889 pares validation.jsonl: 236 pares test.jsonl: 237 pares Cada linha contém: { "id": "...", "prompt": [{"role": "system", "content": "..."}, {"role": "user", "content": "..."}], "chosen": [{"role": "assistant", "content": "..."}], "rejected": [{"role": "assistant", "content": "..."}], "metadata": {"task_type": "..."… See the full description on the dataset page: https://huggingface.co/datasets/celsowm/valdoria-dpo-qwen35-dataset.texttext-generation1K<n<10K0 likes25 downloads3mo agoHugging Face26CelesteLove /minecraft_qa_es Minecraft Q&A (Spanish) A Spanish, chat-formatted question/answer dataset about Minecraft. Each example is a short conversation with a single user question and a single assistant answer (plus a system prompt). Data format The dataset is provided as JSON Lines (.jsonl): one JSON object per line. Each record has a single key: messages: an array of chat messages, each with: role: one of "system", "user", "assistant" content: the message text Typical structure:… See the full description on the dataset page: https://huggingface.co/datasets/CelesteLove/minecraft_qa_es.textquestion-answering10K<n<100K0 likes24 downloads9mo agoHugging Face27celsowm /modelos_peticoesHTML was converted to Markdown (better for LLMs) texttext-generationn<1K1 likes23 downloads2y agoHugging Face28CelestialWandererOfTheVoid /reddit_dataset_231 Bittensor Subnet 13 Reddit Dataset Miner Data Compliance Agreement In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy. Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks. For more… See the full description on the dataset page: https://huggingface.co/datasets/CelestialWandererOfTheVoid/reddit_dataset_231.texttext-classification100K<n<1M0 likes22 downloads2y agoHugging Face29Emilynnjk /celestial-comprehensive-spiritual-ai 🌟 CELESTIAL Comprehensive Spiritual AI Dataset 🚀 SPEED-OPTIMIZED TRAINING - 45-90 MINUTES! Latest Update: Added speed-optimized training notebook that reduces training time from 21+ hours to 45-90 minutes (15-20x faster!) 📊 Dataset Overview Comprehensive spiritual AI training dataset with 3000+ conversations covering all 50+ CELESTIAL spiritual systems including the newly integrated Sanjay Jumaani numerology method. 🎯 Key Features: ⚡… See the full description on the dataset page: https://huggingface.co/datasets/Emilynnjk/celestial-comprehensive-spiritual-ai.text-generation1K<n<10K0 likes21 downloads5mo agoHugging Face30CelestialWandererOfTheVoid /x_dataset_231 Bittensor Subnet 13 X (Twitter) Dataset Miner Data Compliance Agreement In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy. Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/CelestialWandererOfTheVoid/x_dataset_231.texttext-classification10K<n<100K0 likes20 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.