CoolFace
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01beta3 /GridCorpus_9M_Sudoku_Puzzles_Enriched ╔══════════════════════════════════════════════════════════════════════╗ ║ ║ ║ G R I D C O R P U S ║ ║ ║ ║ "004300209005009001070060043..." ║ ║ │ ║ ║ ▼… See the full description on the dataset page: https://huggingface.co/datasets/beta3/GridCorpus_9M_Sudoku_Puzzles_Enriched.tabularfeature-extraction1M<n<10M1 likes1.3k downloads7mo agoHugging Face02AINovice2005 /carbon-cpu-enriched-sequences carbon-cpu-enriched-sequences A CPU-enriched subset of the carbon pretraining corpus (eukaryote_generator), combining original source fields with normalized sequences and row-level features for quality analysis, GPU enrichment and embedding generation. Information of Features Feature Type Description record_id string NCBI Identifier linking the row back to the source genomic record. It provides the primary record-level identity. begin_of_sequence… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-cpu-enriched-sequences.tabulartext-generation10M<n<100M0 likes721 downloads7d agoHugging Face03BRlkl /chatalpaca-multiturn-enriched-2 chatalpaca-multiturn-enriched-2 Records in data.jsonl: 7924 Source dataset: BRlkl/chatalpaca-multiturn-enriched Generated with scenario-guided Samantha multiturn revision texttext-generation1K<n<10K0 likes147 downloads3mo agoHugging Face04rntc /biomed-fr-v3-enriched-softmin-standard biomed-fr-v3-enriched-softmin-standard This dataset is a quality-upsampled version of rntc/biomed-fr-v3-enriched using soft-min bottleneck sampling. Preprocessing Method Soft-min calculation: Formula: s = (mean(q_k^p))^(1/p) where q_k are the 4 quality scores Parameter p = -2.0 Weight computation: Ratio preference (5 vs 1): R = 10 Gamma exponent: γ = 1.43 (computed as log(R)/log(5)) Weight formula: w = s^γ Floor: w = max(w, median(w) × 0.05) Resampling: Target size:… See the full description on the dataset page: https://huggingface.co/datasets/rntc/biomed-fr-v3-enriched-softmin-standard.tabulartext-generation1M<n<10M0 likes98 downloads1y agoHugging Face05Avinaash /ChartSense_8645_web_enriched ChartSense 8645 web enriched This is a supervised fine tuning dataset that teaches a small language model to behave like a data analyst on chart and data visualization work, i.e. it should be able to see through the user's words, resolve underspecified asks, push back if required, critique flawed charts, and answer the question the chart is a means to. Built for the AutoScientist Challenge by Adaption Labs. This is the web-enriched build. The web slice is grown to its full… See the full description on the dataset page: https://huggingface.co/datasets/Avinaash/ChartSense_8645_web_enriched.texttext-generation10K<n<100K1 likes89 downloads1mo agoHugging Face06jedisct1 /enriched-golang-finetune-dataset Enriched Commit Diff Fine-tuning Dataset Generated from jedisct1/golang on 2026-06-07T17:24:24.423892+00:00. Each kept source commit produces four supervised fine-tuning variants: message_to_diff: original commit message -> original diff diff_to_message: original diff -> original commit message minimized_message_to_diff: concise/minimized commit message -> original diff message_to_shuffled_diff: original commit message -> original diff with file blocks shuffled Output files… See the full description on the dataset page: https://huggingface.co/datasets/jedisct1/enriched-golang-finetune-dataset.texttext-generation1M<n<10M0 likes68 downloads4mo agoHugging Face07rntc /biomed-fr-v3-enriched-softmin-leger biomed-fr-v3-enriched-softmin-leger This dataset is a quality-upsampled version of rntc/biomed-fr-v3-enriched using soft-min bottleneck sampling. Preprocessing Method Soft-min calculation: Formula: s = (mean(q_k^p))^(1/p) where q_k are the 4 quality scores Parameter p = -0.7 Weight computation: Ratio preference (5 vs 1): R = 5 Gamma exponent: γ = 1.00 (computed as log(R)/log(5)) Weight formula: w = s^γ Floor: w = max(w, median(w) × 0.05) Resampling: Target size: Same… See the full description on the dataset page: https://huggingface.co/datasets/rntc/biomed-fr-v3-enriched-softmin-leger.tabulartext-generation1M<n<10M0 likes55 downloads1y agoHugging Face08jedisct1 /enriched-rust-finetune-dataset Enriched Commit Diff Fine-tuning Dataset Generated from jedisct1/rust on 2026-06-07T17:21:21.169834+00:00. Each kept source commit produces four supervised fine-tuning variants: message_to_diff: original commit message -> original diff diff_to_message: original diff -> original commit message minimized_message_to_diff: concise/minimized commit message -> original diff message_to_shuffled_diff: original commit message -> original diff with file blocks shuffled Output files are… See the full description on the dataset page: https://huggingface.co/datasets/jedisct1/enriched-rust-finetune-dataset.texttext-generation1M<n<10M1 likes52 downloads4mo agoHugging Face09BRlkl /chatalpaca-multiturn-enriched-3.5 chatalpaca-multiturn-enriched-3.5 This dataset combines the existing Samantha A10 multiturn corpus with new long-memory and exact-answer specialist conversations. Splits train: 18,801 rows (existing, manual-evaluation, and generated rows) No separate validation split is published; all records remain in train. Total: 18,801 rows Composition Existing source artifact: BRlkl/chatalpaca-multiturn-enriched-2.1 New long-memory rows: 8,000 New arithmetic… See the full description on the dataset page: https://huggingface.co/datasets/BRlkl/chatalpaca-multiturn-enriched-3.5.texttext-generation10K<n<100K0 likes48 downloads1mo agoHugging Face10Avinaash /ChartSense_8827_corpus_enriched ChartSense 8827 corpus enriched This is a supervised fine tuning dataset that teaches a small language model to behave like a data analyst on chart and data visualization work, i.e. it should be able to see through the user's words, resolve underspecified asks, push back if required, critique flawed charts, and answer the question the chart is a means to. Built for the AutoScientist Challenge by Adaption Labs. This is the corpus-enriched build. The corpus slice is grown to its… See the full description on the dataset page: https://huggingface.co/datasets/Avinaash/ChartSense_8827_corpus_enriched.texttext-generation10K<n<100K0 likes43 downloads1mo agoHugging Face11BRlkl /chatalpaca-multiturn-enriched-3 chatalpaca-multiturn-enriched-3 This dataset combines the existing Samantha A10 multiturn corpus with new long-memory and exact-answer specialist conversations. Splits train: 18,914 rows (existing, manual-evaluation, and generated rows) No separate validation split is published; all records remain in train. Total: 18,914 rows Composition Existing source artifact: BRlkl/chatalpaca-multiturn-enriched-2.1 New long-memory rows: 8,000 New arithmetic… See the full description on the dataset page: https://huggingface.co/datasets/BRlkl/chatalpaca-multiturn-enriched-3.texttext-generation10K<n<100K0 likes37 downloads2mo agoHugging Face12BRlkl /chatalpaca-multiturn-enriched-3-26b chatalpaca-multiturn-enriched-3-26b This dataset combines the existing Samantha A10 multiturn corpus with new long-memory and exact-answer specialist conversations. Splits train: 8,038 rows validation: 20 locked manual-evaluation rows from the prior dataset Total: 8,058 rows Composition Existing source artifact: BRlkl/chatalpaca-multiturn-enriched-2.1 New long-memory rows: 80 New arithmetic rows: 10 New logic rows: 10 New conversational-QA rows:… See the full description on the dataset page: https://huggingface.co/datasets/BRlkl/chatalpaca-multiturn-enriched-3-26b.texttext-generation1K<n<10K0 likes34 downloads2mo agoHugging Face13BRlkl /chatalpaca-multiturn-enriched-probe ChatAlpaca Multiturn Enriched Probe Deterministic transcript-memory probe dataset for Samantha multiturn latent-state pretraining. Source dataset: BRlkl/chatalpaca-multiturn-enriched Each conversation keeps the original messages and adds state_supervision with one fixed probe for every prefix after the first user/assistant pair. Fixed probe question: What is everything we have talked about so far? Give exact conversation transcript verbatim in following format: [User 1]: X… See the full description on the dataset page: https://huggingface.co/datasets/BRlkl/chatalpaca-multiturn-enriched-probe.texttext-generation1K<n<10K0 likes31 downloads3mo agoHugging Face14BRlkl /chatalpaca-multiturn-enriched-3-4breasoning chatalpaca-multiturn-enriched-3-4breasoning This dataset combines the existing Samantha A10 multiturn corpus with new long-memory and exact-answer specialist conversations. Splits train: 8,039 rows validation: 20 locked manual-evaluation rows from the prior dataset Total: 8,059 rows Composition Existing source artifact: BRlkl/chatalpaca-multiturn-enriched-2.1 New long-memory rows: 80 New arithmetic rows: 10 New logic rows: 10 New… See the full description on the dataset page: https://huggingface.co/datasets/BRlkl/chatalpaca-multiturn-enriched-3-4breasoning.texttext-generation1K<n<10K0 likes29 downloads2mo agoHugging Face15rntc /biomed-fr-v3-enriched-softmin-tres_agressif biomed-fr-v3-enriched-softmin-tres_agressif This dataset is a quality-upsampled version of rntc/biomed-fr-v3-enriched using soft-min bottleneck sampling. Preprocessing Method Soft-min calculation: Formula: s = (mean(q_k^p))^(1/p) where q_k are the 4 quality scores Parameter p = -4.0 Weight computation: Ratio preference (5 vs 1): R = 20 Gamma exponent: γ = 1.86 (computed as log(R)/log(5)) Weight formula: w = s^γ Floor: w = max(w, median(w) × 0.02) Resampling: Target… See the full description on the dataset page: https://huggingface.co/datasets/rntc/biomed-fr-v3-enriched-softmin-tres_agressif.tabulartext-generation1M<n<10M0 likes24 downloads1y agoHugging Face16rntc /biomed-fr-v3-enriched-softmin-tres_agressif_min2 biomed-fr-v3-enriched-softmin-tres_agressif_min2 This dataset is a quality-upsampled version of rntc/biomed-fr-v3-enriched using soft-min bottleneck sampling. Preprocessing Method Preprocessing steps: Pre-filtering: Removed all rows with any quality score < 2.0 (i.e., removed rows with any score = 1) Soft-min calculation: Formula: s = (mean(q_k^p))^(1/p) where q_k are the 4 quality scores Parameter p = -4.0 Weight computation: Ratio preference (5 vs 1): R = 20 Gamma… See the full description on the dataset page: https://huggingface.co/datasets/rntc/biomed-fr-v3-enriched-softmin-tres_agressif_min2.tabulartext-generation1M<n<10M0 likes23 downloads1y agoHugging Face17sovereign3b /ZK-Enriched ZK-Enriched: AI-Generated Analysis of Zero-Knowledge Cryptography Code AI-generated explanations and concept extraction from 22 open-source zero-knowledge cryptography projects. Created autonomously on the Dria decentralized inference network. Dataset Statistics Metric Value Total entries 18,503 Code analyses 13,884 Documentation summaries 4,619 Total content ~9.0M tokens Avg explanation length 1,753 characters Avg concepts length 257 characters Avg… See the full description on the dataset page: https://huggingface.co/datasets/sovereign3b/ZK-Enriched.texttext-generation10K<n<100K1 likes11 downloads6mo agoHugging Face18mmmaurer /enriched-generated-arguments Info This is a version of a generated arguments corpus enriched with linguistic features and argument quality dimensions. The linguistic features were extracted with elfen. The argument quality dimensions were extracte with these adapters. Citation If you use this enriched version of the generated arguments corpus, please cite @inproceedings{doenmez-maurer-2025-ai, title = "AI Argues Differently: Distinct Argumentative and Linguistic Patterns of LLMs in Persuasive… See the full description on the dataset page: https://huggingface.co/datasets/mmmaurer/enriched-generated-arguments.tabulartext-classification10K<n<100K0 likes10 downloads1y agoHugging Face19FatimahEmadEldin /icaire-ai-glossary-enriched ICAIRE AI Glossary — Enriched (Mustalih Living) Bilingual Arabic-English AI glossary based on the ICAIRE canonical vocabulary, enriched through a multi-layer LLM pipeline into a fully structured multimodal dataset: metaphors, detailed explanations, UML diagrams, typed knowledge-graph edges, and narrator-voice story-track assignments. Dataset structure Each term (1,242 total) is one JSON record with these fields: Field Type Description english_term string… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/icaire-ai-glossary-enriched.text-generation1K<n<10K0 likes7 downloads5mo agoHugging Face20Nini0la /edgeimci-beta0-1k-multitask-enriched-2258-v1 EdgeIMCI Beta0-1K Multitask Enriched 2258 Dataset summary This is the public, message-normalized publication of the dataset used to fine-tune the selected EdgeIMCI 4B-alpha-second research checkpoint from the pinned Qwen/Qwen3-4B base model. It combines structured extraction, routing, bounded clinical-response, project self-knowledge, scope-safety, and proposition/negation examples for the EdgeIMCI sick-child assessment workflow. The dataset is research evidence… See the full description on the dataset page: https://huggingface.co/datasets/Nini0la/edgeimci-beta0-1k-multitask-enriched-2258-v1.texttext-generation1K<n<10K0 likes8h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.