datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GridCorpus_9M_Sudoku_Puzzles_Enriched
╔══════════════════════════════════════════════════════════════════════╗
║ ║
║ G R I D C O R P U S ║
║ ║
║ "004300209005009001070060043..." ║
║ │ ║
║ ▼… See the full description on the dataset page: https://huggingface.co/datasets/beta3/GridCorpus_9M_Sudoku_Puzzles_Enriched.carbon-cpu-enriched-sequences
carbon-cpu-enriched-sequences
A CPU-enriched subset of the carbon pretraining corpus (eukaryote_generator), combining original source fields with normalized sequences
and row-level features for quality analysis, GPU enrichment and embedding generation.
Information of Features
Feature
Type
Description
record_id
string
NCBI Identifier linking the row back to the source genomic record. It provides the primary record-level identity.
begin_of_sequence… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-cpu-enriched-sequences.chatalpaca-multiturn-enriched-2
chatalpaca-multiturn-enriched-2
Records in data.jsonl: 7924
Source dataset: BRlkl/chatalpaca-multiturn-enriched
Generated with scenario-guided Samantha multiturn revision
biomed-fr-v3-enriched-softmin-standard
biomed-fr-v3-enriched-softmin-standard
This dataset is a quality-upsampled version of rntc/biomed-fr-v3-enriched using soft-min bottleneck sampling.
Preprocessing Method
Soft-min calculation:
Formula: s = (mean(q_k^p))^(1/p) where q_k are the 4 quality scores
Parameter p = -2.0
Weight computation:
Ratio preference (5 vs 1): R = 10
Gamma exponent: γ = 1.43 (computed as log(R)/log(5))
Weight formula: w = s^γ
Floor: w = max(w, median(w) × 0.05)
Resampling:
Target size:… See the full description on the dataset page: https://huggingface.co/datasets/rntc/biomed-fr-v3-enriched-softmin-standard.ChartSense_8645_web_enriched
ChartSense 8645 web enriched
This is a supervised fine tuning dataset that teaches a small language model to
behave like a data analyst on chart and data visualization work, i.e.
it should be able to see through the user's words,
resolve underspecified asks,
push back if required,
critique flawed charts, and
answer the question the chart is a means to.
Built for the AutoScientist Challenge by Adaption Labs.
This is the web-enriched build. The web slice is grown to its full… See the full description on the dataset page: https://huggingface.co/datasets/Avinaash/ChartSense_8645_web_enriched.enriched-golang-finetune-dataset
Enriched Commit Diff Fine-tuning Dataset
Generated from jedisct1/golang on 2026-06-07T17:24:24.423892+00:00.
Each kept source commit produces four supervised fine-tuning variants:
message_to_diff: original commit message -> original diff
diff_to_message: original diff -> original commit message
minimized_message_to_diff: concise/minimized commit message -> original diff
message_to_shuffled_diff: original commit message -> original diff with file blocks shuffled
Output files… See the full description on the dataset page: https://huggingface.co/datasets/jedisct1/enriched-golang-finetune-dataset.biomed-fr-v3-enriched-softmin-leger
biomed-fr-v3-enriched-softmin-leger
This dataset is a quality-upsampled version of rntc/biomed-fr-v3-enriched using soft-min bottleneck sampling.
Preprocessing Method
Soft-min calculation:
Formula: s = (mean(q_k^p))^(1/p) where q_k are the 4 quality scores
Parameter p = -0.7
Weight computation:
Ratio preference (5 vs 1): R = 5
Gamma exponent: γ = 1.00 (computed as log(R)/log(5))
Weight formula: w = s^γ
Floor: w = max(w, median(w) × 0.05)
Resampling:
Target size: Same… See the full description on the dataset page: https://huggingface.co/datasets/rntc/biomed-fr-v3-enriched-softmin-leger.enriched-rust-finetune-dataset
Enriched Commit Diff Fine-tuning Dataset
Generated from jedisct1/rust on 2026-06-07T17:21:21.169834+00:00.
Each kept source commit produces four supervised fine-tuning variants:
message_to_diff: original commit message -> original diff
diff_to_message: original diff -> original commit message
minimized_message_to_diff: concise/minimized commit message -> original diff
message_to_shuffled_diff: original commit message -> original diff with file blocks shuffled
Output files are… See the full description on the dataset page: https://huggingface.co/datasets/jedisct1/enriched-rust-finetune-dataset.chatalpaca-multiturn-enriched-3.5
chatalpaca-multiturn-enriched-3.5
This dataset combines the existing Samantha A10 multiturn corpus with new long-memory and exact-answer specialist conversations.
Splits
train: 18,801 rows (existing, manual-evaluation, and generated rows)
No separate validation split is published; all records remain in train.
Total: 18,801 rows
Composition
Existing source artifact: BRlkl/chatalpaca-multiturn-enriched-2.1
New long-memory rows: 8,000
New arithmetic… See the full description on the dataset page: https://huggingface.co/datasets/BRlkl/chatalpaca-multiturn-enriched-3.5.ChartSense_8827_corpus_enriched
ChartSense 8827 corpus enriched
This is a supervised fine tuning dataset that teaches a small language model to
behave like a data analyst on chart and data visualization work, i.e.
it should be able to see through the user's words,
resolve underspecified asks,
push back if required,
critique flawed charts, and
answer the question the chart is a means to.
Built for the AutoScientist Challenge by Adaption Labs.
This is the corpus-enriched build. The corpus slice is grown to its… See the full description on the dataset page: https://huggingface.co/datasets/Avinaash/ChartSense_8827_corpus_enriched.chatalpaca-multiturn-enriched-3
chatalpaca-multiturn-enriched-3
This dataset combines the existing Samantha A10 multiturn corpus with new long-memory and exact-answer specialist conversations.
Splits
train: 18,914 rows (existing, manual-evaluation, and generated rows)
No separate validation split is published; all records remain in train.
Total: 18,914 rows
Composition
Existing source artifact: BRlkl/chatalpaca-multiturn-enriched-2.1
New long-memory rows: 8,000
New arithmetic… See the full description on the dataset page: https://huggingface.co/datasets/BRlkl/chatalpaca-multiturn-enriched-3.chatalpaca-multiturn-enriched-3-26b
chatalpaca-multiturn-enriched-3-26b
This dataset combines the existing Samantha A10 multiturn corpus with new long-memory and exact-answer specialist conversations.
Splits
train: 8,038 rows
validation: 20 locked manual-evaluation rows from the prior dataset
Total: 8,058 rows
Composition
Existing source artifact: BRlkl/chatalpaca-multiturn-enriched-2.1
New long-memory rows: 80
New arithmetic rows: 10
New logic rows: 10
New conversational-QA rows:… See the full description on the dataset page: https://huggingface.co/datasets/BRlkl/chatalpaca-multiturn-enriched-3-26b.chatalpaca-multiturn-enriched-probe
ChatAlpaca Multiturn Enriched Probe
Deterministic transcript-memory probe dataset for Samantha multiturn latent-state pretraining.
Source dataset: BRlkl/chatalpaca-multiturn-enriched
Each conversation keeps the original messages and adds state_supervision with one fixed probe for every prefix after the first user/assistant pair.
Fixed probe question:
What is everything we have talked about so far? Give exact conversation transcript verbatim in following format: [User 1]: X… See the full description on the dataset page: https://huggingface.co/datasets/BRlkl/chatalpaca-multiturn-enriched-probe.chatalpaca-multiturn-enriched-3-4breasoning
chatalpaca-multiturn-enriched-3-4breasoning
This dataset combines the existing Samantha A10 multiturn corpus with new long-memory and exact-answer specialist conversations.
Splits
train: 8,039 rows
validation: 20 locked manual-evaluation rows from the prior dataset
Total: 8,059 rows
Composition
Existing source artifact: BRlkl/chatalpaca-multiturn-enriched-2.1
New long-memory rows: 80
New arithmetic rows: 10
New logic rows: 10
New… See the full description on the dataset page: https://huggingface.co/datasets/BRlkl/chatalpaca-multiturn-enriched-3-4breasoning.biomed-fr-v3-enriched-softmin-tres_agressif
biomed-fr-v3-enriched-softmin-tres_agressif
This dataset is a quality-upsampled version of rntc/biomed-fr-v3-enriched using soft-min bottleneck sampling.
Preprocessing Method
Soft-min calculation:
Formula: s = (mean(q_k^p))^(1/p) where q_k are the 4 quality scores
Parameter p = -4.0
Weight computation:
Ratio preference (5 vs 1): R = 20
Gamma exponent: γ = 1.86 (computed as log(R)/log(5))
Weight formula: w = s^γ
Floor: w = max(w, median(w) × 0.02)
Resampling:
Target… See the full description on the dataset page: https://huggingface.co/datasets/rntc/biomed-fr-v3-enriched-softmin-tres_agressif.biomed-fr-v3-enriched-softmin-tres_agressif_min2
biomed-fr-v3-enriched-softmin-tres_agressif_min2
This dataset is a quality-upsampled version of rntc/biomed-fr-v3-enriched using soft-min bottleneck sampling.
Preprocessing Method
Preprocessing steps:
Pre-filtering: Removed all rows with any quality score < 2.0 (i.e., removed rows with any score = 1)
Soft-min calculation:
Formula: s = (mean(q_k^p))^(1/p) where q_k are the 4 quality scores
Parameter p = -4.0
Weight computation:
Ratio preference (5 vs 1): R = 20
Gamma… See the full description on the dataset page: https://huggingface.co/datasets/rntc/biomed-fr-v3-enriched-softmin-tres_agressif_min2.ZK-Enriched
ZK-Enriched: AI-Generated Analysis of Zero-Knowledge Cryptography Code
AI-generated explanations and concept extraction from 22 open-source zero-knowledge cryptography projects. Created autonomously on the Dria decentralized inference network.
Dataset Statistics
Metric
Value
Total entries
18,503
Code analyses
13,884
Documentation summaries
4,619
Total content
~9.0M tokens
Avg explanation length
1,753 characters
Avg concepts length
257 characters
Avg… See the full description on the dataset page: https://huggingface.co/datasets/sovereign3b/ZK-Enriched.enriched-generated-arguments
Info
This is a version of a generated arguments corpus enriched with linguistic features and argument quality dimensions.
The linguistic features were extracted with elfen.
The argument quality dimensions were extracte with these adapters.
Citation
If you use this enriched version of the generated arguments corpus, please cite
@inproceedings{doenmez-maurer-2025-ai,
title = "AI Argues Differently: Distinct Argumentative and Linguistic Patterns of LLMs in Persuasive… See the full description on the dataset page: https://huggingface.co/datasets/mmmaurer/enriched-generated-arguments.icaire-ai-glossary-enriched
ICAIRE AI Glossary — Enriched (Mustalih Living)
Bilingual Arabic-English AI glossary based on the ICAIRE canonical vocabulary,
enriched through a multi-layer LLM pipeline into a fully structured
multimodal dataset: metaphors, detailed explanations, UML diagrams, typed
knowledge-graph edges, and narrator-voice story-track assignments.
Dataset structure
Each term (1,242 total) is one JSON record with these fields:
Field
Type
Description
english_term
string… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/icaire-ai-glossary-enriched.edgeimci-beta0-1k-multitask-enriched-2258-v1
EdgeIMCI Beta0-1K Multitask Enriched 2258
Dataset summary
This is the public, message-normalized publication of the dataset used to fine-tune the selected EdgeIMCI 4B-alpha-second research checkpoint from the pinned Qwen/Qwen3-4B base model. It combines structured extraction, routing, bounded clinical-response, project self-knowledge, scope-safety, and proposition/negation examples for the EdgeIMCI sick-child assessment workflow.
The dataset is research evidence… See the full description on the dataset page: https://huggingface.co/datasets/Nini0la/edgeimci-beta0-1k-multitask-enriched-2258-v1.
