datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ncbi-genbank-complete
Dataset Card for NCBI GenBank Complete
Dataset Summary
GenBank® is the NIH genetic sequence database, an annotated collection of all publicly available DNA sequences. GenBank is part of the International Nucleotide Sequence Database Collaboration (INSDC), which comprises the DNA DataBank of Japan (DDBJ), the European Nucleotide Archive (ENA), and GenBank at NCBI. These three organizations exchange data on a daily basis.
This dataset has been processed into a… See the full description on the dataset page: https://huggingface.co/datasets/pulmo/ncbi-genbank-complete.Complete_Data_Source_100K_HOURS
Multi-Language Audio Collection (100K Hours)
This repository is physically reorganized for Absolute 100% Data Visibility.
🏗️ Global Consolidator
Select your language subset to listen to high-quality waveform audio. All shards from legacy and modern pipelines are automatically routed here.
ncbi-refseq-complete
Dataset Card for NCBI RefSeq Complete
Dataset Summary
The NCBI Reference Sequence (RefSeq) complete dataset provides a comprehensive, integrated, non-redundant, well-annotated set of sequences, including genomic DNA, transcripts, and proteins. It serves as a stable reference for genome annotation, gene identification and characterization, mutation and polymorphism analysis, expression studies, and comparative analyses.
This dataset has been processed into a… See the full description on the dataset page: https://huggingface.co/datasets/huggingworld/ncbi-refseq-complete.arxiv-complete
arXiv Complete Corpus
A snapshot of arXiv's metadata, version history, submission files and rendered
documents. It covers 3,148,796 papers and includes file contents, paths, sizes
and SHA-256 digests. Metadata comes from arXiv's OAI-PMH arXivRaw interface;
files come from the GCS mirror, S3 source archives and direct PDF fetches.
This release holds a PDF for 99.47% of papers and 99.54% of versions reported
with a non-zero submission size. It is a one-off snapshot; coverage gaps… See the full description on the dataset page: https://huggingface.co/datasets/secemp9/arxiv-complete.Jailbreak_Complete_DS_labeledOmni-CAD-Subset-CompleteBraTS-2024-Complete
BraTS 2024 Complete Prepared Dataset
Brain Tumor Segmentation (Leave a like 💖 if this helped you)
Dataset Description
This is an organized and verified version of the BraTS 2024 challenge datasets, including three tumor types.
Included Datasets
Dataset
Type
Cases
Source
BraTS-GLI
Glioma
1,809
Synapse (Dec 2024)
BraTS-MEN-RT
Meningioma + RT
571
Synapse (Feb 2025)
BraTS-PED
Pediatric
348
Cancer Imaging Archive… See the full description on the dataset page: https://huggingface.co/datasets/Spirit-26/BraTS-2024-Complete.dementor-complete-experiment-results
Dementor complete experiment results
Audited outputs for the configuration-defined Dementor completion campaign.
Audited scope
Behavioral imitation adapters: 1,104 total (528 SFT, 528 DPO, 48 self-SFT controls).
Behavioral-fidelity evaluation: 1,104 adapters on 200 held-out prompts, with embedding and
primary LLM-judge scores, plus 48 target-reference response sets.
Activation steering: 29 models, seven benchmarks, and two operators (original and fpall),
totaling… See the full description on the dataset page: https://huggingface.co/datasets/dementor-research/dementor-complete-experiment-results.pexel-0808-complete-final-testGithub Page: https://github.com/UmiMarch/OpenVideo
license: cc-by-4.0
task_categories:
- video-text-to-text
size_categories:
- 100K<n<1M
features_mavos_completeComplete-FABLE.5-traces-2M
license: mit
pretty_name: Claude Library — Fable 5 · Opus · Sonnet
annotations_creators:
machine-generated
language:
en
language_creators:
found
machine-generated
multilinguality:
monolingual
size_categories:
10K<n<100K
task_categories:
text-generation
task_ids:
language-modeling
tags:
agent-traces
claude
claude-fable-5
claude-opus
claude-sonnet
chain-of-thought
tool-use
coding-agents
content-verified
maintained-mirror
deduplicated
parquet
configs:
config_name:… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Complete-FABLE.5-traces-2M.tvelve_map_complete_datasetzcache-results-completedutch-tts-labeled-complete
Dutch TTS Dataset - Complete Labeled
A comprehensive Dutch text-to-speech dataset with 596,508 audio samples totaling 234GB of audio data.
Quick Preview
The default config shows a 100-row sample for the dataset viewer. To access the full dataset, use the full config.
Dataset Description
This dataset contains Dutch speech recordings with rich metadata including:
Emotion labels (neutral, happy, sad, angry)
Speaker IDs (239,388 unique speakers)… See the full description on the dataset page: https://huggingface.co/datasets/AITRADER/dutch-tts-labeled-complete.wiki_book_corpus_complete_processed_bert_dataset
Dataset Card for "wiki_book_corpus_complete_processed_bert_dataset"
More Information needed
coco_captioning_complete_formatCOMPLETE_COAL_DATASETleetcode-complete
Complete LeetCode Problems Dataset
This dataset contains a comprehensive collection of LeetCode problems (including premium) with AI-generated solutions in JSONL format. It is regularly updated to include new problems as they are added to LeetCode.
Splits
The dataset is divided into the following splits:
train: Contains approximately 80% of the problems for training
validation: Contains approximately 10% of the problems for validation
test: Contains approximately… See the full description on the dataset page: https://huggingface.co/datasets/whiskwhite/leetcode-complete.booksum-complete-cleaned
Description:
This repository contains the Booksum dataset introduced in the paper BookSum: A Collection of Datasets for Long-form Narrative Summarization
.
This dataset includes both book and chapter summaries from the BookSum dataset (unlike the kmfoda/booksum one which only contains the chapter dataset). Some mismatched summaries have been corrected. Uneccessary columns have been discarded. Contains minimal text-to-summary rows. As there are multiple summaries for a given text… See the full description on the dataset page: https://huggingface.co/datasets/ubaada/booksum-complete-cleaned.ISEAR-dataset-completewiki_book_corpus_complete_raw_dataset
Dataset Card for "wiki_book_corpus_complete_raw_dataset"
More Information needed
GOD_Coder_Complete_DataSet
GOD_Coder_Complete_DataSet
Subtitle
A large-scale complete-project coding dataset by gss1147 / WithIn Us AI, built to train language models into stronger professional software-engineering assistants.
Dataset Summary
GOD_Coder_Complete_DataSet is a large synthetic supervised fine-tuning dataset designed to help turn a general language model into a professional complete-project AI coder.
The dataset focuses on teaching models how to:
diagnose… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/GOD_Coder_Complete_DataSet.Visdrone_fisheye-v51-complete
Dataset Card for visdrone_fisheye-v51-complete
This is a FiftyOne dataset with 20942 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Abeyankar/Visdrone_fisheye-v51-complete")
# Launch the App
session =… See the full description on the dataset page: https://huggingface.co/datasets/Abeyankar/Visdrone_fisheye-v51-complete.BraTS-2024-Complete
BraTS 2024 Complete Prepared Dataset
Brain Tumor Segmentation
Dataset Description
This is an organized and verified version of the BraTS 2024 challenge datasets, including three tumor types.
Included Datasets
Dataset
Type
Cases
Source
BraTS-GLI
Glioma
1,809
Synapse (Dec 2024)
BraTS-MEN-RT
Meningioma + RT
571
Synapse (Feb 2025)
BraTS-PED
Pediatric
348
Cancer Imaging Archive
Total: 2,728 multi-parametric MRI cases… See the full description on the dataset page: https://huggingface.co/datasets/Aff77/BraTS-2024-Complete.vie-speech-corpus-completeBenetech_PlotQa_DVQA_combined_matcha_completemeta-llama_Llama-3.1-8B-Instruct-jdgfct-Completenesswikipedia-paragraph-embeddings-en-gist-complete
Dataset Summary
Paragraph embeddings for every article in English Wikipedia (not the Simple English version).
Based on wikimedia/wikipedia, 20231101.en.
Embeddings were generated with avsolatorio/GIST-small-Embedding-v0
and are quantized to int8.
You can load the data with the following:
from datasets import load_dataset
ds = load_dataset(path="Abrak/wikipedia-paragraph-embeddings-en-gist-complete", data-dir="20231101.en")
Dataset Structure
The structure of the… See the full description on the dataset page: https://huggingface.co/datasets/Abrak/wikipedia-paragraph-embeddings-en-gist-complete.sat-mathematics-complete-questionscomplete-voiceai-speech-dataset
Silencio Voice AI Sample Dataset
Speaker-attributed spontaneous speech. 363 labelled contributors across 111 self-reported origin varieties, recorded on their own devices in their own environments. Every clip carries origin region / variety, mother tongue, gender, device, OS, recording environment.
Hours
16.05
Clips
1,305
Speakers
363
Origin varieties
111
Languages
21
Configs
44
Speaker metadata
origin region / variety, mother tongue, gender, device, OS… See the full description on the dataset page: https://huggingface.co/datasets/SilencioNetwork/complete-voiceai-speech-dataset.
