datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AI-CUDA-Engineer-Archive
The AI CUDA Engineer Archive 👷: Agentic CUDA Kernel Discovery, Optimization & Composition
We release The AI CUDA Engineer archive, a dataset consisting of approximately 30,000 CUDA kernels generated by The AI CUDA Engineer. It is released under the CC-By-4.0 license and can be accessed via HuggingFace and interactively visualized here. The dataset is based on the Kernel tasks provided in KernelBench and includes a torch reference implementation, torch, NCU and Clang-tidy… See the full description on the dataset page: https://huggingface.co/datasets/SakanaAI/AI-CUDA-Engineer-Archive.pdfa-eng-wds
Dataset Card for PDF Association dataset (PDFA)
Dataset Summary
PDFA dataset is a document dataset filtered from the SafeDocs corpus, aka CC-MAIN-2021-31-PDF-UNTRUNCATED. The original purpose of that corpus is for comprehensive pdf documents analysis. The purpose of that subset differs in that regard, as focus has been done on making the dataset machine learning-ready for vision-language models.
An example page of one pdf document, with added bounding… See the full description on the dataset page: https://huggingface.co/datasets/JBrightmanAI/pdfa-eng-wds.smollm-chunked
FAISS Indices and Chunked Datasets for SmolLM and SmolLM2 corpora
This repository contains part of the FAISS indices and chunked datasets used for novelty detection for SmolLM and SmolLM2, as presented in the paper LLM generation novelty through the lens of semantic similarity.
Full Documentation
For complete usage instructions, installation guide, and tutorial, please refer to:
Main Tutorial README
Data Distribution
Due to Hugging Face storage quota… See the full description on the dataset page: https://huggingface.co/datasets/enguyen/smollm-chunked.English-Zomi-OPUS_Tatoeba_v20230412
English–Zomi Parallel Corpus (1.78M)
This dataset contains 1.78 million English–Zomi sentence pairs, created to support
machine translation, linguistic research, and large‑scale language model training.
It is fully open and permissively licensed for commercial and non‑commercial use.
🌐 Linguistic Background: Zomi, Tedim Chin, and ISO Codes
Zomi is the endonym (self‑chosen name) of the people and their language.However, Zomi does not yet have an official ISO 639‑3 code.… See the full description on the dataset page: https://huggingface.co/datasets/ZomiLearner/English-Zomi-OPUS_Tatoeba_v20230412.gutenberg_english
Dataset Card for Project Gutenber - English Language eBooks
A collection of non-english language eBooks (48284 rows, 80%+ of all english language books available on the site) from the Project Gutenberg site with metadata removed.
Originally colected for https://github.com/LAION-AI/Open-Assistant (follows the OpenAssistant training format)
The METADATA column contains catalogue meta information on each book as a serialized JSON:
key
original column
language
-
text_id… See the full description on the dataset page: https://huggingface.co/datasets/sedthh/gutenberg_english.pdfa-eng-wds
Dataset Card for PDF Association dataset (PDFA)
Dataset Summary
PDFA dataset is a document dataset filtered from the SafeDocs corpus, aka CC-MAIN-2021-31-PDF-UNTRUNCATED. The original purpose of that corpus is for comprehensive pdf documents analysis. The purpose of that subset differs in that regard, as focus has been done on making the dataset machine learning-ready for vision-language models.
An example page of one pdf document, with added bounding boxes… See the full description on the dataset page: https://huggingface.co/datasets/pixparse/pdfa-eng-wds.CC_eng_urlenglish_quotes_copy
Dataset Card for "english_quotes_copy"
More Information needed
awesome-loop-engineering
Awesome Loop Engineering Dataset
A structured dataset of 1022 papers, official docs, tools, benchmarks, patterns, critiques, and implementation guides for recurring AI-agent systems.
Resource Atlas ·
GitHub field guide ·
Resource selection ·
Report a correction
Dataset Summary
Each row connects an original source to its contribution, novelty, impact, publication details, lifecycle stages, audience, evidence type, link status, and… See the full description on the dataset page: https://huggingface.co/datasets/cy0307/awesome-loop-engineering.agieval-gaokao-english
Dataset Card for "agieval-gaokao-english"
Dataset taken from https://github.com/microsoft/AGIEval and processed as in that repo, following dmayhem93/agieval-* datasets on the HF hub.
This dataset contains the contents of the Gaokao-English subtask of AGIEval, as accessed in https://github.com/ruixiangcui/AGIEval/commit/5c77d073fda993f1652eaae3cf5d04cc5fd21d40 .
Citation:
@misc{zhong2023agieval,
title={AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models}… See the full description on the dataset page: https://huggingface.co/datasets/hails/agieval-gaokao-english.mls_eng
Dataset Card for English MLS
Dataset Summary
This is a streamable version of the English version of the Multilingual LibriSpeech (MLS) dataset.
The data archives were restructured from the original ones from OpenSLR to make it easier to stream.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls_eng.msmarco-v2.1-embed-english-v3
TREC-RAG 2024 Corpus (MSMARCO 2.1) - Encoded with Cohere Embed English v3
This dataset contains the embeddings for the TREC-RAG Corpus 2024 embedded with the Cohere Embed V3 English model.
It contains embeddings for 113,520,750 passages, embeddings for 1677 queries from TREC-Deep Learning 2021-2023, as well as top-1000 hits for all queries using a brute-force (flat) index.
Search over the Index
We have a pre-build index that only requires 300 MB available at… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/msmarco-v2.1-embed-english-v3.ShareGPT-Chinese-English-90k
ShareGPT-Chinese-English-90k Bilingual Human-Machine QA Dataset
A high-quality Chinese-English parallel bilingual human-machine QA dataset, covering user questions in real and complex scenarios. It is used for training high-quality dialogue models (more robust in instruction distribution than those datasets generated by repeatedly calling API interfaces to simulate machine-generated Q&A, like Moss)
Features:
Provides fully semantically equivalent Chinese-English parallel corpus… See the full description on the dataset page: https://huggingface.co/datasets/shareAI/ShareGPT-Chinese-English-90k.english_quotes
Dataset Card for English quotes
I-Dataset Summary
english_quotes is a dataset of all the quotes retrieved from goodreads quotes. This dataset can be used for multi-label text classification and text generation. The content of each quote is in English and concerns the domain of datasets for NLP and beyond.
II-Supported Tasks and Leaderboards
Multi-label text classification : The dataset can be used to train a model for text-classification, which consists of… See the full description on the dataset page: https://huggingface.co/datasets/Abirate/english_quotes.sphere_cohere_embed-english-v3.0fastplus-125m-dataset-eng-6
ISAI - 이사이
I’m an independent developer building and maintaining AI projects on my own.
Everything from model development to server costs, datasets, and feature updates is managed personally.
Any support you can provide greatly helps keep this project running and allows for continuous improvements.
If you find this project helpful, please consider supporting my work. Thank you.
혼자서 AI 프로젝트를 개발하고 운영하고 있습니다.
모델 개발부터 데이터셋 준비, 서버 비용 감당, 기능 업데이트까지 모두 직접 진행하고 있습니다.
보내주시는 따뜻한 후원은 안정적인… See the full description on the dataset page: https://huggingface.co/datasets/aixk/fastplus-125m-dataset-eng-6.rendered-wikipedia-english
Dataset Card for Team-PIXEL/rendered-wikipedia-english
Dataset Summary
This dataset contains the full English Wikipedia from February 1, 2018, rendered into images of 16x8464 resolution.
The original text dataset was built from a Wikipedia dump. Each example in the original text dataset contained the content of one full Wikipedia article with cleaning to strip markdown and unwanted sections (references, etc.). Each rendered example contains a subset of one full article.… See the full description on the dataset page: https://huggingface.co/datasets/Team-PIXEL/rendered-wikipedia-english.so-combined-engThis dataset was created using LeRobot.
Dataset Description
The English version of this dataset integrates 598 open-source community datasets into a single unified corpus, comprising 22,709 episodes and approximately 9.4 million frames across 563 distinct tasks. Several transformations were applied to ensure standardization and data quality:
Camera view normalizationBecause community datasets do not follow a consistent naming scheme for camera viewpoints, we used the… See the full description on the dataset page: https://huggingface.co/datasets/dunnolab/so-combined-eng.ghana-english-asr-2700hrs
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
🇬🇭 Ghana English ASR Dataset
A speech dataset of Ghanaian English extracted from Ghanaian news media broadcasts,
designed for training and fine-tuning Automatic Speech Recognition (ASR) models on
West African English accents.… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-english-asr-2700hrs.beir-embed-english-v3
BEIR embeddings with Cohere embed-english-v3.0 model
This datasets contains all query & document embeddings for BEIR, embedded with the Cohere embed-english-v3.0 embedding model.
Overview of datasets
This repository hosts all 18 datasets from BEIR, including query and document embeddings. The following table gives an overview of the available datasets.
See the next section how to load the individual datasets.
Dataset
nDCG@10
#Documents
arguana
53.98
8,674… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/beir-embed-english-v3.logits-english-512
Dataset Card for "logits-english-512"
More Information needed
CRCD
Comprehensive Robotic Cholecystectomy Dataset (CRCD)
The Comprehensive Robotic Cholecystectomy Dataset (CRCD) is a large-scale, multimodal dataset for robot-assisted surgery (RAS) research.It provides synchronized endoscopic videos, da Vinci surgical robot kinematics, and pedal usage signals, making it one of the most comprehensive open datasets for studying robotic cholecystectomy procedures.
CRCD supports research in:
Medical robotics and surgical automation
Computer vision… See the full description on the dataset page: https://huggingface.co/datasets/SITL-Eng/CRCD.emilia-yodas-english-neucodec
Dataset Card for NeuCodec Emilia-YODAS
Dataset Summary
The NeuCodec Emilia-YODAS dataset is an English-language dataset containing >30M audio samples (>78k hours), taken from the English-language subset of Emilia-YODAS and compressed with NeuCodec.
Usage
import torch
from datasets import load_dataset
from neucodec import NeuCodec
# load dataset and model
dataset = load_dataset("neuphonic/emilia-yodas-english-neucodec", split="train"… See the full description on the dataset page: https://huggingface.co/datasets/neuphonic/emilia-yodas-english-neucodec.FineWiki-eng-mdsrefinedweb-embed-english-v3.0english_dialects
Dataset Card for "english_dialects"
Dataset Summary
This dataset consists of 31 hours of transcribed high-quality audio of English sentences recorded by 120 volunteers speaking with different accents of the British Isles. The dataset is intended for linguistic analysis as well as use for speech technologies. The speakers self-identified as native speakers of Southern England, Midlands, Northern England, Welsh, Scottish and Irish varieties of English.
The recording scripts… See the full description on the dataset page: https://huggingface.co/datasets/ylacombe/english_dialects.english-openlist
English OpenList
The largest open-source, validated English word list for NLP and games.
Dataset Description
English OpenList is a comprehensive, continuously updated dictionary of valid English words. It provides:
~345,000 validated English words, plus a candidate pool of ~9.3 million awaiting evidence
Validation provenance for every word: which sources attested it, and when
Daily updates from authoritative dictionary sources
Version history with changelogs for… See the full description on the dataset page: https://huggingface.co/datasets/ryanjosephkamp/english-openlist.sentence_alignment_dataset-Sinhala-Tamil-English
Dataset summary
This is a gold-standard benchmark dataset for sentence alignment, between Sinhala-English-Tamil languages. Data had been crawled from the following news websites. The aligned documents annotated in the dataset NLPC-UOM/document_alignment_dataset-Sinhala-Tamil-English had been considered to annotate the aligned sentences.
News Source
url
Army
https://www.army.lk/
Hiru
http://www.hirunews.lk
ITN
https://www.newsfirst.lk
Newsfirst
https://www.itnnews.lk… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/sentence_alignment_dataset-Sinhala-Tamil-English.cqadupstack-english
CQADupstackEnglishRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
CQADupStack: A Benchmark Data Set for Community Question-Answering Research
Task category
t2t
Domains
Written
Reference
http://nlp.cis.unimelb.edu.au/resources/cqadupstack/
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["CQADupstackEnglishRetrieval"])
evaluator = mteb.MTEB(task)… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cqadupstack-english.ghana-english-speech-600hrs
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
🇬🇭 Ghana English ASR Dataset
A speech dataset of Ghanaian English extracted from Ghanaian news media broadcasts,
designed for training and fine-tuning Automatic Speech Recognition (ASR) models on
West African English accents.… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-english-speech-600hrs.
