CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SakanaAI /AI-CUDA-Engineer-Archive The AI CUDA Engineer Archive 👷: Agentic CUDA Kernel Discovery, Optimization & Composition We release The AI CUDA Engineer archive, a dataset consisting of approximately 30,000 CUDA kernels generated by The AI CUDA Engineer. It is released under the CC-By-4.0 license and can be accessed via HuggingFace and interactively visualized here. The dataset is based on the Kernel tasks provided in KernelBench and includes a torch reference implementation, torch, NCU and Clang-tidy… See the full description on the dataset page: https://huggingface.co/datasets/SakanaAI/AI-CUDA-Engineer-Archive.tabular10K<n<100K227 likes136k downloads2y agoHugging Face02JBrightmanAI /pdfa-eng-wds Dataset Card for PDF Association dataset (PDFA) Dataset Summary PDFA dataset is a document dataset filtered from the SafeDocs corpus, aka CC-MAIN-2021-31-PDF-UNTRUNCATED. The original purpose of that corpus is for comprehensive pdf documents analysis. The purpose of that subset differs in that regard, as focus has been done on making the dataset machine learning-ready for vision-language models. An example page of one pdf document, with added bounding… See the full description on the dataset page: https://huggingface.co/datasets/JBrightmanAI/pdfa-eng-wds.image-to-text10M<n<100M0 likes47k downloads2mo agoHugging Face03enguyen /smollm-chunked FAISS Indices and Chunked Datasets for SmolLM and SmolLM2 corpora This repository contains part of the FAISS indices and chunked datasets used for novelty detection for SmolLM and SmolLM2, as presented in the paper LLM generation novelty through the lens of semantic similarity. Full Documentation For complete usage instructions, installation guide, and tutorial, please refer to: Main Tutorial README Data Distribution Due to Hugging Face storage quota… See the full description on the dataset page: https://huggingface.co/datasets/enguyen/smollm-chunked.tabulartext-retrieval100M<n<1B1 likes35k downloads7mo agoHugging Face04ZomiLearner /English-Zomi-OPUS_Tatoeba_v20230412 English–Zomi Parallel Corpus (1.78M) This dataset contains 1.78 million English–Zomi sentence pairs, created to support machine translation, linguistic research, and large‑scale language model training. It is fully open and permissively licensed for commercial and non‑commercial use. 🌐 Linguistic Background: Zomi, Tedim Chin, and ISO Codes Zomi is the endonym (self‑chosen name) of the people and their language.However, Zomi does not yet have an official ISO 639‑3 code.… See the full description on the dataset page: https://huggingface.co/datasets/ZomiLearner/English-Zomi-OPUS_Tatoeba_v20230412.tabulartranslation1M<n<10M0 likes15k downloads7mo agoHugging Face05sedthh /gutenberg_english Dataset Card for Project Gutenber - English Language eBooks A collection of non-english language eBooks (48284 rows, 80%+ of all english language books available on the site) from the Project Gutenberg site with metadata removed. Originally colected for https://github.com/LAION-AI/Open-Assistant (follows the OpenAssistant training format) The METADATA column contains catalogue meta information on each book as a serialized JSON: key original column language - text_id… See the full description on the dataset page: https://huggingface.co/datasets/sedthh/gutenberg_english.texttext-generation10K<n<100K39 likes12k downloads4y agoHugging Face06pixparse /pdfa-eng-wds Dataset Card for PDF Association dataset (PDFA) Dataset Summary PDFA dataset is a document dataset filtered from the SafeDocs corpus, aka CC-MAIN-2021-31-PDF-UNTRUNCATED. The original purpose of that corpus is for comprehensive pdf documents analysis. The purpose of that subset differs in that regard, as focus has been done on making the dataset machine learning-ready for vision-language models. An example page of one pdf document, with added bounding boxes… See the full description on the dataset page: https://huggingface.co/datasets/pixparse/pdfa-eng-wds.textimage-to-text1K<n<10K161 likes9k downloads2y agoHugging Face07leiwx52 /CC_eng_urltext100M<n<1B0 likes7.1k downloads2y agoHugging Face08ybelkada /english_quotes_copy Dataset Card for "english_quotes_copy" More Information needed text1K<n<10K1 likes6.6k downloads3y agoHugging Face09cy0307 /awesome-loop-engineering Awesome Loop Engineering Dataset A structured dataset of 1022 papers, official docs, tools, benchmarks, patterns, critiques, and implementation guides for recurring AI-agent systems. Resource Atlas · GitHub field guide · Resource selection · Report a correction &nbsp; Dataset Summary Each row connects an original source to its contribution, novelty, impact, publication details, lifecycle stages, audience, evidence type, link status, and… See the full description on the dataset page: https://huggingface.co/datasets/cy0307/awesome-loop-engineering.imagetext-classification1K<n<10K3 likes5.2k downloads22h agoHugging Face10hails /agieval-gaokao-english Dataset Card for "agieval-gaokao-english" Dataset taken from https://github.com/microsoft/AGIEval and processed as in that repo, following dmayhem93/agieval-* datasets on the HF hub. This dataset contains the contents of the Gaokao-English subtask of AGIEval, as accessed in https://github.com/ruixiangcui/AGIEval/commit/5c77d073fda993f1652eaae3cf5d04cc5fd21d40 . Citation: @misc{zhong2023agieval, title={AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models}… See the full description on the dataset page: https://huggingface.co/datasets/hails/agieval-gaokao-english.textn<1K1 likes3.7k downloads3y agoHugging Face11parler-tts /mls_eng Dataset Card for English MLS Dataset Summary This is a streamable version of the English version of the Multilingual LibriSpeech (MLS) dataset. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls_eng.audioautomatic-speech-recognition10M<n<100M40 likes3.5k downloads2y agoHugging Face12CohereLabs /msmarco-v2.1-embed-english-v3 TREC-RAG 2024 Corpus (MSMARCO 2.1) - Encoded with Cohere Embed English v3 This dataset contains the embeddings for the TREC-RAG Corpus 2024 embedded with the Cohere Embed V3 English model. It contains embeddings for 113,520,750 passages, embeddings for 1677 queries from TREC-Deep Learning 2021-2023, as well as top-1000 hits for all queries using a brute-force (flat) index. Search over the Index We have a pre-build index that only requires 300 MB available at… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/msmarco-v2.1-embed-english-v3.tabular100M<n<1B7 likes3.4k downloads6mo agoHugging Face13shareAI /ShareGPT-Chinese-English-90k ShareGPT-Chinese-English-90k Bilingual Human-Machine QA Dataset A high-quality Chinese-English parallel bilingual human-machine QA dataset, covering user questions in real and complex scenarios. It is used for training high-quality dialogue models (more robust in instruction distribution than those datasets generated by repeatedly calling API interfaces to simulate machine-generated Q&A, like Moss) Features: Provides fully semantically equivalent Chinese-English parallel corpus… See the full description on the dataset page: https://huggingface.co/datasets/shareAI/ShareGPT-Chinese-English-90k.question-answering10K<n<100K288 likes3.1k downloads9mo agoHugging Face14Abirate /english_quotes Dataset Card for English quotes I-Dataset Summary english_quotes is a dataset of all the quotes retrieved from goodreads quotes. This dataset can be used for multi-label text classification and text generation. The content of each quote is in English and concerns the domain of datasets for NLP and beyond. II-Supported Tasks and Leaderboards Multi-label text classification : The dataset can be used to train a model for text-classification, which consists of… See the full description on the dataset page: https://huggingface.co/datasets/Abirate/english_quotes.texttext-classification1K<n<10K109 likes2.8k downloads4y agoHugging Face15nreimers /sphere_cohere_embed-english-v3.0text1M<n<10M0 likes2.7k downloads3y agoHugging Face16aixk /fastplus-125m-dataset-eng-6 ISAI - 이사이 I’m an independent developer building and maintaining AI projects on my own. Everything from model development to server costs, datasets, and feature updates is managed personally. Any support you can provide greatly helps keep this project running and allows for continuous improvements. If you find this project helpful, please consider supporting my work. Thank you. 혼자서 AI 프로젝트를 개발하고 운영하고 있습니다. 모델 개발부터 데이터셋 준비, 서버 비용 감당, 기능 업데이트까지 모두 직접 진행하고 있습니다. 보내주시는 따뜻한 후원은 안정적인… See the full description on the dataset page: https://huggingface.co/datasets/aixk/fastplus-125m-dataset-eng-6.100K<n<1M0 likes2.6k downloads3mo agoHugging Face17Team-PIXEL /rendered-wikipedia-english Dataset Card for Team-PIXEL/rendered-wikipedia-english Dataset Summary This dataset contains the full English Wikipedia from February 1, 2018, rendered into images of 16x8464 resolution. The original text dataset was built from a Wikipedia dump. Each example in the original text dataset contained the content of one full Wikipedia article with cleaning to strip markdown and unwanted sections (references, etc.). Each rendered example contains a subset of one full article.… See the full description on the dataset page: https://huggingface.co/datasets/Team-PIXEL/rendered-wikipedia-english.text10M<n<100M4 likes2.6k downloads4y agoHugging Face18dunnolab /so-combined-engThis dataset was created using LeRobot. Dataset Description The English version of this dataset integrates 598 open-source community datasets into a single unified corpus, comprising 22,709 episodes and approximately 9.4 million frames across 563 distinct tasks. Several transformations were applied to ensure standardization and data quality: Camera view normalizationBecause community datasets do not follow a consistent naming scheme for camera viewpoints, we used the… See the full description on the dataset page: https://huggingface.co/datasets/dunnolab/so-combined-eng.tabularrobotics1M<n<10M6 likes2.5k downloads10mo agoHugging Face19ghanaopenai /ghana-english-asr-2700hrs This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. 🇬🇭 Ghana English ASR Dataset A speech dataset of Ghanaian English extracted from Ghanaian news media broadcasts, designed for training and fine-tuning Automatic Speech Recognition (ASR) models on West African English accents.… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-english-asr-2700hrs.audioautomatic-speech-recognition100K<n<1M7 likes2.2k downloads3mo agoHugging Face20CohereLabs /beir-embed-english-v3 BEIR embeddings with Cohere embed-english-v3.0 model This datasets contains all query & document embeddings for BEIR, embedded with the Cohere embed-english-v3.0 embedding model. Overview of datasets This repository hosts all 18 datasets from BEIR, including query and document embeddings. The following table gives an overview of the available datasets. See the next section how to load the individual datasets. Dataset nDCG@10 #Documents arguana 53.98 8,674… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/beir-embed-english-v3.text10M<n<100M9 likes2.2k downloads6mo agoHugging Face21amitness /logits-english-512 Dataset Card for "logits-english-512" More Information needed 1M<n<10M0 likes2.1k downloads3y agoHugging Face22SITL-Eng /CRCD Comprehensive Robotic Cholecystectomy Dataset (CRCD) The Comprehensive Robotic Cholecystectomy Dataset (CRCD) is a large-scale, multimodal dataset for robot-assisted surgery (RAS) research.It provides synchronized endoscopic videos, da Vinci surgical robot kinematics, and pedal usage signals, making it one of the most comprehensive open datasets for studying robotic cholecystectomy procedures. CRCD supports research in: Medical robotics and surgical automation Computer vision… See the full description on the dataset page: https://huggingface.co/datasets/SITL-Eng/CRCD.imagerobotics100K<n<1M4 likes2k downloads10mo agoHugging Face23neuphonic /emilia-yodas-english-neucodecgated Dataset Card for NeuCodec Emilia-YODAS Dataset Summary The NeuCodec Emilia-YODAS dataset is an English-language dataset containing >30M audio samples (>78k hours), taken from the English-language subset of Emilia-YODAS and compressed with NeuCodec. Usage import torch from datasets import load_dataset from neucodec import NeuCodec # load dataset and model dataset = load_dataset("neuphonic/emilia-yodas-english-neucodec", split="train"… See the full description on the dataset page: https://huggingface.co/datasets/neuphonic/emilia-yodas-english-neucodec.tabular10M<n<100M17 likes1.9k downloads1y agoHugging Face24QuangDuy /FineWiki-eng-mds0 likes1.9k downloads10mo agoHugging Face25InsightHub /refinedweb-embed-english-v3.00 likes1.9k downloads2y agoHugging Face26ylacombe /english_dialects Dataset Card for "english_dialects" Dataset Summary This dataset consists of 31 hours of transcribed high-quality audio of English sentences recorded by 120 volunteers speaking with different accents of the British Isles. The dataset is intended for linguistic analysis as well as use for speech technologies. The speakers self-identified as native speakers of Southern England, Midlands, Northern England, Welsh, Scottish and Irish varieties of English. The recording scripts… See the full description on the dataset page: https://huggingface.co/datasets/ylacombe/english_dialects.audiotext-to-speech10K<n<100K37 likes1.8k downloads3y agoHugging Face27ryanjosephkamp /english-openlist English OpenList The largest open-source, validated English word list for NLP and games. Dataset Description English OpenList is a comprehensive, continuously updated dictionary of valid English words. It provides: ~345,000 validated English words, plus a candidate pool of ~9.3 million awaiting evidence Validation provenance for every word: which sources attested it, and when Daily updates from authoritative dictionary sources Version history with changelogs for… See the full description on the dataset page: https://huggingface.co/datasets/ryanjosephkamp/english-openlist.texttext-classification100K<n<1M2 likes1.8k downloads3d agoHugging Face28NLPC-UOM /sentence_alignment_dataset-Sinhala-Tamil-English Dataset summary This is a gold-standard benchmark dataset for sentence alignment, between Sinhala-English-Tamil languages. Data had been crawled from the following news websites. The aligned documents annotated in the dataset NLPC-UOM/document_alignment_dataset-Sinhala-Tamil-English had been considered to annotate the aligned sentences. News Source url Army https://www.army.lk/ Hiru http://www.hirunews.lk ITN https://www.newsfirst.lk Newsfirst https://www.itnnews.lk… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/sentence_alignment_dataset-Sinhala-Tamil-English.sentence-similarity3 likes1.7k downloads3y agoHugging Face29mteb /cqadupstack-english CQADupstackEnglishRetrieval An MTEB dataset Massive Text Embedding Benchmark CQADupStack: A Benchmark Data Set for Community Question-Answering Research Task category t2t Domains Written Reference http://nlp.cis.unimelb.edu.au/resources/cqadupstack/ How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["CQADupstackEnglishRetrieval"]) evaluator = mteb.MTEB(task)… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cqadupstack-english.texttext-retrieval10K<n<100K1 likes1.6k downloads1y agoHugging Face30ghanaopenai /ghana-english-speech-600hrs This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. 🇬🇭 Ghana English ASR Dataset A speech dataset of Ghanaian English extracted from Ghanaian news media broadcasts, designed for training and fine-tuning Automatic Speech Recognition (ASR) models on West African English accents.… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-english-speech-600hrs.audioautomatic-speech-recognition100K<n<1M1 likes1.6k downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.