CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SakanaAI /AI-CUDA-Engineer-Archive The AI CUDA Engineer Archive 👷: Agentic CUDA Kernel Discovery, Optimization & Composition We release The AI CUDA Engineer archive, a dataset consisting of approximately 30,000 CUDA kernels generated by The AI CUDA Engineer. It is released under the CC-By-4.0 license and can be accessed via HuggingFace and interactively visualized here. The dataset is based on the Kernel tasks provided in KernelBench and includes a torch reference implementation, torch, NCU and Clang-tidy… See the full description on the dataset page: https://huggingface.co/datasets/SakanaAI/AI-CUDA-Engineer-Archive.tabular10K<n<100K227 likes136k downloads2y agoHugging Face02markov-ai /cad-environments CAD Environments CAD Environments is a multimodal dataset of complete, human-performed workflows in desktop CAD software. The current release contains 51 task workflows totaling 99.03 hours, covering eight software groups across mechanical design, architecture, MEP, structural design, and general 3D modeling. Each workflow preserves the full task context—not just the final model—including the problem statement, reference and input files, a gold output, evaluation rubrics, a… See the full description on the dataset page: https://huggingface.co/datasets/markov-ai/cad-environments.imagen<1K17 likes57k downloads2mo agoHugging Face03enguyen /smollm-chunked FAISS Indices and Chunked Datasets for SmolLM and SmolLM2 corpora This repository contains part of the FAISS indices and chunked datasets used for novelty detection for SmolLM and SmolLM2, as presented in the paper LLM generation novelty through the lens of semantic similarity. Full Documentation For complete usage instructions, installation guide, and tutorial, please refer to: Main Tutorial README Data Distribution Due to Hugging Face storage quota… See the full description on the dataset page: https://huggingface.co/datasets/enguyen/smollm-chunked.tabulartext-retrieval100M<n<1B1 likes35k downloads7mo agoHugging Face04asahi417 /seamless-align-enA-viA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes20k downloads2y agoHugging Face05asahi417 /seamless-align-enA-esA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes20k downloads2y agoHugging Face06asahi417 /seamless-align-enA-frA.speaker-embedding.hubert-xltabular1M<n<10M0 likes16k downloads2y agoHugging Face07asahi417 /seamless-align-enA-jaA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes15k downloads2y agoHugging Face08ZomiLearner /English-Zomi-OPUS_Tatoeba_v20230412 English–Zomi Parallel Corpus (1.78M) This dataset contains 1.78 million English–Zomi sentence pairs, created to support machine translation, linguistic research, and large‑scale language model training. It is fully open and permissively licensed for commercial and non‑commercial use. 🌐 Linguistic Background: Zomi, Tedim Chin, and ISO Codes Zomi is the endonym (self‑chosen name) of the people and their language.However, Zomi does not yet have an official ISO 639‑3 code.… See the full description on the dataset page: https://huggingface.co/datasets/ZomiLearner/English-Zomi-OPUS_Tatoeba_v20230412.tabulartranslation1M<n<10M0 likes15k downloads7mo agoHugging Face09asahi417 /seamless-align-deA-enA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes11k downloads2y agoHugging Face10asahi417 /seamless-align-enA-hiA.speaker-embedding.hubert-xltabular100K<n<1M0 likes10k downloads2y agoHugging Face11asahi417 /seamless-align-enA-frA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes9.9k downloads2y agoHugging Face12asahi417 /seamless-align-enA-esA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes9.9k downloads2y agoHugging Face13asahi417 /seamless-align-enA-zhA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes9.7k downloads2y agoHugging Face14asahi417 /seamless-align-enA-frA.speaker-embedding.w2vbert-600mtabular1M<n<10M0 likes8.9k downloads2y agoHugging Face15asahi417 /seamless-align-enA-zhA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes8.1k downloads2y agoHugging Face16asahi417 /seamless-align-enA-zhA.speaker-embedding.hubert-xltabular100K<n<1M0 likes7k downloads2y agoHugging Face17asahi417 /seamless-align-enA-koA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes6.9k downloads2y agoHugging Face18asahi417 /seamless-align-enA-hiA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes6.5k downloads2y agoHugging Face19asahi417 /seamless-align-enA-viA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes6.3k downloads2y agoHugging Face20asahi417 /seamless-align-enA-jaA.speaker-embedding.hubert-xltabular100K<n<1M0 likes5.9k downloads2y agoHugging Face21asahi417 /seamless-align-enA-hiA.speaker-embedding.xlsr-2btabular10K<n<100K0 likes5.7k downloads2y agoHugging Face22asahi417 /seamless-align-enA-jaA.speaker-embedding.xlsr-2btabular10K<n<100K0 likes5.6k downloads2y agoHugging Face23SetFit /enron_spamThis is a version of the Enron Spam Email Dataset, containing emails (subject + message) and a label whether it is spam or ham. tabular10K<n<100K21 likes5.5k downloads5y agoHugging Face24cy0307 /awesome-loop-engineering Awesome Loop Engineering Dataset A structured dataset of 1022 papers, official docs, tools, benchmarks, patterns, critiques, and implementation guides for recurring AI-agent systems. Resource Atlas · GitHub field guide · Resource selection · Report a correction &nbsp; Dataset Summary Each row connects an original source to its contribution, novelty, impact, publication details, lifecycle stages, audience, evidence type, link status, and… See the full description on the dataset page: https://huggingface.co/datasets/cy0307/awesome-loop-engineering.imagetext-classification1K<n<10K3 likes5.2k downloads2d agoHugging Face25asahi417 /seamless-align-enA-koA.speaker-embedding.hubert-xltabular100K<n<1M0 likes4k downloads2y agoHugging Face26CohereLabs /msmarco-v2.1-embed-english-v3 TREC-RAG 2024 Corpus (MSMARCO 2.1) - Encoded with Cohere Embed English v3 This dataset contains the embeddings for the TREC-RAG Corpus 2024 embedded with the Cohere Embed V3 English model. It contains embeddings for 113,520,750 passages, embeddings for 1677 queries from TREC-Deep Learning 2021-2023, as well as top-1000 hits for all queries using a brute-force (flat) index. Search over the Index We have a pre-build index that only requires 300 MB available at… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/msmarco-v2.1-embed-english-v3.tabular100M<n<1B7 likes3.4k downloads6mo agoHugging Face27asahi417 /seamless-align-deA-enA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes3.3k downloads2y agoHugging Face28asahi417 /seamless-align-enA-koA.speaker-embedding.xlsr-2btabular10K<n<100K0 likes3.2k downloads2y agoHugging Face29asahi417 /seamless-align-enA-esA.speaker-embedding.hubert-xltabular100K<n<1M0 likes3.1k downloads2y agoHugging Face30bs-modeling-metadata /c4-en-html-with-metadatatabular10M<n<100M14 likes2.8k downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.