datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AI-CUDA-Engineer-Archive
The AI CUDA Engineer Archive 👷: Agentic CUDA Kernel Discovery, Optimization & Composition
We release The AI CUDA Engineer archive, a dataset consisting of approximately 30,000 CUDA kernels generated by The AI CUDA Engineer. It is released under the CC-By-4.0 license and can be accessed via HuggingFace and interactively visualized here. The dataset is based on the Kernel tasks provided in KernelBench and includes a torch reference implementation, torch, NCU and Clang-tidy… See the full description on the dataset page: https://huggingface.co/datasets/SakanaAI/AI-CUDA-Engineer-Archive.cad-environments
CAD Environments
CAD Environments is a multimodal dataset of complete, human-performed workflows in desktop CAD software. The current release contains 51 task workflows totaling 99.03 hours, covering eight software groups across mechanical design, architecture, MEP, structural design, and general 3D modeling.
Each workflow preserves the full task context—not just the final model—including the problem statement, reference and input files, a gold output, evaluation rubrics, a… See the full description on the dataset page: https://huggingface.co/datasets/markov-ai/cad-environments.smollm-chunked
FAISS Indices and Chunked Datasets for SmolLM and SmolLM2 corpora
This repository contains part of the FAISS indices and chunked datasets used for novelty detection for SmolLM and SmolLM2, as presented in the paper LLM generation novelty through the lens of semantic similarity.
Full Documentation
For complete usage instructions, installation guide, and tutorial, please refer to:
Main Tutorial README
Data Distribution
Due to Hugging Face storage quota… See the full description on the dataset page: https://huggingface.co/datasets/enguyen/smollm-chunked.seamless-align-enA-viA.speaker-embedding.xlsr-2bseamless-align-enA-esA.speaker-embedding.w2vbert-600mseamless-align-enA-frA.speaker-embedding.hubert-xlseamless-align-enA-jaA.speaker-embedding.w2vbert-600mEnglish-Zomi-OPUS_Tatoeba_v20230412
English–Zomi Parallel Corpus (1.78M)
This dataset contains 1.78 million English–Zomi sentence pairs, created to support
machine translation, linguistic research, and large‑scale language model training.
It is fully open and permissively licensed for commercial and non‑commercial use.
🌐 Linguistic Background: Zomi, Tedim Chin, and ISO Codes
Zomi is the endonym (self‑chosen name) of the people and their language.However, Zomi does not yet have an official ISO 639‑3 code.… See the full description on the dataset page: https://huggingface.co/datasets/ZomiLearner/English-Zomi-OPUS_Tatoeba_v20230412.seamless-align-deA-enA.speaker-embedding.xlsr-2bseamless-align-enA-hiA.speaker-embedding.hubert-xlseamless-align-enA-frA.speaker-embedding.xlsr-2bseamless-align-enA-esA.speaker-embedding.xlsr-2bseamless-align-enA-zhA.speaker-embedding.w2vbert-600mseamless-align-enA-frA.speaker-embedding.w2vbert-600mseamless-align-enA-zhA.speaker-embedding.xlsr-2bseamless-align-enA-zhA.speaker-embedding.hubert-xlseamless-align-enA-koA.speaker-embedding.w2vbert-600mseamless-align-enA-hiA.speaker-embedding.w2vbert-600mseamless-align-enA-viA.speaker-embedding.w2vbert-600mseamless-align-enA-jaA.speaker-embedding.hubert-xlseamless-align-enA-hiA.speaker-embedding.xlsr-2bseamless-align-enA-jaA.speaker-embedding.xlsr-2benron_spamThis is a version of the Enron Spam Email Dataset, containing emails (subject + message) and a label whether it is spam or ham.
awesome-loop-engineering
Awesome Loop Engineering Dataset
A structured dataset of 1022 papers, official docs, tools, benchmarks, patterns, critiques, and implementation guides for recurring AI-agent systems.
Resource Atlas ·
GitHub field guide ·
Resource selection ·
Report a correction
Dataset Summary
Each row connects an original source to its contribution, novelty, impact, publication details, lifecycle stages, audience, evidence type, link status, and… See the full description on the dataset page: https://huggingface.co/datasets/cy0307/awesome-loop-engineering.seamless-align-enA-koA.speaker-embedding.hubert-xlmsmarco-v2.1-embed-english-v3
TREC-RAG 2024 Corpus (MSMARCO 2.1) - Encoded with Cohere Embed English v3
This dataset contains the embeddings for the TREC-RAG Corpus 2024 embedded with the Cohere Embed V3 English model.
It contains embeddings for 113,520,750 passages, embeddings for 1677 queries from TREC-Deep Learning 2021-2023, as well as top-1000 hits for all queries using a brute-force (flat) index.
Search over the Index
We have a pre-build index that only requires 300 MB available at… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/msmarco-v2.1-embed-english-v3.seamless-align-deA-enA.speaker-embedding.w2vbert-600mseamless-align-enA-koA.speaker-embedding.xlsr-2bseamless-align-enA-esA.speaker-embedding.hubert-xlc4-en-html-with-metadata
