datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
awesome-loop-engineering
Awesome Loop Engineering Dataset
A structured dataset of 1022 papers, official docs, tools, benchmarks, patterns, critiques, and implementation guides for recurring AI-agent systems.
Resource Atlas ·
GitHub field guide ·
Resource selection ·
Report a correction
Dataset Summary
Each row connects an original source to its contribution, novelty, impact, publication details, lifecycle stages, audience, evidence type, link status, and… See the full description on the dataset page: https://huggingface.co/datasets/cy0307/awesome-loop-engineering.Awesome_Spatial_VQA_Benchmarksawesome-egocentric-atlas
Use this dataset
from datasets import load_dataset
ds = load_dataset("cy0307/awesome-egocentric-atlas", split="train")
print(len(ds), "resources")
print(ds[0])
papers = load_dataset(
"csv",
data_files="https://huggingface.co/datasets/cy0307/awesome-egocentric-atlas/resolve/main/awesome-egocentric-papers.csv",
split="train",
)
print(len(papers), "paper-linked resources")
Each row is one catalogued resource. Columns:
Column
Description
name
Resource name… See the full description on the dataset page: https://huggingface.co/datasets/cy0307/awesome-egocentric-atlas.awesome-markdown-ebooks
Awesome-markdown-ebooks
Your GitHub PDFs, Now AI-Ready.
Project repo: https://github.com/OpenDataLab/awesome-markdown-ebooks
Awesome_Spatial_VQA_Benchmarks_ViewSpatial-Benchawesome-llm-datasets-only-Chineseawesome-egocentric-atlas
Use this dataset
from datasets import load_dataset
ds = load_dataset("cy0307/awesome-egocentric-atlas", split="train")
print(len(ds), "resources")
print(ds[0])
papers = load_dataset(
"csv",
data_files="https://huggingface.co/datasets/cy0307/awesome-egocentric-atlas/resolve/main/awesome-egocentric-papers.csv",
split="train",
)
print(len(papers), "paper-linked resources")
Each row is one catalogued resource. Columns:
Column
Description
name
Resource name… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz042/awesome-egocentric-atlas.awesome-japanese-corpus
Awesome Japanese Corpus
4つの日本語データソースを text、from、from_license の3列に正規化した
Parquet データセットです。
FineWeb以外の3ソースは、生成処理を停止した時点までに取得済みの公開データを
採用しています。FineWebは最大3シャードを並列先読みする1時間限定の処理で
サンプリングしています。空文字は除外しています。Infini-News は取得済みの
年度について language_iso639_3 == "jpn" の行を採用しています。
Sources
hotchpotch/fineweb-2-edu-japanese (odc-by)
ruggsea/infini-news-corpus の language_iso639_3 == "jpn" (cc-by-4.0)
turing-motors/MOMIJI (cc-by-4.0)
AhmedSSabir/Japanese-wiki-dump-sentence-dataset… See the full description on the dataset page: https://huggingface.co/datasets/nakasyou/awesome-japanese-corpus.multi_news_parquet
Multi-News (Parquet Version)
This is a cleaned and updated version of the Multi-News dataset,converted to Parquet format for compatibility with the latest versionsof the Hugging Face datasets library.
Usage
from datasets import load_dataset
dataset = load_dataset("Awesome075/multi_news_parquet")
print(dataset)
Acknowledgments
This dataset is based on the original Multi-News datasetcreated by Alexander Fabbri et al. (2019).
The current version is a Parquet… See the full description on the dataset page: https://huggingface.co/datasets/Awesome075/multi_news_parquet.Rustins_Super_Mega_Awesome_VEDU_Model
Rustin's Super Mega Awesome VEDU Model
A reproducible, heavily-documented pipeline that maps Ventenata dubia ("VEDU", an invasive
winter-annual grass) across Montana from satellite + environmental data.
Science reference: docs/VEDU_48_predictors_detailed.md
Data decisions & gotchas: docs/CONTRADICTIONS.md
Parity with the Earth Engine build: docs/GEE_PARITY.md
Continue-the-build guide: docs/HANDOFF.md
Label inventory: docs/DATA_SOURCES.md
What it produces
57… See the full description on the dataset page: https://huggingface.co/datasets/UniversityOfMontanaSAL/Rustins_Super_Mega_Awesome_VEDU_Model.GraphRAG-BenchThis is the official dataset for GraphRAG-Bench: Challenging Domain-Specific Reasoning for Evaluating Graph Retrieval-Augmented Generation.
It contains 5 question types spanning 16 disciplines and a corpus of 7 million words from 20 computer science textbooks.
license
This dataset can only be used for academic research and cannot be used for any commercial purposes. It is prohibited to distribute or modify the content of the dataset. Any legal consequences resulting from violating… See the full description on the dataset page: https://huggingface.co/datasets/Awesome-GraphRAG/GraphRAG-Bench.awesome-ai-agent-papers
Hand-picked research papers on the AI agent ecosystem, published in 2026.
More awesome collections for developers
Awesome AI Agent Papers
A curated collection of research papers published in 2026 and sourced from arXiv, covering core topics from the AI agent ecosystem like multi-agent coordination, memory & RAG, tooling, evaluation & observability, and security.
Whether you're an AI engineer building agent systems, a… See the full description on the dataset page: https://huggingface.co/datasets/molmohsen/awesome-ai-agent-papers.awesome-ai-index
Awesome AI Index
Machine-readable JSON. No paywalls. Updated daily via GitHub Actions.
Curated catalog of AI tools, models, papers, frameworks, and resources for engineers and researchers.
Contents
What's Inside
Why This Exists
Quick Start
Top Open Models
Top Proprietary Models
Agent Frameworks
RAG Frameworks & Tools
Fine-Tuning Tools
Inference Optimization
Vector Databases
LLM Orchestration
Prompt Engineering
AI Code Assistants
AI Image Generation
AI Video… See the full description on the dataset page: https://huggingface.co/datasets/alpha-one-index/awesome-ai-index.awesome-degraded-segmentation
A Survey on Degraded Image Segmentation
A comprehensive survey on robust image segmentation under various degradation conditions
Paper | Paper List | Project Page
Abstract
Segmentation is the core of visual understanding — the foundation of Physical AI and World Model.
Image segmentation is a fundamental task in computer vision with wide-ranging applications. While deep learning models have achieved remarkable success under ideal conditions… See the full description on the dataset page: https://huggingface.co/datasets/Linwei-Chen/awesome-degraded-segmentation.awesome-listawesome-tagesschau
Awesome Tagesschau Dataset
This repository presents the new Awesome Tagesschau dataset. The dataset contains various news articles from the German Tagesschau website:
tagesschau.de is the central news portal of ARD (ARD is a joint organisation of Germany's regional public-service broadcasters).
The website provides a round-the-clock overview of the current news situation, supplemented by background information.
The editorial team relies on the extensive network of ARD… See the full description on the dataset page: https://huggingface.co/datasets/stefan-it/awesome-tagesschau.awesome_spatial_vl_msSSP-DETR-PCB
SSP-DETR PCB Dataset
This repository contains the PCB defect-detection dataset used for the SSP-DETR experiments.
Dataset Summary
Split
Images
Annotations
Train
394
1,668
Test
99
411
The six categories are missing_hole, mouse_bite, open_circuit, short, spur, and spurious_copper.
Layout
.
|-- annotations/
| |-- instances_train.json
| `-- instances_test.json
`-- images/
|-- train/
`-- test/
Annotations follow the COCO… See the full description on the dataset page: https://huggingface.co/datasets/awesome-aning/SSP-DETR-PCB.awesome-text2video-prompts
Rapidata Video Generation Preference Dataset
If you get value from this dataset and would like to see more in the future, please consider liking it.
This dataset contains prompts for video generation for 14 different categories. They were collected with a combination of manual prompting and ChatGPT 4o. We provide one example sora video generation for each video.
Overview
Categories and Comments
Object Interactions Scenes: Basic scenes with… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/awesome-text2video-prompts.awesome-mllm-benchmarks-samples
Awesome MLLM Benchmarks – Sample Data
🌐 Interactive Dashboard ·
💻 GitHub
This dataset hosts the sample data (images, questions, answers, metadata) used by the Awesome MLLM Benchmarks interactive dashboard. It provides curated preview samples from 130+ multimodal LLM benchmarks across 20+ categories.
Overview
Stat
Count
Benchmarks with samples
125
Total subtasks
248
Total files (images + metadata)
~8,000
Categories
20+… See the full description on the dataset page: https://huggingface.co/datasets/lchen1019/awesome-mllm-benchmarks-samples.awesome-graph-engineering
Awesome Graph Engineering Resource Atlas
A versioned collection of research, standards, frameworks, protocols, reliability systems, evaluations, and critiques for graph-structured multi-agent systems and programmable AI-agent organizations.
This dataset mirrors Awesome Graph Engineering. The GitHub JSONL file is canonical; the Hub exposes the same records through Dataset Viewer, direct downloads, datasets, and pandas.
Working definition
Graph engineering is the… See the full description on the dataset page: https://huggingface.co/datasets/cy0307/awesome-graph-engineering.awesome-computational-primatology
Awesome Computational Primatology
Papers per year — this hand-curated list (left) vs. the raw "primate + ML" literature on OpenAlex (right).
This repository contains the corpus of projects at the intersection of deep learning and non-human primatology since around the time AlexNet was published (~2012). This repo is intended for papers that provide novel approaches or applications in computational primatology. We occasionally include datasets which contain both… See the full description on the dataset page: https://huggingface.co/datasets/fparodi/awesome-computational-primatology.llama-3.1-awesome-chatgpt-prompts
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/llama-3.1-awesome-chatgpt-prompts.awesome-dataset-sinhala
Mixed Sinhala Dataset (1M+ Rows) | මිශ්ර සිංහල දත්ත කට්ටලය
(Please find the English description below the Sinhala description)
🇬🇧 English
This is a comprehensive dataset containing over one million rows of Sinhala text data. It is highly suitable for training Artificial Intelligence (AI) models and conducting Natural Language Processing (NLP) research.
Dataset Details
Language: Sinhala (si)
Total Rows: 1,079,909
Format: Parquet (Optimized for Hugging… See the full description on the dataset page: https://huggingface.co/datasets/sh4lu-z/awesome-dataset-sinhala.awesome_hunyuanImage_prompts
Awesome HunyuanImage Prompts
Built with Hugging Face AI Sheets
Do you want to master HunyuanImage 3, one of the best open models for generating images? This resource is for you.
HunyuanImage 3's Prompt Handbook is a great resource for learning how to prompt the model. Unfortunately, it is in Chinese, so I have created this resource for the open community.
It contains:
All prompts in HunyuanImage 3's Prompt Handbook are organized by category.
Their translation into English… See the full description on the dataset page: https://huggingface.co/datasets/dvilasuero/awesome_hunyuanImage_prompts.awesome-japanese-nlp-multilabel-dataset
Dataset overview
This is a dataset for Japanese natural language processing with multi-label annotations of research field labels for GitHub repositories in the NLP domain.
Please refer to this paper for the specific method of constructing the dataset. It is written in Japanese.
Input and Output
Input: Information from GitHub repositories (description, README text, PDF text, screenshot images)
Output: Multi-label classification of NLP research fields
Problem Setting of the… See the full description on the dataset page: https://huggingface.co/datasets/taishi-i/awesome-japanese-nlp-multilabel-dataset.awesome-cybersecurity-datasets
🛡️ AUTHORITATIVE 2026 FORK & INTERACTIVE CATALOG
This dataset is an official mirror of the curated Awesome Cybersecurity Datasets list.
🌐 View the Official Interactive Searchable Catalog
💻 Contribute on the Official GitHub Repository
Awesome Cybersecurity Datasets
🛡️ AUTHORITATIVE 2026 FORK
The original Awesome-Cybersecurity-Datasets repository was abandoned in 2021 and left to rot.
This is the actively maintained, state-of-the-art continuation by SystemHelpdesk. We… See the full description on the dataset page: https://huggingface.co/datasets/Jordan123234/awesome-cybersecurity-datasets.awesome-python-apps
Dataset Card for "awesome-python-apps"
This contains .py files for the following repos taken from awesome-python-applications (on GitHub here)
abilian-sbe clone_repos.sh invesalius3 photonix sk1-wx
ambar CONTRIBUTING.md isso picard soundconverter
apatite CTFd kibitzrpi-hole soundgrain
ArchiveBox Cura KindleEar planet stargate… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/awesome-python-apps.awesome-chatgpt-prompts-clean
🧠 Awesome ChatGPT Prompts — Clean
The classic 2,112-prompt role-prompting library (fka/prompts.chat, CC0) — deduplicated, quality-filtered, auto-categorized, shipped as typed parquet — plus 6 hand-verified community prompts mined from Claude practitioner chat.
Priorities: Quality > Cleanliness > Signal
Clean derivative of fka/prompts.chat (2,124 rows). License unchanged: CC0-1.0 ✅ no restrictions.
🧹 Quality Pipeline
Step
Removed
Reason
Raw… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/awesome-chatgpt-prompts-clean.awesome-taiwan-knowledge
Awesome Taiwan Knowledge (ATK) Dataset
The Awesome Taiwan Knowledge (ATK) Dataset is a comprehensive collection of questions and answers designed to evaluate artificial intelligence models' understanding of Taiwan-specific information. This unique dataset addresses the growing need for culturally nuanced AI performance metrics, particularly for models claiming global competence.
Key Features:
Taiwan-Centric Content: Covers a wide range of topics uniquely relevant to… See the full description on the dataset page: https://huggingface.co/datasets/aigrant/awesome-taiwan-knowledge.
