datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
catalog
Mesh-LLM Catalog
This dataset is the Hugging Face-backed catalog for Mesh-LLM.
The runtime catalog entries live under entries/**/*.json. The Dataset Viewer
uses catalog_rows.jsonl, a flat generated table with one row per model variant.
The catalog deliberately excludes raw blob URLs. Entries should resolve to
Hugging Face repositories and canonical Mesh refs.
ine-catalog
INE
Este repositorio contiene todas las tablas¹ del Instituto Nacional de Estadística exportadas a ficheros Parquet.
Puedes encontrar cualquiera de las tablas o sus metadatos en la carpeta tablas.
Cada tabla está identificado un una ID. Puedes encontrar la ID de la tabla tanto en el INE (es el número que aparece en la URL) or en el archivo tablas.jsonl de este repositorio que puedes explorar en el Data Viewer.
Por ejemplo, la tabla de Índices nacionales de clases se corresponde… See the full description on the dataset page: https://huggingface.co/datasets/datania/ine-catalog.european-open-data-catalogue
European Open Data Catalogue
This repository publishes independently versioned metadata and licensed source snapshots:
A discovery catalogue with 15565 dataset entries from
ISTAT, Eurostat, OECD, ILO, DoveVannoINostriSoldi (DVNS) and Cruscotto Italia.
3 independently pinned availability indexes with
911,795 joint combinations across 35 datasets, built from complete
source responses within the explicitly declared scope.
Licensed Cruscotto source snapshots, stored separately from… See the full description on the dataset page: https://huggingface.co/datasets/Gramscii-IT/european-open-data-catalogue.news-category-datasetDataset from https://www.kaggle.com/datasets/rmisra/news-category-dataset
minif2f-lean4Fixing the errors in some formal statements and informal proofs of minif2f-lean4.
News_Category_Dataset_v2CategoricalHarmfulQA
CatQA: A categorical harmful questions dataset
CatQA is used in LLM safety realignment research:
Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic (Paper, Code)
How to download
from datasets import load_dataset
dataset = load_dataset("declare-lab/CategoricalHarmfulQA")
What is CatQA?
To comprehensively evaluate the model across a wide range of harmful categories, we construct a new safety… See the full description on the dataset page: https://huggingface.co/datasets/declare-lab/CategoricalHarmfulQA.student-question-categoriesThis is the IITJEE NEET AIIMS Students Questions Data dataset.
It categorizes university entry questions into 4 categories: Physics, Chemistry, Biology, and Mathematics.
ni-20-categoriesRFP_Categorizationliving-catalog
Living catalog
This is the directory of record. Every public Council of AI dataset, Space, model, API, and RAG pointer is a row. Rebuild by running publish_living_catalog.py (overnight keeper calls it).
Viewer Space: https://huggingface.co/spaces/csoai/living-catalog
Living board: GET https://councilof.ai/api/gspc (counts are derived there, never typed here as scores)
Verify: https://councilof.ai/gspc-verify (free forever)
What this is not
Not a certification… See the full description on the dataset page: https://huggingface.co/datasets/csoai/living-catalog.japanese-corpus-categorized
日本語コーパス
mc4-jaなどのwebコーパスをクリーニング後、教師なし学習モデルでテキストを約1万件にクラスタリングしたコーパスです。
著作権法で認められた情報解析目的で使用できます。
一部のファイルしかparquet化されていないので、ご注意ください。ファイルリストはoutフォルダ内にあります
git lfsなどでダウンロードください。
sharegpt-deduplicated
Dataset Card for Dataset Name
Dataset Description
Dataset Summary
This dataset is a deduplicated version of sharegpt4.
The deduplication process has two steps:
The literal duplicates (both input and outputs) are removed
The remaining (5749) instances are embedded with the SentenceTransformer library ("paraphrase-multilingual-mpnet-base-v2" model).
Then, we compute the cosine similarity among all the possible pairs, and consider paraphrases those pairs with a… See the full description on the dataset page: https://huggingface.co/datasets/CaterinaLac/sharegpt-deduplicated.LMOD-Cataract-1K-surgical-analysis-cot
Cataract-1K LLM-Generated Surgical Instructions
Dataset Overview
This dataset is derived from the Cataract-1K dataset (part of the LMOD benchmark) and enhanced using Qwen3-VL-30B-A3B-Thinking, a large vision-language model with reasoning capabilities. It is designed for training medical AI systems to provide actionable surgical guidance with transparent reasoning.
Generation Process
Source Data: Cataract-1K processed frames with segmentation annotations… See the full description on the dataset page: https://huggingface.co/datasets/mehti/LMOD-Cataract-1K-surgical-analysis-cot.MOUNT-Cattle
Updates/News 📣
🎉 News (Feb. 2026): The dataset paper FSMC-Pose has been accepted for CVPR 2026 Findings!
🔗 News: Please find the open-source dataset on Hugging Face: MOUNT-Cattle.
🔥 Downloads reached 2.4k within 7 days of release.
📌 Overview
Mounting posture is an important visual indicator of estrus in dairy cattle. MOUNT-Cattle is a mounting dataset, covering 1,176 mounting instances, which follows the COCO format… See the full description on the dataset page: https://huggingface.co/datasets/eelianafang/MOUNT-Cattle.games-catalog
Council of AI — games catalog
Catalog door. Games load into Council Space. No contest engine on this card.
Live: https://councilof.ai/gspc-arena
Council OS: https://councilof.ai/os
Council Space: https://councilof.ai/gspc-arena
Measurement, not certification. Empty slots are not for sale. No scores on this card.
Jail is a measured floor, not a 16th pane.
The live board is the authority
GET https://councilof.ai/api/gspc — quote totals.public_count. This Hub… See the full description on the dataset page: https://huggingface.co/datasets/csoai/games-catalog.reachy-mini-app-categoriesCC-Cat
CC_Cat
Extract from CC-WARC snapshots.
Mainly includes texts with 149 languages.
PDF/IMAGE/AUDIO/VIDEO raw downloading link.
Notice
Since my computing resources are limited, this dataset will update by one-day of CC snapshots timestampts.
After a snapshot is updated, the deduplicated version will be uploaded.
If you are interested in providing computing resources or have cooperation needs, please contact me.
carreyallthetime@gmail.com
CAT-Conversational-Dataset-IndicAmazon_Categoryplaycat-cat-behavior-new-data-set
PlayCat Cat Behavioral Enrichment Dataset
The definitive multilingual research dataset on cat behavioral enrichment by PlayCat Research
Dataset Summary
The PlayCat Cat Behavioral Enrichment Dataset is the largest open, bilingual (Korean-English) collection dedicated to feline environmental enrichment research. It contains 12,262 deduplicated entries spanning peer-reviewed academic papers, patents, veterinary Q&A, and community knowledge on cat behavior enrichment… See the full description on the dataset page: https://huggingface.co/datasets/playcat/playcat-cat-behavior-new-data-set.IndustryInstruction_Hospitality-Catering
IndustryInstruction: Hospitality Catering
This repository contains the IndustryInstruction: Hospitality Catering domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Hospitality-Catering.cat-v3xxl
🐱 cat-v3xxl (XXXL)
Part of the cat-v3 dataset family — synthetic instruction-tuning data that teaches models to be
helpful, accurate, and delightfully cat-flavoured.
About cat-v3xxl (XXXL)
At 1,075,000 rows (5.375× the XL variant), this is the large-scale training dataset for serious fine-tuning runs. The full topic bank is sampled densely, providing high repetition for core topics and meaningful coverage of rare ones.
Sharded into 250,000-row JSONL files for easy… See the full description on the dataset page: https://huggingface.co/datasets/Nix-ai/cat-v3xxl.cat-v3.6
Dataset: Nix-ai/cat-v3.6
This is a procedurally generated synthetic dataset, part of the cat-v3.6 family of datasets.
Dataset Statistics
Total Expected Rows: 817,089
Unique Topics: 273
Sets per Topic: 2,993
Detail Multiplier: 1.00x base
File Format: jsonl
Generation Rules applied to this tier:
Topics: Procedurally combined without using any character names.
Scaling System: Each version mathematically scales Topics by 1.375x, Details by 2.15x, and Sets by… See the full description on the dataset page: https://huggingface.co/datasets/Nix-ai/cat-v3.6.lm_code_github-eval_subsetSwissSPARK_Catalogs
⚠️ Caution: The dataset is subject to continuous changes. We are currently actively developing it.
Dataset Card: Auxiliary Dataset for a Swiss Sustainable Procurement Analysis & Reporting Kit
Dataset Description
This dataset contains a specific snapshot of the sustainability procurement criteria catalogs available at IntelliProcure/sustainability_criteria. This version was used to annotate Swiss calls for tender, which form the core of the… See the full description on the dataset page: https://huggingface.co/datasets/IntelliProcure/SwissSPARK_Catalogs.harbor-release-catalog
February 2025 Harbor Software Catalog
Approved stable releases published in February 2025.
Artifact
Version
Published
Downloads
Buoy Mapper
1.0.0
2025-02-26
5,175
Dock Ledger
3.2.1
2025-02-24
11,980
Harbor Status API
1.4.0
2025-02-14
18,420
Total releases: 3
Total downloads: 35,575
catena
Catena: an open, cited corpus of the Catholic tradition
A clean, machine-readable, public-domain corpus of the Catholic tradition where
every unit of text carries a canonical citation, the text is stored verbatim
(never paraphrased or truncated), and the tradition's own citation graph is
included as loadable data. Built for grounded, cite-or-refuse retrieval.
Source, ingest code, a live web explorer, and an MCP grounding server:
https://github.com/AlvaroBalbin/catena… See the full description on the dataset page: https://huggingface.co/datasets/TheAlvaroBalbin/catena.mantinc-catalan-drift
Mantinc — Catalan Drift Benchmark
Descripció (ca)
Mantinc és un banc de proves que avalua si un model de llenguatge continua
responent en català quan el missatge, la conversa prèvia o el context recuperat
l'empenyen a fer-ho en una altra llengua, normalment el castellà o l'anglès.
Dataset Description
Mantinc is a benchmark that measures whether a language model keeps answering
in Catalan when the prompt, prior conversation, or retrieved context… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/mantinc-catalan-drift.collaborative_catalog
