datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ITBench-AA
ITBench-AA
Artificial Analysis' release of the public scenarios from
IBM's ITBench benchmark, used for
the ITBench-AA leaderboard.
This repo currently contains the SRE subset (sre config). Each row is a
Kubernetes incident scenario with its expected contributing-factor entities. An
agent under evaluation is given access to an offline snapshot of the affected
cluster (alerts, events, traces, topology) and must identify the entity
(Deployment, Pod, ConfigMap, etc.) responsible for… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/ITBench-AA.gspc-art5
GSPC — art5 safeguard bank (Art5Bench)
Council of AI measurement bank. Measurement, not certification.
Bank. Frozen split. Live n is the matching axis on GET https://councilof.ai/api/gspc, not a Hub score. Not a certificate. Art 50 (EUR-Lex): 2 August 2026 live; marking grace 2 December 2026.
Live measurement. This bank stands behind the art5-safeguard row of the live GSPC board: GET https://councilof.ai/api/gspc?axis=art5-safeguard (family, kind, status and n are on that row… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-art5.ArtiMuse-10K
ArtiMuse:
Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding
[🌐 Project Page]
[🚀 Online Demo]
[💻 Code]
[📄 Paper]
[[🧩 Checkpoints: 🤗 Hugging Face | 🤖 ModelScope]]
🌟 Building upon on ArtiMuse, we introduce UniPercept, a comprehensive follow-up work that provides a meticulous study on perceptual-level image understanding. It spans Image Aesthetics Assessment (IAA), Image Quality Assessment (IQA), and Image Structure & Texture… See the full description on the dataset page: https://huggingface.co/datasets/Thunderbolt215215/ArtiMuse-10K.IndustryInstruction_Artificial-Intelligence
IndustryInstruction: Artificial Intelligence
This repository contains the IndustryInstruction: Artificial Intelligence domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Artificial-Intelligence.constitution-2.0
Rules for agents that remember
If you are building multi-agent systems, you already have orchestration, tools, and
a memory store. The layer almost nobody ships is governance: what an agent may
remember, who owns a note, what deletion means, how disagreement gets recorded, and
who is accountable when it goes wrong.
This dataset is one complete answer to that, in the public domain. 47 articles,
one row each, byte-exact and hash-pinned. Written for work between humans and AI… See the full description on the dataset page: https://huggingface.co/datasets/article11/constitution-2.0.devcenter-articles
Overview
This dataset consists of ~600 articles from the MongoDB Developer Center.
Dataset Structure
The dataset consists of the following fields:
sourceName: The source of the article. This value is devcenter for the entire dataset.
url: Link to the article
action: Action taken on the article. This value is created for the entire dataset.
body: Content of the article in Markdown format
format: Format of the content. This value is md for all articles.
metadata: Metadata… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/devcenter-articles.ArtiMuse-10K
ArtiMuse:
Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding
[🌐 Project Page]
[🚀 Online Demo]
[💻 Code]
[📄 Paper]
[[🧩 Checkpoints: 🤗 Hugging Face | 🤖 ModelScope]]
🌟 Building upon on ArtiMuse, we introduce UniPercept, a comprehensive follow-up work that provides a meticulous study on perceptual-level image understanding. It spans Image Aesthetics Assessment (IAA), Image Quality Assessment (IQA), and Image Structure & Texture… See the full description on the dataset page: https://huggingface.co/datasets/to-hope-we-bound/ArtiMuse-10K.tw-processed-law-article
Dataset Card for tw-processed-law-article
tw-processed-law-article 是一個中華民國法規條文之結構化資料集,以條文為單位展開,包含 230,974 筆條文,涵蓋 11,462 部不同法規,橫跨憲法、法律與命令三個層級。每筆資料包含法規名稱、層級、條文內容、廢止註記與最後修正日期等欄位,適用於法律檢索系統、條文問答模型,或作為其他法律衍生資料集之結構化底層語料。
Dataset Details
Dataset Description
本資料集整理自中華民國全國法規資料庫(law.moj.gov.tw)之公開法規條文。原始法規資料經處理後以單條條文為一筆資料(one row per article),每筆附帶法規名稱、層級分類、條文全文、廢止註記、最後修正日期與 API 更新日期等元資料。
層級分佈:
層級
筆數
說明
命令
180,757
由主管機關訂定之命令、規則、辦法等
法律
49,977
經立法院三讀通過之法律
憲法… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-processed-law-article.LLM-MARS_QAThis dataset was developed by a team from Skoltech's Intelligent Space Robotics Laboratory.
The dataset was used to train LLM for Question Answering Module based on game results.
Note that this model is part of a multi-agent artificial intelligence system for the dog robot described in the LLM-MARS paper.
Paper preprint BibTeX cite:
@misc{lykov2023llmmarslargelanguagemodel,
title={LLM-MARS: Large Language Model for Behavior Tree Generation and NLP-enhanced Dialogue in Multi-Agent Robot… See the full description on the dataset page: https://huggingface.co/datasets/ArtemLykov/LLM-MARS_QA.art
🧠 Argument Reasoning Tasks (ART) Dataset
Evaluating natural language argumentative reasoning in large language models.
📖 Overview
The Argument Reasoning Tasks (ART) dataset is a large-scale benchmark designed to evaluate the ability of large language models (LLMs) to perform natural language argumentative reasoning.
It contains multiple-choice questions where models must identify missing argument components, given an argument context and reasoning structure.… See the full description on the dataset page: https://huggingface.co/datasets/debela-arg/art.phoronix-articles
Phoronix Articles Dataset: The Archive of Open-Source Computing Journalism
The definitive dataset of Phoronix - your gateway to years of open-source hardware/software evolution, performance analysis, and Linux ecosystem journalism.
🚀 What's Inside?
This dataset contains the complete archive of Phoronix articles - from bleeding-edge hardware launches to deep-dive Linux kernel analysis. Perfect for researchers, developers, and AI enthusiasts who need high-quality technical… See the full description on the dataset page: https://huggingface.co/datasets/hybridfree/phoronix-articles.artemisia-oab-constitucional
Artemísia — Dataset OAB 2ª Fase (Direito Constitucional)
Dataset de treinamento para a Artemísia, tutora de IA especializada na preparação para a 2ª fase da OAB em Direito Constitucional (banca FGV).
Sobre o Dataset
Campo
Valor
Total de exemplos
5,000
Modelo de embedding
paraphrase-multilingual-MiniLM-L12-v2 (384 dims)
Formato
JSONL (com embeddings) + CSV + JSONL leve (fine-tuning)
Idioma
Português Brasileiro
Domínio
Direito Constitucional — OAB… See the full description on the dataset page: https://huggingface.co/datasets/Arquimedes-msc/artemisia-oab-constitucional.Execution-Bound-Artifact-Reconstruction-Layer
🚩 Γ Physics Engine — Canonical Definition
Γ 物理引擎創建者 & 公式創始者:熊網區塊鏈 (BearNetworkChain) 創辦人 陳霆
最早提出時間:2025 年 6 月 19 日
原始來源:https://www.facebook.com/share/p/19cadcMTGo/
Chen, Ting. (2026). BearNetworkchain Execution Specification. Zenodo
📌 0. 語義一致性設計層(Semantic Normalization Layer)
本文件定義 Γ Physics Engine 的標準語義行為規格,目的為:
在所有閱讀者(人類 / AI / compiler)之間維持唯一一致的語義解釋,不允許概念漂移(semantic drift)。
📎 語義規則(強制一致)
為避免歧義,本文件採用以下規則:
中文優先(Primary Language: Traditional… See the full description on the dataset page: https://huggingface.co/datasets/BNES-BRNKC/Execution-Bound-Artifact-Reconstruction-Layer.brazilian-math-physics-qa
Brazilian Math & Physics QA
English | Português do Brasil
English
Summary
Brazilian Portuguese question-answer pairs covering mathematics, physics, chemistry, and related educational subjects. Each record contains a user question and an assistant answer in chat/SFT format.
Examples: 19.082
Train: 18.148
Validation: 934
Language: Brazilian Portuguese (pt-BR)
Schema
{"id":"qa_...","subject":"fisica","category":"mecanica-geral"… See the full description on the dataset page: https://huggingface.co/datasets/artificialguybr/brazilian-math-physics-qa.Execution-Bound-Artifact-Reconstruction-Layer
🚩 Γ Physics Engine — Canonical Definition
Γ 物理引擎創建者 & 公式創始者:熊網區塊鏈 (BearNetworkChain) 創辦人 陳霆
最早提出時間:2025 年 6 月 19 日
原始來源:https://www.facebook.com/share/p/19cadcMTGo/
Chen, Ting. (2026). BearNetworkchain Execution Specification. Zenodo
📌 0. 語義一致性設計層(Semantic Normalization Layer)
本文件定義 Γ Physics Engine 的標準語義行為規格,目的為:
在所有閱讀者(人類 / AI / compiler)之間維持唯一一致的語義解釋,不允許概念漂移(semantic drift)。
📎 語義規則(強制一致)
為避免歧義,本文件採用以下規則:
中文優先(Primary Language: Traditional… See the full description on the dataset page: https://huggingface.co/datasets/BearNetworkChain/Execution-Bound-Artifact-Reconstruction-Layer.brazilian-math-physics-qa-vision
Brazilian Math & Physics QA — Image Dependent
English | Português do Brasil
English
Summary
Brazilian Portuguese educational question-answer pairs whose problem statement or solution depends on one or more images.
Examples: 3,808
Referenced image URLs: 5,094 unique
Language: Brazilian Portuguese (pt-BR)
Schema
{"id":"vqa_...","subject":"matematica","category":"geometria","title":"...","messages":[{"role":"user","content":"...… See the full description on the dataset page: https://huggingface.co/datasets/artificialguybr/brazilian-math-physics-qa-vision.formatted_artofwarfinal1devcenter-articles-embedded
Overview
This dataset consists of chunked and embedded versions of a subset of articles from the MongoDB Developer Center.
Dataset Structure
The dataset consists of the following fields:
sourceName: The source of the article. This value is devcenter for the entire dataset.
url: Link to the article
action: Action taken on the article. This value is created for the entire dataset.
body: Content of the chunk in Markdown format
format: Format of the content. This value is… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/devcenter-articles-embedded.longchau_vaccine_articlesOver 600 vaccination related articles from nhathuoclongchau.com.vn
ArtiMuse-10K
ArtiMuse:
Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding
[🌐 Project Page]
[🚀 Online Demo]
[💻 Code]
[📄 Paper]
[[🧩 Checkpoints: 🤗 Hugging Face | 🤖 ModelScope]]
🌟 Building upon on ArtiMuse, we introduce UniPercept, a comprehensive follow-up work that provides a meticulous study on perceptual-level image understanding. It spans Image Aesthetics Assessment (IAA), Image Quality Assessment (IQA), and Image Structure & Texture… See the full description on the dataset page: https://huggingface.co/datasets/gltscc/ArtiMuse-10K.
