datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
spain-reference-personas-frontier
Spain Reference Personas Frontier
Spain Reference Personas Frontier is an open synthetic reference population and benchmark substrate for evaluating and designing socially grounded AI systems for Spain.
It is not observed microdata, not a survey, not a prediction of real citizens, and not a substitute for fieldwork, administrative data, or domain-specific validation.
The package is designed for simulation, evaluation, prompt conditioning, subgroup analysis, service design… See the full description on the dataset page: https://huggingface.co/datasets/apol/spain-reference-personas-frontier.ai-glossary-reference
AI Glossary
10,200 unique artificial intelligence and machine learning terms with concise definitions, a category, a difficulty level, and links to related terms. 149,760 words of definitions across 38 categories.
Usage
from datasets import load_dataset
ds = load_dataset("whoashish115/ai-glossary-reference", split="train")
The same rows are also provided as ai_glossary.jsonl and ai_glossary.csv (related terms joined with ; ).
Fields
Field… See the full description on the dataset page: https://huggingface.co/datasets/whoashish115/ai-glossary-reference.DOD-Enterprise-DevSecOps-Reference-Design-AWS-Managed-Services
DoD Enterprise DevSecOps AWS Managed Services Reference Design Question-Answer Dataset
Maintainer: Terry Eppler
Owner: US Federal Government
Dataset Summary
This dataset contains document-grounded question-and-answer records based on DoD Enterprise DevSecOps Reference Design: AWS Managed Services (DoD IaC Baseline), Version 0.2, September 2021.
The source presents a draft Department of Defense reference design for implementing a DevSecOps software factory… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DOD-Enterprise-DevSecOps-Reference-Design-AWS-Managed-Services.task401_numeric_fused_head_reference
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task401_numeric_fused_head_reference
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task401_numeric_fused_head_reference.MLIR-Functional-Reference-30
MLIR-Functional-Reference-30
Hand-authored functional-correctness reference set for arith, linalg+memref, and stablehlo (n=30, 10 per dialect).
Composition
Instances: 30
Format: one JSON record per line in data/test.jsonl
Schema: fields = canonical_fn_name, canonical_signature, dialect, expected_output, expected_output_pattern, expected_stdout_regex, id, inputs, iree_inputs, memref_inputs, memref_print, nl, result_type, scalar_inputs, source_benchmark, source_id… See the full description on the dataset page: https://huggingface.co/datasets/plawanrath/MLIR-Functional-Reference-30.bash-reference-manual-general-QAs
Dataset generated from bash reference manual.
book information like date and bash version are available within the very first rows of the dataset
this dataset is pretty small in general, but covering almost all of the definition and technical terms, commands and flags in the book
columns : "Question", "Answer"
persian-llm-reference
Persian LLM Reference — manifest snapshot
Bilingual, receipt-gated registry of Persian (Farsi) language models, datasets, benchmarks, and leaderboards.
This Hub dataset mirrors a release snapshot. It is not a competing source of truth.
Canonical surfaces (use these)
Role
Surface
Canonical reference (machine)
API v1 /api/v1/reference.json
Human interface
PLR Atlas
GitHub SSOT
manifest on main
Versioned release
GitHub Releases
HF Dataset (mirror)… See the full description on the dataset page: https://huggingface.co/datasets/Noetfield/persian-llm-reference.
