datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kimi-k3-full-mxfp4-kld-reference-32x2048
Kimi K3 full-MXFP4 KLD reference logits
This dataset contains the canonical full-vocabulary reference logits for
quantization comparisons of Kimi K3. The source is the original full MXFP4
checkpoint served as W4A16 on TP16 with vLLM dev/gg-k3, SparkInfer, and
InstantTensor.
Contents
32 independent 2048-token windows
65,504 scored next-token positions (32 * 2047)
vocabulary size 163,840
one [2047, 163840] F32 safetensors tensor per window
tensor key: logits
total… See the full description on the dataset page: https://huggingface.co/datasets/festr2/kimi-k3-full-mxfp4-kld-reference-32x2048.spain-reference-personas-frontier
Spain Reference Personas Frontier
Spain Reference Personas Frontier is an open synthetic reference population and benchmark substrate for evaluating and designing socially grounded AI systems for Spain.
It is not observed microdata, not a survey, not a prediction of real citizens, and not a substitute for fieldwork, administrative data, or domain-specific validation.
The package is designed for simulation, evaluation, prompt conditioning, subgroup analysis, service design… See the full description on the dataset page: https://huggingface.co/datasets/apol/spain-reference-personas-frontier.ai-glossary-reference
AI Glossary
10,200 unique artificial intelligence and machine learning terms with concise definitions, a category, a difficulty level, and links to related terms. 149,760 words of definitions across 38 categories.
Usage
from datasets import load_dataset
ds = load_dataset("whoashish115/ai-glossary-reference", split="train")
The same rows are also provided as ai_glossary.jsonl and ai_glossary.csv (related terms joined with ; ).
Fields
Field… See the full description on the dataset page: https://huggingface.co/datasets/whoashish115/ai-glossary-reference.DOD-Enterprise-DevSecOps-Reference-Design-AWS-Managed-Services
DoD Enterprise DevSecOps AWS Managed Services Reference Design Question-Answer Dataset
Maintainer: Terry Eppler
Owner: US Federal Government
Dataset Summary
This dataset contains document-grounded question-and-answer records based on DoD Enterprise DevSecOps Reference Design: AWS Managed Services (DoD IaC Baseline), Version 0.2, September 2021.
The source presents a draft Department of Defense reference design for implementing a DevSecOps software factory… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DOD-Enterprise-DevSecOps-Reference-Design-AWS-Managed-Services.task401_numeric_fused_head_reference
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task401_numeric_fused_head_reference
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task401_numeric_fused_head_reference.MLIR-Functional-Reference-30
MLIR-Functional-Reference-30
Hand-authored functional-correctness reference set for arith, linalg+memref, and stablehlo (n=30, 10 per dialect).
Composition
Instances: 30
Format: one JSON record per line in data/test.jsonl
Schema: fields = canonical_fn_name, canonical_signature, dialect, expected_output, expected_output_pattern, expected_stdout_regex, id, inputs, iree_inputs, memref_inputs, memref_print, nl, result_type, scalar_inputs, source_benchmark, source_id… See the full description on the dataset page: https://huggingface.co/datasets/plawanrath/MLIR-Functional-Reference-30.bash-reference-manual-general-QAs
Dataset generated from bash reference manual.
book information like date and bash version are available within the very first rows of the dataset
this dataset is pretty small in general, but covering almost all of the definition and technical terms, commands and flags in the book
columns : "Question", "Answer"
persian-llm-reference
Persian LLM Reference — manifest snapshot
Bilingual, receipt-gated registry of Persian (Farsi) language models, datasets, benchmarks, and leaderboards.
This Hub dataset mirrors a release snapshot. It is not a competing source of truth.
Canonical surfaces (use these)
Role
Surface
Canonical reference (machine)
API v1 /api/v1/reference.json
Human interface
PLR Atlas
GitHub SSOT
manifest on main
Versioned release
GitHub Releases
HF Dataset (mirror)… See the full description on the dataset page: https://huggingface.co/datasets/Noetfield/persian-llm-reference.
