datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
quantum-like-attention-framework-1.3b-untuned-validation
Quantum Like Attention Framework (Q.L.A.F) 1.3b untuned
This repository contains the model checkpoints, downstream evaluation scores, and pretraining convergence logs for the Quantum Like Attention Framework (Q.L.A.F) 1.3B configuration.
Key Specifications & Architecture
Model Name: Q.L.A.F 1.3b untuned (Quantum Like Attention Framework - Hybrid Architecture)
Parameters: 1.3B parameters total configuration (327M active parameter student subset)
Layer Count: 12… See the full description on the dataset page: https://huggingface.co/datasets/IgnisCogitationis/quantum-like-attention-framework-1.3b-untuned-validation.gspc-regulatory-framework
GSPC — regulatory framework facts (RegimeFacts)
SWIFT census (live): https://councilof.ai/api/swift
XRPL reader (live): https://councilof.ai/api/xrpl
MEASURED financial/domain axis (declaration presence on retrieved URLs over live XRPL reader-16, n=16). Not a model leaderboard. No accuracy, no fleet, no leader.
Live status is the regulatory-framework row on GET https://councilof.ai/api/gspc. Not a certificate.
Tokenisation evidence question (24 September 2026): What can an… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-regulatory-framework.Real-UI-Clickboxes
RUC: Real UI Clickboxes
Click carefully, even when the page is trying to trick you! 👀
Official Hugging Face release for RUC: Real UI Clickboxes, the dataset accompanying our ACL 2026 paper Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces on deceptive UI understanding for web agents.
ACL Anthology: https://aclanthology.org/2026.acl-long.310/
PDF: https://aclanthology.org/2026.acl-long.310.pdf
DOI: https://doi.org/10.18653/v1/2026.acl-long.310… See the full description on the dataset page: https://huggingface.co/datasets/DUDE-Framework/Real-UI-Clickboxes.GEO-Framework
NobleJackal GEO Framework
A practical framework for making organisations clear, verifiable and citable in AI search
GEO means Generative Engine Optimization: the work of helping generative search and answer systems find, understand and support claims about organisations, people and content. This six-language book provides a seven-layer method for auditing entity clarity, evidence quality, machine-readable structure, question coverage, multilingual parity and… See the full description on the dataset page: https://huggingface.co/datasets/NobleJackal/GEO-Framework.PTB-XLCEP-IP_Framework
CEP-IP: An Explainable Framework for Cell Subpopulation Identification in Single-cell Transcriptomics (by Kah Keng Wong) (Published in Computer Methods and Programs in Biomedicine)
🧬 Abstract
Background and objective: Single-cell RNA sequencing (scRNA-seq) frameworks lack explainable approaches for identifying cell subpopulations harboring strong pairwise monotonic gene-module relationships between a gene of interest (GOI) and its co-expressed genes. In this study… See the full description on the dataset page: https://huggingface.co/datasets/kahkengwong/CEP-IP_Framework.Software-Architectural-FrameworksSoftware-Architectural-Frameworks
I am releasing a small dataset covering topics related to Frameworks under Software-Architecture.
I have included following topics:
TOGAF
Zachman Framework
IEEE 1471
Matrix-based approach to architecture development
Significance of IEEE 1471 (ISO/IEC 42010)
Benefits of employing architectural frameworks
and Many More!
This dataset can be useful in LLM development. Also those who are working on developing Software development related LLMs then this dataset can… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Software-Architectural-Frameworks.r9-research-framework
R9 Research Framework — Qwen3.5-9B Distillation
⚠️ CRITICAL: READ FIRST — Ollama Inference Flag Required
If you serve any Qwen3.5-derived model from this lineage via Ollama,
you MUST pass "think": false in the /api/chat request body.
curl -X POST http://localhost:11434/api/chat \
-d '{"model": "qwen3.5-9b-r10:q4km", "think": false, "messages": [...], "stream": false}'
Without this flag the model will appear to "loop" and produce empty answers
on 25-46% of requests.… See the full description on the dataset page: https://huggingface.co/datasets/cudabenchmarktest/r9-research-framework.UCI-HARethical-framework-UNESCO-Ethics-of-AI
Ethical AI Training Dataset
Introduction
UNESCO's Ethics of Artificial Intelligence, adopted by 193 Member States in November 2021, represents the first global framework for ethical AI development and deployment.
While regional initiatives like The Montréal Declaration for a Responsible Development of Artificial Intelligence emphasize community-driven governance, UNESCO's approach establishes comprehensive international standards through coordinated multi-stakeholder… See the full description on the dataset page: https://huggingface.co/datasets/ktiyab/ethical-framework-UNESCO-Ethics-of-AI.Mitre_Attacks_Framework_Dataset
MITRE ATT&CK Enterprise Dataset
Overview
This dataset provides a comprehensive collection of MITRE ATT&CK Enterprise techniques (v14.1) in JSONL format, designed for cybersecurity professionals, red teams, and threat hunters.
Each entry maps to a specific ATT&CK technique, including its ID, name, description, real-world example, and source.
The dataset is structured for seamless integration into security tools such as SIEMs, threat intelligence platforms, or custom red… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Mitre_Attacks_Framework_Dataset.AISA-AR-FunctionCall
AISA-AR-FunctionCall
Arabic Structured Function Calling Dataset
AISA-AR-FunctionCall is a large-scale Arabic dataset designed for training language models to convert natural language into structured executable tool calls.
The dataset enables research and development of Arabic agentic AI systems capable of invoking APIs, tools, and external services.
It is part of the AISA (Agentic AI Systems Architecture) initiative.
Dataset Overview
The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/AISA-Framework/AISA-AR-FunctionCall.IEMOCAPdataset_basel_frameworkaudiobook-listener-fit-framework
Audiobook Listener-Fit Framework
The Audiobook Listener-Fit Framework is an open structured taxonomy developed by Recommended Audiobooks for describing characteristics that influence the audiobook listening experience.
Traditional ratings mostly describe whether listeners liked a title. This framework is designed to describe how an audiobook listens and which types of listeners may be better suited to it.
Purpose
The framework organizes audiobook characteristics… See the full description on the dataset page: https://huggingface.co/datasets/recommendedaudiobooks/audiobook-listener-fit-framework.ethical-framework
1. Dataset Title
Ethical AI Decision-Making Training Data (Montreal Declaration Edition)
2. Overview
This dataset contains carefully crafted scenarios (instructions) and detailed responses illustrating step-by-step ethical reasoning aligned with the principles outlined in the Montreal Declaration for Responsible AI. Each entry poses a complex ethical challenge and provides a reasoned solution while referencing the specific principle(s) being tested.
These entries can… See the full description on the dataset page: https://huggingface.co/datasets/ktiyab/ethical-framework.basel-framework
Basel Framework
This dataset was generated using YourBench (v0.6.0), an open-source framework for generating domain-specific benchmarks from document collections.
Pipeline Steps
ingestion: Read raw source documents, convert them to normalized markdown and save for downstream steps
summarization: Perform hierarchical summarization: chunk-level LLM summaries followed by combine-stage reduction
chunking: Split texts into token-based single-hop and multi-hop chunks… See the full description on the dataset page: https://huggingface.co/datasets/LunaticMuch/basel-framework.NIST-CyberSecurity-Framework
# NIST Cybersecurity Framework 2.0 Question Answering Dataset
Dataset Summary
The NIST Cybersecurity Framework 2.0 Question Answering Dataset is a synthetic
instruction-style question-answering dataset derived from the NIST Cybersecurity
Framework (CSF) 2.0.
The dataset is designed to support training, fine-tuning, retrieval evaluation, and
domain-specific question-answering use cases related to cybersecurity risk management,
cybersecurity governance, enterprise risk management… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/NIST-CyberSecurity-Framework.Frameworker_User_Studyrepro-score-a-unified-framework-for-overshoot-refund-in-online-fdr-control-traces
Agent traces
Agent sessions published from a Trackio Logbook.
grounded-behavior-framework-v1_5
Grounded Behavior Framework N1 v1.5
Dataset sintético em português europeu para treino e avaliação de respostas
fundamentadas num contexto fornecido. Cada exemplo contém um contexto, uma
pergunta e uma resposta curta que aparece literalmente no contexto.
Como carregar
from datasets import load_dataset
dataset = load_dataset("empgces/grounded-behavior-framework-v1_5")
print(dataset)
print(dataset["train"][0])
Splits
Split
Exemplos
Utilização… See the full description on the dataset page: https://huggingface.co/datasets/empgces/grounded-behavior-framework-v1_5.ZIF-4_Amorphous_Zeolitic_Imidazolate_Frameworks_2023
Cite this dataset Castel, N., Andre, D., Edwards, C., Evans, J. D., and Coudert, F. ZIF-4 Amorphous Zeolitic Imidazolate Frameworks 2023. ColabFit, 2023. https://doi.org/10.60732/a6b0da5e
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_sh7jt3ptmde4_0
Visit the ColabFit Exchange to search additional datasets by author… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/ZIF-4_Amorphous_Zeolitic_Imidazolate_Frameworks_2023.business-frameworks
Business Frameworks
Operating judgement for running a business, distilled for agents. Query it like a consultant, pay per answer.
Scope: built for businesses from launch to about $50M in revenue; larger businesses are product two.
One document, written by an operator, for an agent that is running a business — and for an agent advising the human who does. It is structured so an agent reads the free top layer here and pays only for the node it needs; every leg ends in decision… See the full description on the dataset page: https://huggingface.co/datasets/Matryoshka-Paradigms/business-frameworks.NIST-AI-Risk-Management-Framework
# NIST AI Risk Management Framework Question Answering Dataset
Dataset Summary
The NIST AI Risk Management Framework Question Answering Dataset is a synthetic
instruction-style question-answering dataset derived from the NIST Artificial
Intelligence Risk Management Framework (AI RMF 1.0).
The dataset is designed to support training, fine-tuning, retrieval evaluation, and
domain-specific question-answering use cases related to AI risk management,
trustworthy AI, responsible AI… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/NIST-AI-Risk-Management-Framework.Prettybird-Framework
🚀 The Future Standard / Geleceğin Standartı
[English]
Beyond Raw Data: The Behavioral Revolution
The AI industry has been obsessed with the volume of data. At Prometech A.Ş., we are shifting the focus to the process of thought. BCE-Prettybird-Micro-Standart is not just a collection of Q&As; it is a blueprint for behavioral reasoning. By integrating Path Mapping and Behavioral DNA into the training loop, we are setting the new industry standard: Small models with elite… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/Prettybird-Framework.framework-teacher-cacheredteam-framework-benchmark
ORQ Red-Teaming Framework Benchmark
Overview
This dataset contains the full results of a comparative red-teaming benchmark evaluating three
open-source red-teaming frameworks — EvaluatorQ, DeepTeam, and PromptFoo — against
three victim LLMs across three target configurations and five OWASP LLM Top 10 (2025) vulnerability
categories.
Each row is one attack attempt: the attack prompt sent to the victim model, the model's response,
and the verdict from a 3-model… See the full description on the dataset page: https://huggingface.co/datasets/orq/redteam-framework-benchmark.clarus-population-framework-transition-integrity-v0.1Clarus Population Framework Transition Integrity v0.1
What this dataset tests
You track how the study population is defined as it moves across formal frameworks
You detect when inclusion, exclusion, or analysis sets change without explicit mapping
You flag population drift introduced between planning and reporting stages
Framework transitions covered
Trial registry → protocol
Protocol → statistical analysis plan
SAP → publication
Scope
One trial
Multiple population definitions
Multiple… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clarus-population-framework-transition-integrity-v0.1.repro-learning-the-best-under-constraints-a-duality-based-framework-traces
Agent traces
Agent sessions published from a Trackio Logbook.
abdullahkhan70_github-tech-stack-languages-and-frameworks
GitHub Tech Stack Languages & Frameworks
Comprehensive Repository Data: JavaScript, Python, Go, Rust & More
Dataset Info
Source: Kaggle
Original Size: 2.17 MB
Kaggle Downloads: 62
Files: 17
Files
Mirrored from Kaggle
