datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Technical-Architectures-Large
Technical Architectures Large (294k Samples)
Overview
Generating complex, syntactically valid diagram code from natural language requirements is a major challenge for AI models. This dataset bridges that gap by providing over 293,000+ distinct enterprise software architectures generated using two cutting-edge models: GPT-OSS-120B and Qwen3-Coder-Next-FP8.
Unlike simple "toy" examples, these architectures model realistic enterprise systems complete with client… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Technical-Architectures-Large.agentbattler-bench
AgentBattler Mini Ledger V5
Immutable evidence for 15/15 accepted Mini Ledger V5 runs across 3 harness × model conditions.
What is here
snapshots/mini-ledger-v5-r5-droid-2026-08-04t13-03-57-476z/site/terminal-campaign.json: compact website and analysis input.
snapshots/mini-ledger-v5-r5-droid-2026-08-04t13-03-57-476z/campaign.json: source-revision-preserving campaign index with host paths removed.
snapshots/mini-ledger-v5-r5-droid-2026-08-04t13-03-57-476z/runs/:… See the full description on the dataset page: https://huggingface.co/datasets/techfren/agentbattler-bench.Neurips2026_evaluating_accuracy_KV-cache_reuse_techniques
BoxOffice Verified Seeds
This dataset contains the released BoxOffice seed datasets used in the
benchmark pipeline described in the accompanying paper. The release includes
ten verified seeds:
7
11
13
17
19
23
29
31
47
73
For each seed, we provide:
a full JSONL file containing warmup rows plus evaluation rows
an eval JSONL file containing only the evaluation rows
a manifest JSON file
a validation JSON file with directional warmup counts
Layout
viewer/
normalized… See the full description on the dataset page: https://huggingface.co/datasets/Boxoffice1280/Neurips2026_evaluating_accuracy_KV-cache_reuse_techniques.IndustryInstruction_Technology-Research
IndustryInstruction: Technology & Research
This repository contains the IndustryInstruction: Technology & Research domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Technology-Research.Tech-Sentences-For-ASR-Training
TechVoice Dataset
Work in Progress – This dataset is actively being expanded with new recordings.
Dataset Statistics
Metric
Current
Target
Progress
Duration
38m 43s
5h 0m 0s
██░░░░░░░░░░░░░░░░░░ 12.9%
Words
10,412
50,000
████░░░░░░░░░░░░░░░░ 20.8%
Total Recordings: 205 samples
Total Characters: 74,312
A specialized speech dataset for fine-tuning Automatic Speech Recognition (ASR) models on technical and developer vocabulary. Contains human-recorded… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Tech-Sentences-For-ASR-Training.ru-tech-jobs
Russian Tech Jobs
Dataset Description
Dataset contains job vacancy posts from one of Telegram channels focusing on IT and tech recruitment.
Data Fields
id: Unique identifier of the post.
date: ISO timestamp of when the job was posted.
text: The raw text of the post with markdown.
views: The view count of the post at the time of scraping.
tags: Special tags thats will be taken from text.
How to use
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/exeyarikus/ru-tech-jobs.factbutcher-benchmarkEnglish · Русский
FactButcher Russian Fact-Checking Dataset
This dataset contains 423 claims in Russian and the results of checking them.
Some claims are true, some are false, and some allow more than one defensible
answer.
You can see what kinds of claims people bring to fact-checkers, investigate a
few of them yourself, or use the complete collection to compare different
fact-checking tools.
A few examples
Claim
Result
Origin
Слоны боятся мышей
False… See the full description on the dataset page: https://huggingface.co/datasets/teplitsa-soc-tech/factbutcher-benchmark.breathing-techniques
Breathing Techniques
16 evidence-based breathing practices for emotional regulation, with contraindications, session guidance, difficulty levels, and primary benefits.
Quick Start
from datasets import load_dataset
ds = load_dataset("buley/breathing-techniques")
print(ds["train"][0])
Structure
Field
Description
id
Unique identifier
name
Technique name
category
Foundational/Calming, Energizing, Advanced, Specialized
difficulty_level
Beginner… See the full description on the dataset page: https://huggingface.co/datasets/buley/breathing-techniques.khmer-nlp-technical-corpus
khmer-nlp-technical-corpus — Khmer Strategic NLP Corpus
Dataset Summary
This dataset contains peer-grade long-form technical treatises (3,000+ words each) in the Khmer language (km / ភាសាខ្មែរ). Every article is normalized and features neural BiGRU+CRF word segmentation with Zero-Width Space (\u200B) injection to prevent token fragmentation in sub-word tokenizers.
Dataset Statistics
Total Documents: 3
Train Documents: 3
Total Words: 8,002
Total… See the full description on the dataset page: https://huggingface.co/datasets/guanvireak/khmer-nlp-technical-corpus.KurdishCorpus-clean
The Largest Documented Kurmanji and Multi-Dialect Kurdish Dataset — v1.1
A large-scale, multi-dialect Kurdish text corpus prioritizing native Kurmanji
(Northern Kurdish) fluency, with substantial Sorani (Central Kurdish) and a
Zazaki baseline. Built for language modeling, tokenizer training, and
general-purpose Kurdish NLP.
This release contains only openly-licensed or presumptively-free
redistributable content. A parallel research-tier subset (copyrighted
commercial… See the full description on the dataset page: https://huggingface.co/datasets/kurdish-tech/KurdishCorpus-clean.VGIBench
VGIBench
VGIBench is a video question-answering benchmark of human-validated multiple-choice
questions over long-form videos. The questions are designed to mitigate the common
mistakes of today's video benchmarks and to expose pragmatic challenges for
modern state-of-the-art models.
This is the public split (439 questions), released with answers so anyone can
score a model with exact-match. A held-out
private split is evaluated open-ended by the benchmark maintainers as a… See the full description on the dataset page: https://huggingface.co/datasets/Seldon-Technologies/VGIBench.ppRegionTestppAllTestai-detector-ref-enbell-labs-technical-archive
Bell Labs Documents and Stuff
This is a conservative public-release subset of the internal BELLA continued-pretraining corpus. It keeps the Bell-system technical material that survived a stricter final pass for public dataset hosting and removes records that still looked risky, off-scope, or too low-signal for a Hugging Face corpus listing.
What is in the release
Split
Documents
train
1220
validation
29
test
42
The release contains 1291 documents out… See the full description on the dataset page: https://huggingface.co/datasets/hunterbown/bell-labs-technical-archive.Technical-Architectures-Large
Technical Architectures Large (210k+ Samples)
Overview
Generating complex, syntactically valid diagram code from natural language requirements is a major challenge for AI models. This dataset bridges that gap by providing over 210,000 distinct enterprise software architectures generated using two cutting-edge models: GPT-OSS-120B and Qwen3-Coder-Next-FP8.
Unlike simple "toy" examples, these architectures model realistic enterprise systems complete with client… See the full description on the dataset page: https://huggingface.co/datasets/autoshift/Technical-Architectures-Large.ppRegionTraincityUnlTrainppLatValppCityTestqubikdatasetsampleamodal-counting-benchmark
Amodal Counting Benchmark
Product: amodal-counting, "count
what detectors can't see": visibility-corrected object counting through crowds, clutter, and
occlusion, reported as a calibrated interval rather than a bare point estimate.
This dataset is the exact evaluation population the product's own amodal bench command scores
against: procedurally generated scenes with known ground-truth occupancy, a simulated detector
with a known detectability curve, and the naive-vs-corrected… See the full description on the dataset page: https://huggingface.co/datasets/Dhi-Technologies/amodal-counting-benchmark.ppCityTrainppLngTrainppAllValppRegionValppZipTrainppLatTrainppZipTestppLatTest
