datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IndustryCorpus_technology[中文主页]
Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise.
To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_technology.information_technology_instruct_mcq_2481IndustryInstruction_Technology-Research
IndustryInstruction: Technology & Research
This repository contains the IndustryInstruction: Technology & Research domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Technology-Research.forge-3b-dpo-data
FORGE-3B DPO Preference Data
Tokenized (prompt, chosen, rejected) preference triples for DPO post-training
of FORGE-3B, built per the FORGE paper Section 6.2 / Appendix A.2.
This is data preparation output only — no model was trained to produce this.
Stats
Total pairs: 0 (paper target: ~200,000)
Domains: 0/4
Context length: 4096 tokens (paper Appendix A.2, DPO block)
Format: unpacked — one (prompt, chosen, rejected) triple per training example
Chat template:… See the full description on the dataset page: https://huggingface.co/datasets/Phase-Technologies/forge-3b-dpo-data.techno_datasetarayun_173-invariance-technology-system-law-agi
ARAYUN_173 Dataset
This dataset provides structured, machine-readable representations of the ARAYUN_173 research series.
Structure
Each record contains:
id
source
type
title
content
keywords
doi
language
Purpose
This dataset represents Paper 4 of the ARAYUN_173 research series in structured, machine-readable form.
It corresponds to:
ARAYUN_173: Invariance Technology and Invariant System-Law Architecture for AGI
DOI:
10.5281/zenodo.18179361… See the full description on the dataset page: https://huggingface.co/datasets/ARAYUN173/arayun_173-invariance-technology-system-law-agi.website-technology-evidence-dataset
Website Technology Evidence Dataset
A balanced synthetic dataset for classifying observable website fingerprint evidence into a likely web technology.
Provenance
All records are synthetically generated from documented, recognizable public fingerprints. No claim is made that these records were collected from real websites.
Dataset
38 technologies
6,840 examples
180 examples per technology
train: 5,472
validation: 684
test: 684
balanced classes… See the full description on the dataset page: https://huggingface.co/datasets/newazhala/website-technology-evidence-dataset.global_mmlu_lite_pt
🌎 Global-MMLU Lite (Portuguese)
A Focused Benchmark for Portuguese-Language Reasoning in Large Language Models
Global-MMLU Lite (Portuguese) is a curated subset of the Global-MMLU Lite benchmark designed to evaluate the reasoning, knowledge, and multiple-choice question-answering capabilities of large language models in Portuguese, providing a diverse and computationally efficient collection of translated and adapted QA samples across domains such as general knowledge, science… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/global_mmlu_lite_pt.Lab-2-language-technologyphyenem_2025
🇧🇷 ENEM 2025 — Brazilian National High School Exam Dataset
A High-Quality Benchmark for Portuguese Academic Reasoning in Large Language Models
ENEM 2025 Dataset is a curated collection of question-answer pairs derived from the 2025 edition of the Brazilian National High School Exam (ENEM), designed to evaluate and improve the reasoning, reading comprehension, and multiple-choice answering capabilities of large language models in Brazilian Portuguese; as one of the largest… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/enem_2025.simple_bench
📊 Simple Bench Dataset
A Compact Benchmark for Structured Reasoning and Multiple-Choice Evaluation in Large Language Models
Simple Bench Dataset is a structured evaluation collection derived from the Simple Bench benchmark, designed to assess reasoning, comprehension, and multiple-choice question-answering capabilities of large language models through concise yet non-trivial problems that require logical inference rather than simple retrieval; each sample consists of a natural… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/simple_bench.VGIBench
VGIBench
VGIBench is a video question-answering benchmark of human-validated multiple-choice
questions over long-form videos. The questions are designed to mitigate the common
mistakes of today's video benchmarks and to expose pragmatic challenges for
modern state-of-the-art models.
This is the public split (439 questions), released with answers so anyone can
score a model with exact-match. A held-out
private split is evaluated open-ended by the benchmark maintainers as a… See the full description on the dataset page: https://huggingface.co/datasets/Seldon-Technologies/VGIBench.global_mmlu_lite
🌍 Global-MMLU Lite Dataset
A Lightweight Benchmark for Multi-Domain Reasoning in Large Language Models
Global-MMLU Lite is a curated and efficient subset of the Global Massive Multitask Language Understanding (MMLU) benchmark, designed to evaluate and fine-tune large language models across a wide range of academic and professional domains through high-quality multiple-choice question answering; preserving the diversity and rigor of the original benchmark while significantly… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/global_mmlu_lite.schemaforge-engineering-and-technology-15
spectrum.ieee.org
Auto-refined by SchemaForge
Metadata
Topic: Engineering and Technology
Quality Score: 0.95
Source: Autonomous web scraper
Extracted Facts
IEEE Spectrum is the flagship publication of the IEEE.
IEEE is the world's largest professional organization devoted to engineering and applied sciences.
The Institute content is only available for IEEE Members.
Downloading full PDF issues is exclusive for IEEE Members.
Following topics is a… See the full description on the dataset page: https://huggingface.co/datasets/GudduButt/schemaforge-engineering-and-technology-15.OllamaDocsThis is a dataset generated from the documentation of Ollama as of
01/02/2025. The docs were fed into a model and then for every 10
words, another question was generated (roughly).
Was created with https://github.com/technovangelist/llm_dataset_builder
global_mmlu_lite_en
🌍 Global-MMLU Lite (English Only)
A Focused Benchmark for English-Language Reasoning in Large Language Models
Global-MMLU Lite (English Only) is a curated subset of the Global-MMLU Lite benchmark specifically designed to evaluate the reasoning, knowledge, and multiple-choice question-answering capabilities of large language models within the English language, providing a diverse yet computationally efficient collection of structured QA samples spanning domains such as science… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/global_mmlu_lite_en.schemaforge-ai-and-technology-1
news.ycombinator.com
Auto-refined by SchemaForge
Metadata
Topic: AI and Technology
Quality Score: 0.85
Source: Autonomous web scraper
Extracted Facts
New Zealand lost its music media
OpenChamber is an Agentic Development Environment
Andrew Wiles proved Fermat's Last Theorem in 1995
Tinnitus disappeared after being made a friend
Windows 11's Weather app wastes 1 GB of RAM
Microsoft Word for Windows 1.1a has a Native X64 Port
Georgia police… See the full description on the dataset page: https://huggingface.co/datasets/GudduButt/schemaforge-ai-and-technology-1.prompt2model-examples
Prompt2Model Toy Examples
Product: Prompt2Model:
a language-guided vision model factory. A typed pipeline (prompt, dataset config, training,
calibration/conformal abstain, ONNX export, an optional distill/quantize step with an
accuracy-floor gate, and a hard-case flywheel).
What this is (and isn't)
This is not a benchmark dataset. Prompt2Model has no natural "own" benchmark corpus the way a
task-specific product does. What's uploaded here is the repository's own… See the full description on the dataset page: https://huggingface.co/datasets/Dhi-Technologies/prompt2model-examples.math_precision_benchmarking
🧮 Math Precision — Benchmarking
A Formal Framework for High-Precision Arithmetic Evaluation in Large Language Models
Math Precision — Benchmarking, developed by Sapiens Technology®, is a rigorous framework for evaluating the true arithmetic capabilities of large language models by generating fully stochastic, high-precision mathematical problems that eliminate memorization and heuristic guessing; operating in a 100-digit floating-point field, it forces extreme numerical precision… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/math_precision_benchmarking.schemaforge-technology-news-13
arstechnica.com
Auto-refined by SchemaForge
Metadata
Topic: Technology News
Quality Score: 0.92
Source: Autonomous web scraper
Extracted Facts
Ars Technica serves the technologist since 1998
Mount Toba eruption had little climate impact
First self-driving vehicle on Mars was a success
DeepMind's hurricane breakthrough surprised weather scientists
Flesh-eating screwworms infected 500 humans in Mexico
Judge ruled Meta caused public nuisance and… See the full description on the dataset page: https://huggingface.co/datasets/GudduButt/schemaforge-technology-news-13.lglobal_mmlu_lite_es
🌎 Global-MMLU Lite (Spanish Only)
A Focused Benchmark for Spanish-Language Reasoning in Large Language Models
Global-MMLU Lite (Spanish Only) is a curated subset of the Global-MMLU Lite benchmark specifically designed to evaluate the reasoning, knowledge, and multiple-choice question-answering capabilities of large language models in Spanish, providing a diverse and computationally efficient collection of fully translated and standardized QA samples across domains such as science… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/global_mmlu_lite_es.ntadsschemaforge-technology-news-14
arstechnica.com
Auto-refined by SchemaForge
Metadata
Topic: Technology News
Quality Score: 0.92
Source: Autonomous web scraper
Extracted Facts
Ars Technica has been serving the technologist since 1998.
The pre-Google web was chaotic.
The Mount Toba eruption had little climate impact.
The first self-driving vehicle on Mars was a success.
Perseverance drove 90% of its distance autonomously.
DeepMind's hurricane breakthrough surprised weather… See the full description on the dataset page: https://huggingface.co/datasets/GudduButt/schemaforge-technology-news-14.qubikdatasetsampleamodal-counting-benchmark
Amodal Counting Benchmark
Product: amodal-counting, "count
what detectors can't see": visibility-corrected object counting through crowds, clutter, and
occlusion, reported as a calibrated interval rather than a bare point estimate.
This dataset is the exact evaluation population the product's own amodal bench command scores
against: procedurally generated scenes with known ground-truth occupancy, a simulated detector
with a known detectability curve, and the naive-vs-corrected… See the full description on the dataset page: https://huggingface.co/datasets/Dhi-Technologies/amodal-counting-benchmark.schemaforge-technology-news-23
www.engadget.com
Auto-refined by SchemaForge
Metadata
Topic: Technology News
Quality Score: 0.85
Source: Autonomous web scraper
Extracted Facts
Engadget publishes technology news
Kalshi sued for data misuse
Zoom screen sharing bug
Abbott partners with Google Health
MacBook can run multiple external monitors
Authors face backlash for AI study
iPhone changed in last five years
Disney to stream Formula E racing
Fix stick drift on PS5 and Xbox… See the full description on the dataset page: https://huggingface.co/datasets/GudduButt/schemaforge-technology-news-23.technode.global-mychem
