datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Treble10-RIR
Treble10-RIR (32 kHz)
The Treble10-RIR dataset is a dataset for automatic speech recognition (ASR), containing high fidelity room-acoustic simulations from 10 different furnished rooms: 2 bathrooms, 2 bedrooms, 2 living rooms with hallway, 2 living rooms without hallway, 2 meeting rooms.
The room volumes range between 14 and 46 m3, resulting in reverberation times between 0.17 and 0.84 s.
Illustrative plots of the rooms and device included in this dataset may be found in the… See the full description on the dataset page: https://huggingface.co/datasets/treble-technologies/Treble10-RIR.industrial-technical-archive
🚀 Latest Updates (July, 2026)
Version: v07.2026 (Verified)
Status: Integrated with 1,000,000+ records.
New Files: product-E-26-07-2026.csv & product-V-26-07-2026.csv.
QTE Technologies: Industrial & Scientific Knowledge Base
Wikidata Entity: Q138411149
IPFS CID: bafybeibogxxuhmzfrsuhcfd4qr4tmc4okhmrcwhp3266hq47ccuyjnjxoq
Official Neural Hub: qtetech.github.io
This is the permanent technical archive for QTE Technologies, ensuring long-term accessibility of… See the full description on the dataset page: https://huggingface.co/datasets/QTE-Technologies/industrial-technical-archive.v2-rc2medical-prescriptionsv2-r2-longafrica-synth-disability-assistive-technology-access-all
Assistive Technology & Devices Access (SSA) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: technology_digital - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-disability-assistive-technology-access-all.physicsussr_typewriter_ocr_bench
Document OCR using GLM-OCR
This dataset contains OCR results from images in technocreep/ussr_typewriter using GLM-OCR, a compact 0.9B OCR model achieving SOTA performance.
Processing Details
Source Dataset: technocreep/ussr_typewriter
Model: zai-org/GLM-OCR
Task: text recognition
Number of Samples: 50
Processing Time: 1.5 min
Processing Date: 2026-02-25 12:54 UTC
Configuration
Image Column: image
Output Column: markdown
Dataset Split: train
Batch Size:… See the full description on the dataset page: https://huggingface.co/datasets/technocreep/ussr_typewriter_ocr_bench.prompt2model-examples
Prompt2Model Toy Examples
Product: Prompt2Model:
a language-guided vision model factory. A typed pipeline (prompt, dataset config, training,
calibration/conformal abstain, ONNX export, an optional distill/quantize step with an
accuracy-floor gate, and a hard-case flywheel).
What this is (and isn't)
This is not a benchmark dataset. Prompt2Model has no natural "own" benchmark corpus the way a
task-specific product does. What's uploaded here is the repository's own… See the full description on the dataset page: https://huggingface.co/datasets/Dhi-Technologies/prompt2model-examples.publicTechnology-Images-DatasetDataset Description:
This Technology image dataset is part of a large-scale STEM image dataset containing 383,697 images, designed to support the development and training of advanced computer vision, educational AI, and multimodal learning systems.
Additionally, this dataset can be used in pipelines for Supervised Fine-Tuning (SFT) and Reinforcement Learning with Human Feedback (RLHF) workflows, improving model performance in scientific image understanding, visual reasoning, concept… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Technology-Images-Dataset.2023_charged_upussr_typewritertechno_albumcsszengardenambiguous-technology
