datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
waqfeya-library
Waqfeya Library
📖 Overview
Waqfeya is one of the primary online resources for Islamic books, similar to Shamela. It hosts more than 10,000 PDF books across over 80 categories.
In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX.
📊 Dataset Contents
The dataset includes 22,443 PDF files (spanning 8,978,634 pages) representing 10,150 Islamic books. Each book is… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/waqfeya-library.shamela-waqfeya-library
Shamela Waqfeya Library
📖 Overview
Shamela Waqfeya is one of the primary online resources for Islamic books, similar to Shamela. It hosts more than 4,500 PDF books across over 40 categories.
In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX.
📊 Dataset Contents
The dataset includes 12,877 PDF files (spanning 5,138,027 pages) representing 4,661 Islamic books.… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/shamela-waqfeya-library.real-or-fake-fake-jobposting-predictionwaqfeya-library-compressed
Waqfeya Library - Compressed
📖 Overview
Waqfeya is one of the primary online resources for Islamic books, similar to Shamela. It hosts more than 10,000 PDF books across over 80 categories.
In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX.
📊 Dataset Contents
This dataset is identical to ieasybooks-org/waqfeya-library, with one key difference: the contents… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/waqfeya-library-compressed.CUDA-L2
CUDA-L2: Surpassing cuBLAS Performance for Matrix Multiplication through Reinforcement Learning
🥳 Introduction
CUDA-L2 is a system that combines large language models (LLMs) and reinforcement learning (RL) to automatically optimize Half-precision General Matrix Multiply (HGEMM) CUDA kernels. CUDA-L2 systematically outperforms major matmul baselines to date, from the widely-used torch.matmul to state-of-the-art NVIDIA closed-source libraries (cuBLAS… See the full description on the dataset page: https://huggingface.co/datasets/ornith-ai/CUDA-L2.real-or-fake-fake-jobposting-predictionorca-whirlpool-historical-data
Orca Whirlpools Historical Data
Decoded Solana mainnet instructions and events from Orca Whirlpools, Orca's concentrated-liquidity AMM: swaps, position open/close, liquidity changes, fee and reward collection, and the Token-2022 (_v2) instruction family.
54 tables, 37,021 rows, one row per decoded instruction or event.
This is a free sample from datastore.sh, which publishes the complete history as versioned Parquet.
Read this before you analyse it
The sample is… See the full description on the dataset page: https://huggingface.co/datasets/DataStore/orca-whirlpool-historical-data.oregon-layoffs-warn-act-notices-daily
Oregon WARN Act layoff notices — every filing we hold since 1988, one CSV, rebuilt daily
1,371 Oregon WARN notices — every one this dataset holds, back to 1988 — free to download in full: no paywalled years, no login, no account · most recent notice filed 2026-09-23
· state source last checked 2026-09-24T14:04Z · official source: Oregon Higher Education Coordinating Commission — WARN notices.
Oregon employers must file a WARN Act notice with the state before a qualifying
mass… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/oregon-layoffs-warn-act-notices-daily.finer-ord
Dataset Card for "FiNER-ORD"
Dataset Summary
The FiNER-Open Research Dataset (FiNER-ORD) consists of a manually annotated dataset of financial news articles (in English)
collected from webz.io.
In total, there are 47851 news articles available in this data at the point of writing this paper.
Each news article is available in the form of a JSON document with various metadata information like
the source of the article, publication date, author of the article, and the title… See the full description on the dataset page: https://huggingface.co/datasets/gtfintechlab/finer-ord.real-or-fake-fake-jobposting-predictionOriginal-alpha-suppression-task-boostreal-or-fake-fake-jobposting-predictiontri-attia-fast-charge-2020-raw
Closed-loop optimization of extreme fast charging for batteries
BSEBench status: raw_mirror_pending_validation
This dataset repository is a raw mirror of the TRI Energy & Materials Data / data.matr.io project "Closed-loop optimization of extreme fast charging for batteries using machine learning". The source describes commercial lithium-ion phosphate (LFP)/graphite cells cycled under fast-charging conditions for the associated Nature publication.
Source page:… See the full description on the dataset page: https://huggingface.co/datasets/bsebench-org/tri-attia-fast-charge-2020-raw.SceneFly
CaR
Compression and Retrieval: Implicit Memory Retrieval for Video World Models
Zhan Peng1,2,
Jie Ma2,
Huiqiang Sun1,
Chong Gao2,3,
Zhijie Xue1,
Zhiyu Pan1,
Zhiguo Cao1*,
Jun Liang2*,
Jing Li2
1Huazhong University of Science and Technology
2HUJING Digital Media & Entertainment Group
3Sun Yat-sen University
*Corresponding author
SceneFly
SceneFly is a curated video dataset organized by synthetic 3D scenes. Each selected video… See the full description on the dataset page: https://huggingface.co/datasets/Orange-3DV-Team/SceneFly.e-commerce-orders
E-commerce Customer Order Behavior Dataset
A synthetic e-commerce dataset containing 10,000 orders with realistic customer behavior patterns, suitable for e-commerce analytics and machine learning tasks.
Dataset Card for E-commerce Orders
Dataset Summary
This dataset simulates customer order behavior in an e-commerce platform, containing detailed information about orders, customers, products, and delivery patterns. The data is synthetically generated with… See the full description on the dataset page: https://huggingface.co/datasets/millat/e-commerce-orders.Original-no-persona-replacement-remainderwooden_window_factory_01_enriched_v2
Real industrial data, AI-ready for Physical AI
ORION WWF1 – Certified Sample Pack v2.0 (Enriched)
Version
Status
Sector
Pipeline
v2.0-Enriched
🟢 Level 3 Certified
Industrial-Manufacturing
Orion Unified V5.2
🌟 The Evolution: Beyond Anonymization
The ORION WWF1 v2.0 Enriched pack represents the professional evolution of our baseline industrial dataset. While previous versions focused on privacy-first anonymization, v2.0 transforms raw video… See the full description on the dataset page: https://huggingface.co/datasets/Orion-The-Lab/wooden_window_factory_01_enriched_v2.sales-orders-us
US Sales Orders & Invoices Dataset
Complete sales pipeline for a simulated US retail SME: customers, orders, order lines, invoices, and payment allocations. Customer behaviour is stochastic with seasonal patterns — not uniform random. Payment timing varies by customer reliability.
Tables
Table
Rows
companies
1
customers
163
payment_allocations
2,271
sales_invoice_lines
11,065
sales_invoices
2,441
sales_order_lines
11,065
sales_orders
2,441
Total… See the full description on the dataset page: https://huggingface.co/datasets/mindweave/sales-orders-us.Original-hybrid-shared-no-persona-remainderSWEbench-Verified-eval150-u355-M2.7-orch-cost-20260923
Fixed Solo350 u355 + MiniMax-M2.7: orchestration cost study
Best observed cost tradeoff: compact coordinator decisions plus soft review at the existing hard limit (at most 12 worker turns). M2.7 metered token cost falls 59.3%, while mean solved tasks decrease from 90.00 to 87.67/150. Accuracy equivalence was not established.
This closed study contains 6 designs and 16 complete independent runs on the same 150 tasks (2400 scored task/run pairs), each with an independent audit.… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-u355-M2.7-orch-cost-20260923.oruk-bench-leaderboard
oruk-bench leaderboard
Results for 64 speech emotion recognition systems measured on one held-out multilingual
evaluation with a single scoring implementation: open checkpoints, closed APIs, audio LLMs, and
text-only baselines, all on the same protocol.
This dataset is the results table, not the audio. The evaluation clips are assembled from
several emotional-speech corpora whose licences differ, so they are not redistributable; the
benchmark card documents
provenance and how… See the full description on the dataset page: https://huggingface.co/datasets/oruk/oruk-bench-leaderboard.abbas829_global-organized-crime-index-20212025
Global Organized Crime Index (2021–2025)
Global Organized Crime Index dataset for year 2021, 2023 & 2025
Dataset Info
Source: Kaggle
Original Size: 0.04 MB
Kaggle Downloads: 47
Files: 1
Files
global_oc_index.csv
Mirrored from Kaggle
Original-hybrid-correct-train-no-persona-meanfood-delivery-orders
Food Delivery Platform Orders Dataset (Free Sample)
This is a free sample with 2,530 rows. The full dataset has 33,630 rows across 4 tables.
Order records for a simulated food delivery platform operating in a mid-size
US city. 120 restaurants, 45 delivery drivers, 8,000 customers, and 30,000
orders over 12 months.
Features realistic patterns: dinner rush (6-9 PM), weekend peaks, weather-driven
demand, restaurant ratings, delivery time estimates vs actuals, and two anomalies —
a… See the full description on the dataset page: https://huggingface.co/datasets/mindweave/food-delivery-orders.shamela-waqfeya-library-compressed
Shamela Waqfeya Library - Compressed
📖 Overview
Shamela Waqfeya is one of the primary online resources for Islamic books, similar to Shamela. It hosts more than 4,500 PDF books across over 40 categories.
In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX.
📊 Dataset Contents
This dataset is identical to ieasybooks-org/shamela-waqfeya-library, with one key… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/shamela-waqfeya-library-compressed.Original-hybrid-train-no-persona-meanOriginal-circuit-discoveryInternational_Classification_Diseases_Clinical_Modification_icd10cm_order_April_2024SWEbench-Verified-eval150-M2.7-Qwen3.5-9B-orch-7arms-2repeats-w32-20260920
SWE-bench Verified eval150 — M2.7 × Qwen3.5-9B, seven arms, two repeats, 32 concurrency
Campaign 2026-09-20. 14/14 independent full150 runs audited. Complete accuracy evidence.
Evaluation mode is orch: MiniMax-M2.7 orchestrator and the specified Qwen3.5-9B worker. Training mode is labeled independently. All runs use 32 concurrent episodes, 10GiB Docker sandboxes, four TP1 workers and one TP4/EP4 coordinator. Frozen regression-gated prompts, decoding and canonical verifier match… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-Qwen3.5-9B-orch-7arms-2repeats-w32-20260920.Original-baseline-bias-unbias
