datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cmevs-erp-eval
CM-EVS: A Coverage-Curated Panoramic RGB-D Dataset for Indoor Scene Understanding
CM-EVS is a curated panoramic RGB-D dataset built under a single principle: maximize the geometric coverage of a 3D scene with the fewest equirectangular (ERP) frames possible. The release is structured as one redistributable Blender indoor data archive plus four license-aware adapter packages that regenerate matched frames locally from upstream sources whose terms forbid redistribution.
v1.0… See the full description on the dataset page: https://huggingface.co/datasets/anon-cmevs-2026/cmevs-erp-eval.eccv2026-cad-challenge-data
ECCV 2026 CAD Challenge Data
This challenge is part of the workshop The Path to Manufacturing: Evolving
3D Generation to Intelligent Computer-Aided Design.
Workshop homepage: https://3dgen-cad-workshop.github.io/
Challenge submission Space: https://huggingface.co/spaces/jingwei-xu-00/eccv2026-cad-challenge
Dataset rendering and preparation code (only .step files are required): https://github.com/DavidXu-JJ/eccv2026-cad-challenge-data-render
This repository contains the public… See the full description on the dataset page: https://huggingface.co/datasets/jingwei-xu-00/eccv2026-cad-challenge-data.brazil-2026-electoral-divergence
AFOS — Brazil 2026 Electoral Divergence Dataset
🌐 English · Português · Español
English
Open, auditable daily dataset that cross-references prediction markets (Polymarket) × polling institutes (TSE-registered) × press coverage for Brazil's 2026 presidential cycle, with explicit divergence between sources instead of smoothed averages.
Maintained by AFOS Analytics — open-source civic infrastructure for electoral political-risk intelligence. This is the public… See the full description on the dataset page: https://huggingface.co/datasets/AFOS-Analytics1/brazil-2026-electoral-divergence.jepa-qwen3-32b-pure-baselines-2026-05-25
JEPA-Align: Qwen3-32B Safety Defense Matrix
The complete 11-condition Qwen3-32B experiment for Predictive Representation
Alignment (PRA), the paired-view objective introduced in Predictive
Representation Alignment Improves Generalization in LLM Safety.
PRA aligns adversarially rewritten prompts with clean prompts expressing the
same intent. This release contains trained adapters, attack traces, benign
capability evaluations, machine-readable results, and paper-ready tables for… See the full description on the dataset page: https://huggingface.co/datasets/memo-ozdincer/jepa-qwen3-32b-pure-baselines-2026-05-25.2026.RA.Negotiation-Campaigns
Rational-Agent Negotiation Campaigns
This public dataset contains the complete selected evidence for the
ii_mats/experiments/rational_agents negotiation experiments. It includes raw
episode JSON, post-hoc annotations, Markdown and HTML transcripts, committed
instances, run manifests, campaign selection and exclusion ledgers,
machine-readable analysis tables, figures, and integrity manifests.
No contaminated, duplicated, stale, failed, or superseded run is included as
selected… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Negotiation-Campaigns.usa-2026-midterms-divergence
AFOS — US 2026 Midterms Divergence (v1.0.0, pre-electoral)
What the market priced and what the polls measured about the same election, day by day, before it happened.
⚠️ Pre-electoral bundle. The election is on November 3, 2026. There is no certified result to score these forecasts against, so this is not the gold-standard version. See DATASHEET.md. The v2, after certification, closes the cycle.
What is inside
443 national generic-ballot polls from 70… See the full description on the dataset page: https://huggingface.co/datasets/AFOS-Analytics1/usa-2026-midterms-divergence.hub-trending-models-2026-03-06binance-futures-ohlcv-2018-2026
🐱 币安期货 Main4 数据集 (BTC/ETH/BNB/SOL)
本项目由交易猫基金会支持(交易猫基金会 CA:0x8a99b8d53eff6bc331af529af74ad267f3167777)。
精简版币安期货历史数据 - 只包含 4 个主要币种,适合快速下载和研究使用。
📊 数据概览
文件
记录数
压缩大小
时间范围
说明
candles_1m_main4_*.bin.zst
998 万
366 MB
2020-01 ~ 2026-01
1分钟K线 (Binary)
futures_metrics_main4_*.bin.zst
152 万
49 MB
2021-12 ~ 2026-01
期货指标 (Binary)
schema_*.sql.zst
-
6.3 KB
-
TimescaleDB Schema
总计: 1150 万条记录,压缩后约 415MB
🎯 包含币种
币种
K线记录数
K线时间范围… See the full description on the dataset page: https://huggingface.co/datasets/tradecatlabs/binance-futures-ohlcv-2018-2026.2026.RA.Five-Seat-Frontier-Negotiation
Five-Seat Frontier Negotiation
⚠ ERRATUM (2026-08-10) — the omniscient-oracle arms in this bundle carry a spoiled ballot
OmniscientBestResponsePolicy — the computable seat in every *_oracle lineup here — cast its forced-final
vote on whichever live offer it valued most instead of on the one offer under the up/down vote. The protocol
rejects that as a legality error, the seat spends its single retry repeating itself, and the turn is recorded as
a pass: a silent… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Five-Seat-Frontier-Negotiation.dlam-ts-project-data-2026
operations_forecasting_2026
Multivariate hourly forecasting for anonymized operations units.
Target
Predict the future hourly operational load index for each series_id. Higher values indicate more operational pressure in that unit.
Forecast Contract
Frequency: h
Series: 96
Timesteps per series: 4992
Target column: target
Training history length used by the baseline templates: 168
Rollout block length: 24
Required prediction horizon: validation: 336, test: 336… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/dlam-ts-project-data-2026.layoffs-2026-us-warn-act-notices-daily
US layoffs in 2026: 2,795 WARN notices, 253,791 workers - rebuilt every morning
Data as of 2026-09-24. One row per WARN Act layoff notice filed with a US state
agency in 2026, normalised into one schema across 45 states. Free, CC BY 4.0,
no login, no API key, no delayed tier.
2,795
layoff notices filed in 2026 so far
253,791
workers named on them (a floor - see Honest scope)
45
state agencies filed at least one
1,035
most of any state: CA
9,891
largest… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/layoffs-2026-us-warn-act-notices-daily.hle-flowbench-experiments-20260829
HLE FlowBench experiment archive
Private migration snapshot of the local HLE text-only 100-question research
program through 2026-09-01. It preserves the formal and smoke runs, per-question
Codex/Claude/Kimi sessions, workflow attempts and metrics, evaluator state,
scores, monitoring, experiment controllers, reports, analyses, the paused-run
migration package, source Git bundles, and HLE-related host orchestration
sessions.
The current Chinese experiment status, validity… See the full description on the dataset page: https://huggingface.co/datasets/Changyeli03/hle-flowbench-experiments-20260829.eccv2026-cad-challenge-data
ECCV 2026 CAD Challenge Data
This challenge is part of the workshop The Path to Manufacturing: Evolving
3D Generation to Intelligent Computer-Aided Design.
Workshop homepage: https://3dgen-cad-workshop.github.io/
Challenge submission Space: https://huggingface.co/spaces/jingwei-xu-00/eccv2026-cad-challenge
This repository contains the public data package for the challenge. The
evaluation Space accepts STEP predictions for the private evaluation split and
updates the leaderboard… See the full description on the dataset page: https://huggingface.co/datasets/Qiao123rvvr/eccv2026-cad-challenge-data.AGC-Bench
AGC-Bench
AGC-Bench (Artificial General Creativity Benchmark) is a HELM-compatible evaluation suite for measuring creative ability in language and vision-language models. The release includes runnable benchmark scenarios, scoring code, release tables, and scripts that reproduce the paper's 83-model leaderboard. It covers 78 datasets: 67 text-only benchmarks in the primary analysis plus 11 multimodal-only scenarios released as artifacts. Domains span Brainstorming, Problem Solving… See the full description on the dataset page: https://huggingface.co/datasets/agcbench-2026/AGC-Bench.student-burnout-analysis2026
🔥 Predicting Academic Burnout: A Multivariate Analysis of Student Stressors
Exploring how financial pressure, family expectations, and social support shape burnout in university students.
Project Overview & Data Walkthrough
📋 Abstract
Academic burnout is an increasingly recognized phenomenon with far-reaching consequences for student wellbeing and performance. This study investigates the relationship between external environmental stressors —… See the full description on the dataset page: https://huggingface.co/datasets/eliel2003/student-burnout-analysis2026.ridgelora-cross-sensor-sd302d-f-to-m-20260825
RidgeLoRA-FP: SD302A-F to SD302D-M cross-sensor experiment
This public archive contains the leakage-controlled direct cross-sensor
experiment used to evaluate whether Stage-2 synthetic target-sensor images
help recognition on a physically different real sensor.
Locked protocol
Source/condition sensor: NIST SD302A device F.
Target sensor: NIST SD302D device M.
Identity: subject:finger-position; the same fingers exist across both
collections.
Subject split: 160… See the full description on the dataset page: https://huggingface.co/datasets/LamTNguyen/ridgelora-cross-sensor-sd302d-f-to-m-20260825.indeed-job-postings-2026
Indeed Jobs + Company Firmographics — Sample (2026)
A sample of 1,000 Indeed job postings across diverse roles and major US
cities, each enriched with parsed salary ranges, full job descriptions, location,
and joined company firmographics (536 unique employers: rating, size,
revenue, CEO, founded, website, socials).
This is a sample. To pull fresh, larger, or filtered data (60+ countries, native
salary-range filter, free company profiles), run the source actor:
Indeed Jobs… See the full description on the dataset page: https://huggingface.co/datasets/fact-den/indeed-job-postings-2026.tgk-ai-video-generators-2026
Permanent dataset archive: https://doi.org/10.5281/zenodo.22703594
Five AI Video Generators Tested on Dialogue, Action and an Advert
These Guys Know tested Seedance 2.5, MiniMax H3, FLUX 3 Video, Gemini Omni 1.1 Flash and HappyHorse 1.1 on 1 September 2026. Every model received the same three ten-second, 16:9 text-to-video tasks: a father interrupting a computer game, a three-person fight inside a fixed hotel lobby and a Mango Cola advert with an exact product name.
We retained… See the full description on the dataset page: https://huggingface.co/datasets/These-Guys-Know/tgk-ai-video-generators-2026.mer2026-features
MER2026 Track 1 — Quickstart Guide
Hướng dẫn từng bước để chạy training và tạo file submission cho MER-Cross (Track 1) sử dụng pre-extracted features tại HuggingFace: hhieupt/mer2026-features.
Mục lục
Mô tả bài toán và dữ liệu
Yêu cầu hệ thống
Clone repo ban tổ chức
Cài đặt môi trường
Tải dữ liệu từ HuggingFace
Giải nén và tổ chức thư mục
Tạo file config.py
Training
Tạo file submission
Lưu ý và mẹo
1. Mô tả bài toán và dữ liệu
Bài… See the full description on the dataset page: https://huggingface.co/datasets/hhieupt/mer2026-features.diffusion-mcqa-gen-pelatnas-2026
Which Prompt Made This? — Generated Edition
Pelatnas IOAI 2026 · Task Diffusion MCQA (varian trajectory)
Sebuah model text-to-image sedang bekerja. Di tengah prosesnya, gambar belum
menjadi gambar — yang ada hanya latent ter-noise: tensor 4 × 64 × 64 berisi
campuran struktur yang mulai muncul dan derau Gaussian.
Kali ini kalimat itu harfiah. Latent yang kamu terima benar-benar diambil dari
tengah proses generate: sebuah trajectory denoising DDIM 50 langkah dihentikan
sejenak… See the full description on the dataset page: https://huggingface.co/datasets/fassabilf/diffusion-mcqa-gen-pelatnas-2026.2026.RA.Frontier-and-Scale-Cells
Rational-Agent Frontier, Scale, and Framing Cells
This public dataset is a sibling of siddharthmb/2026.RA.Negotiation-Campaigns (the frozen P1-P4 experimental record for the ii_mats/experiments/rational_agents negotiation program) and follows the same conventions: raw per-episode JSON, per-turn oracle annotations, Markdown/HTML transcripts, run manifests, analysis tables, and an integrity manifest over every uploaded file. It packages eight later campaigns that were run against… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Frontier-and-Scale-Cells.aliafzal9323_soxx-ishares-semiconductor-etf-daily-2001-2026
SOXX iShares Semiconductor ETF Daily (2001-2026)
Daily OHLCV price data for iShares Semiconductor ETF (SOXX) spanning 24+ years
Dataset Info
Source: Kaggle
Original Size: 0.11 MB
Kaggle Downloads: 6
Files: 1
Files
SOXX_Daily_Stock_Data.csv
Mirrored from Kaggle
evalita2026
This repository contains the data release for the Cruciverb-IT shared task on automatic crossword solving in Italian, as part of the 2026 EVALITA campaign. Refer to the task website for more details.
The data from both tasks can be downloaded from the 'Files and versions' tab.
Updates:
Minor update to both task_*_scorer.py in order to convert accented letters to their non-accented counterpart during evaluation
Test data is out!!
The test data of both… See the full description on the dataset page: https://huggingface.co/datasets/cruciverb-it/evalita2026.hse-acronym-dictionary-2026
Canonical landing page: https://www.smartqhse.com/datasets/hse-acronym-dictionary-2026
HSE Acronym Dictionary 2026
Authoritative reference of 150+ HSE / EHS / occupational-safety acronyms with one-line definitions. Covers metrics (TRIR, LTIFR, DART, EMR, WBGT), regulations (OSHA, RIDDOR, COSHH, CDM, COMAH, PSM, OSHAD-SF), methodologies (HAZOP, HAZID, LOPA, SIL, FMEA, RCA, BBS, ICAM, TapRooT, Tripod Beta), bodies (NEBOSH, IOSH, IIRSM, IOGP, ACGIH, NIOSH, ANSI, ASSP), and… See the full description on the dataset page: https://huggingface.co/datasets/SmartQHSE/hse-acronym-dictionary-2026.sinhala-tts-dataset-archive-20260429-082457
Sinhala TTS Dataset
Clean, segmented Sinhala speech from the "Unlimited History" YouTube series by @sunchare.
Stats
Metric
Value
Utterances
218
Train
208
Val
10
Hours
0.51
Mean duration
8.5s
Sample rate
22050 Hz
Pipeline
Raw YouTube audio -> HTDemucs -> VoiceFixer + DeepFilterNet3 ->
Diarization -> Silero-VAD -> ASR (faster-whisper: C:\Users\kosal\sinhala-tts\whisper-small-si-ct2) -> Quality filtering (SNR>=20.0dB)
Format… See the full description on the dataset page: https://huggingface.co/datasets/outlawmold/sinhala-tts-dataset-archive-20260429-082457.2026.RA.Five-Seat-MovesOnly-Control
Five-Seat Moves-Only Controls: all seats mute, and one Opus seat mute
This repository holds the public evidence bundles for two control arms of the five-seat frontier negotiation campaign, both run on the identical frozen bank, seeds, model and protocol so that every episode pairs on (instance_id, episode_seed) against the five frozen arms in 2026.RA.Five-Seat-Frontier-Negotiation:
all_llm_moves_only (experiment-name five-seat-moves-only-control-v1, added 2026-08-10): five… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Five-Seat-MovesOnly-Control.2026.RA.Quorum-Rounds-Sweep
2026.RA.Quorum-Rounds-Sweep — what a decision rule does to a table of rational negotiators
Every episode of the quorum x rounds sweep: five computable Bayesian-rational negotiators bargaining over a
package of four issues, replayed on one frozen 24-game bank under three different agreement rules and two
different deadlines, plus a two-factor extension that also dissolves the veto.
What the experiment asks
Five parties must agree on one package out of 256. Each… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Quorum-Rounds-Sweep.lunarness-fashion-data-observatory-2026
Lunarness Fashion Data Observatory 2026
This package is the machine-readable companion to the Lunarness Fashion Data Observatory 2026. It combines a small set of cited public findings about Fashion Week media attention, historical event attendance, online fashion behavior and textile circularity with clearly labeled arithmetic derivations.
The Lunarness page is the canonical editorial presentation. The same versioned package is distributed through the following research and… See the full description on the dataset page: https://huggingface.co/datasets/FoodSecuriry/lunarness-fashion-data-observatory-2026.causalverify-neurips2026
🎯 CausalVerify
An Execution-Grounded Benchmark for LLM Causal Inference Workflows
NeurIPS 2026 — Evaluations and Datasets Track · double-blind review · frozen at tag neurips2026-submission
💡 TL;DR
A benchmark of 259 published economics papers (Experiment A — real-paper text-agreement diagnostic) and 100 fixed-seed synthetic data-generating processes (Experiment B — execution-grounded coefficient recovery), evaluating 7 frontier LLMs. The central… See the full description on the dataset page: https://huggingface.co/datasets/causalverify/causalverify-neurips2026.named-process-safety-incidents-extended-2026
Canonical landing page: https://www.smartqhse.com/datasets/named-process-safety-incidents-extended-2026
Named Process Safety and Industrial Disasters — Extended Reference 2026
Curated reference of 40 named historical process-safety, industrial, and major-fire disasters with dates, fatalities, casual factors, and regulatory consequences. Spans 1917–2024. Covers Bhopal, Piper Alpha, Texas City, Deepwater Horizon, Buncefield, Flixborough, Seveso, Phillips 66 Pasadena, Longford… See the full description on the dataset page: https://huggingface.co/datasets/SmartQHSE/named-process-safety-incidents-extended-2026.
