datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ODELIA-Challenge-2025
ODELIA Challenge Dataset
This dataset is part of the ODELIA project, a European Horizon initiative focused on developing privacy-preserving, AI-driven diagnostic tools using swarm learning.
The dataset provided here represents a curated subset of data from the broader ODELIA consortium. It is designed to facilitate the development, benchmarking, and validation of AI algorithms that can operate effectively across a range of heterogeneous clinical settings.
The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/ODELIA-AI/ODELIA-Challenge-2025.sit-latents-ode-heun-1000-class-0_1000-samples-segment-100-199odexODEX is an Open-Domain EXecution-based NL-to-Code generation data benchmark.
It contains 945 samples with a total of 1,707 human-written test cases,
covering intents in four different natural languages -- 439 in English, 90 in Spanish, 164 in Japanese, and 252 in Russian.sit-latents-ode-heun-1000-class-0_1000-samples-segment-400-499ODEN-Indcorpus
ODEN‑Indcorpus 📚
ODEN‑Indcorpus is a 3.7‑million‑line Odia mixed text collection curated from
fiction, dialogue, encyclopaedia, Q‑A and community writing derived from the ODEN initiative.After thorough normalisation and de‑duplication it serves as a robust substrate for training
Odia‑centric tokenizers, language models and embedding spaces.
Split
Lines
Train
3,373,817
Validation
187,434
Test
187,435
Total
3,748,686
The material ranges from conversational… See the full description on the dataset page: https://huggingface.co/datasets/BBSRguy/ODEN-Indcorpus.ode-preprocessing-hy15-testWaveForcing_Stage2_ODE_pair
WaveForcing Stage 2 — Wan2.1-T2V-14B ODE Endpoint Pairs 2K
本数据集包含 2,176 对文本条件视频生成端点数据:2,048 对训练数据及 128 对留出验证数据。每对数据保存文本提示词、初始高斯噪声 z_ref,以及同一提示词和噪声经 Wan2.1-T2V-14B teacher 去噪得到的最终 latent y_ref,可用于视频扩散模型的成对蒸馏或回归研究。
这里的 “ODE pairs” 指 初始噪声到最终去噪 latent 的端点对。数据没有保存 50 步采样过程中的中间状态,不是完整 ODE 轨迹数据集。
数据划分
Split
数量
pair_index
文件名
Train
2,048
0–2047
000000.pt–002047.pt
Validation
128
2048–2175
002048.pt–002175.pt
总计
2,176
0–2175
2,176 个 .pt 文件
划分由… See the full description on the dataset page: https://huggingface.co/datasets/Osc7/WaveForcing_Stage2_ODE_pair.sit-latents-ode-heun-segment-0-99odexODEX dataset annotated with the ground-truth library documentation, to enable evaluations for retrieval and retrieval-augmented code generation.
Please refer to [code-rag-bench] for more details.
odesia-combined-dipromats-v1only dipromats, downsample P
odesia-dipromats-seq-cls-v1seq cls dipro
odebenchode-enterprise-use-cases
ODE Enterprise Use Case Dataset
15,000 labeled enterprise use cases spanning 31 modules, 215 submodules, 8 industry verticals, 5 channels, and 12 business personas.
Published by Llewellyn Systems Inc — builders of ODE, the Operating System for Decision & Enterprise.
Attribution Required
This dataset is licensed under CC-BY-4.0. You are free to use, share, and adapt this dataset for any purpose — including commercial — as long as you give appropriate credit.
How… See the full description on the dataset page: https://huggingface.co/datasets/LlewellynSystemsInc/ode-enterprise-use-cases.flaws-cloudtrail-security-qa
CloudTrail Security Q&A Dataset
A comprehensive dataset of security-focused questions and answers based on AWS CloudTrail logs, designed for training and evaluating AI agents on cloud security analysis tasks.
Dataset Overview
This dataset contains:
~150 questions across 16 CloudTrail database partitions
Time period: February 2017 - August 2020
4 different AI models used for question generation
DuckDB databases with actual CloudTrail data
Mixed answerable/unanswerable… See the full description on the dataset page: https://huggingface.co/datasets/odemzkolo/flaws-cloudtrail-security-qa.cf_stage2_tf_ode_mixkit_6k_wan1.3bOden-catch-22sit-latents-ode-heun-1000-test2-class-0_1000-samples-segment-10-19odesia-combined-v2combined
sit-latents-ode-heun-1000-class-0_1000-samples-segment-600-699africa-tunisia-realisations-de-l-odesypano-agriculture-de-montagne-progra-53988af1
Realisations De L Odesypano Agriculture De Montagne Progra | Africa (Tunisia Open Data)
31 rows - 1 Africa country/area - detected - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 31 rows from Tunisia Open Data, covering Realisations De L Odesypano Agriculture De Montagne Progra. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-tunisia-realisations-de-l-odesypano-agriculture-de-montagne-progra-53988af1.p2-etf-liquid-neural-ode-resultsafrica-tunisia-plan-de-passation-des-marches-de-l-odesypano-au-titre-de-l-1d23b292
Plan De Passation Des Marches De L Odesypano Au Titre De L | Africa (Tunisia Open Data)
24 rows - 1 Africa country/area - 2019 - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 24 rows from Tunisia Open Data, covering Plan De Passation Des Marches De L Odesypano Au Titre De L. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-tunisia-plan-de-passation-des-marches-de-l-odesypano-au-titre-de-l-1d23b292.llama-email-finetune-1kemail-finetune-1500odesia-combined-v3combined, prompts v2
sit-latents-ode-heun-1000-class-0_1000-samples-segment-500-599africa-tunisia-realisations-de-l-odesypano-protection-des-rces-nalles-pro-e6d7de7e
Realisations De L Odesypano Protection Des Rces Nalles Pro | Africa (Tunisia Open Data)
33 rows - 1 Africa country/area - detected - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 33 rows from Tunisia Open Data, covering Realisations De L Odesypano Protection Des Rces Nalles Pro. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-tunisia-realisations-de-l-odesypano-protection-des-rces-nalles-pro-e6d7de7e.sit-latents-ode-heun-1000-class-0_1000-samples-segment-900-999ode-enterprise-use-cases
ODE Enterprise Use Case Dataset
15,000 labeled enterprise use cases spanning 31 modules, 215 submodules, 8 industry verticals, 5 channels, and 12 business personas.
Published by Llewellyn Systems Inc — builders of ODE, the Operating System for Decision & Enterprise.
Attribution Required
This dataset is licensed under CC-BY-4.0. You are free to use, share, and adapt this dataset for any purpose — including commercial — as long as you give appropriate credit.
How… See the full description on the dataset page: https://huggingface.co/datasets/LlewellynSystems/ode-enterprise-use-cases.ODEN-speech
ODEN‑speech 🗣️🇮🇳
Odia Diverse ENsemble Speech Corpus
ODEN‑speech merges eight publicly‑available Odia (ଓଡ଼ିଆ) speech corpora into a single 16 kHz, speaker‑aware, text‑cleaned dataset suitable for ASR, TTS, representation learning and multilingual research.
✨ Highlights
🗂️ Source
Hours
License
Mozilla Common Voice 17 (Odia)
110 h
MPL‑2.0
LibriTTS (clean + other)
170 h
CC‑BY‑4.0
LJSpeech 1.1
24 h
CC‑BY‑4.0
VCTK (Odia & misc.)
40 h
CC‑BY‑4.0
IndicTTS… See the full description on the dataset page: https://huggingface.co/datasets/BBSRguy/ODEN-speech.
