datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hate_speech18These files contain text extracted from Stormfront, a white supremacist forum. A random set of
forums posts have been sampled from several subforums and split into sentences. Those sentences
have been manually labelled as containing hate speech or not, according to certain annotation guidelines.ode_step5bigann-100m-static-search-eval
bigann-100m-static-search-eval
Static 100M search-evaluation dataset with original vectors, original query/ground-truth files, a rebuilt hnswlib HNSW index, and generated PQ artifacts.
Files
Base: base.u8bin
HNSW index: index_m_32_ef_500
Query: orig_query_10k.u8bin
Ground truth: groundtruth.bin
Checksums: checksums.sha256
Kaggle metadata is provided via dataset-metadata.json during upload and is stored by Kaggle as dataset configuration, not as a listed data… See the full description on the dataset page: https://huggingface.co/datasets/Odeinjul/bigann-100m-static-search-eval.ODELIA-Challenge-2025
ODELIA Challenge Dataset
This dataset is part of the ODELIA project, a European Horizon initiative focused on developing privacy-preserving, AI-driven diagnostic tools using swarm learning.
The dataset provided here represents a curated subset of data from the broader ODELIA consortium. It is designed to facilitate the development, benchmarking, and validation of AI algorithms that can operate effectively across a range of heterogeneous clinical settings.
The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/ODELIA-AI/ODELIA-Challenge-2025.deep-100m-batch-update-eval
deep-100m-batch-update-eval
Batch-update evaluation workload package generated from deep-100m-static-search-eval.
Dataset
Source static dataset: deep-100m-static-search-eval
Vector count: 100,000,000
Dimension: 96
Dtype: float32
Metric: l2
Initial update index: 80,000,000 vectors with external labels equal to A = P[0:80M]
Update order: update_order.u32, a seed-42 permutation of source IDs [0, 100M)
Insert vector source: base_permuted.fbin, where row j equals… See the full description on the dataset page: https://huggingface.co/datasets/Odeinjul/deep-100m-batch-update-eval.deep-100m-static-search-eval
deep-100m-static-search-eval
Static 100M search-evaluation dataset with original vectors, original query/ground-truth files, a rebuilt hnswlib HNSW index, and generated PQ artifacts.
Files
Base: base.fbin
HNSW index: index_m_32_ef_500
Query: orig_query_10k.fbin
Ground truth: groundtruth.bin
Checksums: checksums.sha256
Kaggle metadata is provided via dataset-metadata.json during upload and is stored by Kaggle as dataset configuration, not as a listed data file.… See the full description on the dataset page: https://huggingface.co/datasets/Odeinjul/deep-100m-static-search-eval.spacev-100m-batch-update-eval
spacev-100m-batch-update-eval
Batch-update evaluation workload package generated from spacev-100m-static-search-eval.
Dataset
Source static dataset: spacev-100m-static-search-eval
Vector count: 100,000,000
Dimension: 100
Dtype: int8
Metric: l2
Initial update index: 80,000,000 vectors with external labels equal to A = P[0:80M]
Update order: update_order.u32, a seed-42 permutation of source IDs [0, 100M)
Insert vector source: base_permuted.i8bin, where row j equals… See the full description on the dataset page: https://huggingface.co/datasets/Odeinjul/spacev-100m-batch-update-eval.turing-100m-batch-update-eval
turing-100m-batch-update-eval
Batch-update evaluation workload package generated from turing-100m-static-search-eval.
Dataset
Source static dataset: turing-100m-static-search-eval
Vector count: 100,000,000
Dimension: 100
Dtype: float32
Metric: l2
Initial update index: 80,000,000 vectors with external labels equal to A = P[0:80M]
Update order: update_order.u32, a seed-42 permutation of source IDs [0, 100M)
Insert vector source: base_permuted.fbin, where row j… See the full description on the dataset page: https://huggingface.co/datasets/Odeinjul/turing-100m-batch-update-eval.sit-latents-ode-heun-1000-class-0_1000-samples-segment-100-199bigann-100m-batch-update-eval
bigann-100m-batch-update-eval
Batch-update evaluation workload package generated from bigann-100m-static-search-eval.
Dataset
Source static dataset: bigann-100m-static-search-eval
Vector count: 100,000,000
Dimension: 128
Dtype: uint8
Metric: l2
Initial update index: 80,000,000 vectors with external labels equal to A = P[0:80M]
Update order: update_order.u32, a seed-42 permutation of source IDs [0, 100M)
Insert vector source: base_permuted.u8bin, where row j… See the full description on the dataset page: https://huggingface.co/datasets/Odeinjul/bigann-100m-batch-update-eval.odexODEX is an Open-Domain EXecution-based NL-to-Code generation data benchmark.
It contains 945 samples with a total of 1,707 human-written test cases,
covering intents in four different natural languages -- 439 in English, 90 in Spanish, 164 in Japanese, and 252 in Russian.tti-100m-static-search-eval
tti-100m-static-search-eval
Static 100M search-evaluation dataset with original vectors, original query/ground-truth files, a rebuilt hnswlib HNSW index, and generated PQ artifacts.
This Hugging Face copy provides the static-search-eval package for tti, including static artifacts and PQ artifacts.
Files
Base: base.fbin
HNSW index: index_m_32_ef_500
Query: orig_query_100k.fbin
Ground truth: groundtruth.bin
PQ artifacts: pq_m50.pqcodes/.pqmeta/.pqcodebook… See the full description on the dataset page: https://huggingface.co/datasets/Odeinjul/tti-100m-static-search-eval.turing-100m-static-search-eval
turing-100m-static-search-eval
Static 100M search-evaluation dataset with original vectors, original query/ground-truth files, a rebuilt hnswlib HNSW index, and generated PQ artifacts.
Files
Base: base.fbin
HNSW index: index_m_32_ef_500
Query: orig_query_100k.fbin
Ground truth: groundtruth.bin
Checksums: checksums.sha256
Kaggle metadata is provided via dataset-metadata.json during upload and is stored by Kaggle as dataset configuration, not as a listed data file.… See the full description on the dataset page: https://huggingface.co/datasets/Odeinjul/turing-100m-static-search-eval.deep-10m-static-search-eval
deep-10m-static-search-eval
Static search-evaluation dataset with original vectors, query/ground-truth files, a rebuilt hnswlib HNSW index, and generated PQ artifacts.
Files
Base: base.fbin
HNSW index: index_m_32_ef_500
HNSW layout sidecar: layout-sidecar/index_m_32_ef_500.layout.json`, `layout-sidecar/index_m_32_ef_500.levels.u16`, `layout-sidecar/index_m_32_ef_500.id_mapping.u32
Query: orig_query_10k.fbin
Ground truth: groundtruth.bin
Checksums:… See the full description on the dataset page: https://huggingface.co/datasets/Odeinjul/deep-10m-static-search-eval.Odeasit-latents-ode-heun-1000-class-0_1000-samples-segment-400-499deep-10m-batch-update-eval
deep-10m-batch-update-eval
Batch-update evaluation workload package generated from deep-10m-static-search-eval.
Dataset
Source static dataset: deep-10m-static-search-eval
Vector count: 10,000,000
Dimension: 96
Dtype: float32
Metric: l2
Initial update index: 8,000,000 vectors with external labels equal to A = P[0:8M]
Update order: update_order.u32, a seed-42 permutation of source IDs [0, 10M)
Insert vector source: base_permuted.fbin, where row j equals… See the full description on the dataset page: https://huggingface.co/datasets/Odeinjul/deep-10m-batch-update-eval.ODEN-Indcorpus
ODEN‑Indcorpus 📚
ODEN‑Indcorpus is a 3.7‑million‑line Odia mixed text collection curated from
fiction, dialogue, encyclopaedia, Q‑A and community writing derived from the ODEN initiative.After thorough normalisation and de‑duplication it serves as a robust substrate for training
Odia‑centric tokenizers, language models and embedding spaces.
Split
Lines
Train
3,373,817
Validation
187,434
Test
187,435
Total
3,748,686
The material ranges from conversational… See the full description on the dataset page: https://huggingface.co/datasets/BBSRguy/ODEN-Indcorpus.ode-preprocessing-hy15-testWaveForcing_Stage2_ODE_pair
WaveForcing Stage 2 — Wan2.1-T2V-14B ODE Endpoint Pairs 2K
本数据集包含 2,176 对文本条件视频生成端点数据:2,048 对训练数据及 128 对留出验证数据。每对数据保存文本提示词、初始高斯噪声 z_ref,以及同一提示词和噪声经 Wan2.1-T2V-14B teacher 去噪得到的最终 latent y_ref,可用于视频扩散模型的成对蒸馏或回归研究。
这里的 “ODE pairs” 指 初始噪声到最终去噪 latent 的端点对。数据没有保存 50 步采样过程中的中间状态,不是完整 ODE 轨迹数据集。
数据划分
Split
数量
pair_index
文件名
Train
2,048
0–2047
000000.pt–002047.pt
Validation
128
2048–2175
002048.pt–002175.pt
总计
2,176
0–2175
2,176 个 .pt 文件
划分由… See the full description on the dataset page: https://huggingface.co/datasets/Osc7/WaveForcing_Stage2_ODE_pair.so100_graspThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 4,
"total_frames": 1412,
"total_tasks":1,
"total_videos": 8,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:4"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/odellus/so100_grasp.bigann-1b-static-search-eval-pq
bigann-1b-static-search-eval-pq
Product-quantization artifacts for the BIGANN 1B static search-evaluation dataset. File names follow the Kaggle static-search-eval upload convention: pq_m<M>.pqcodes, pq_m<M>.pqmeta, and pq_m<M>.pqcodebook.
Files
pq_m8.pqcodes
pq_m8.pqmeta
pq_m8.pqcodebook
pq_m16.pqcodes
pq_m16.pqmeta
pq_m16.pqcodebook
pq_m32.pqcodes
pq_m32.pqmeta
pq_m32.pqcodebook
pq_manifest.json
source_manifest.json
Source Data
Base vectors:… See the full description on the dataset page: https://huggingface.co/datasets/Odeinjul/bigann-1b-static-search-eval-pq.sit-latents-ode-heun-segment-0-99so100_test_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 1,
"total_frames": 557,
"total_tasks":1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/odellus/so100_test_2.turing-100m-static-search-eval-lfs-backup-20260720-0825
turing-100m-static-search-eval
Static 100M search-evaluation dataset with original vectors, original query/ground-truth files, a rebuilt hnswlib HNSW index, and generated PQ artifacts.
Files
Base: base.fbin
HNSW index: index_m_32_ef_500
Query: orig_query_100k.fbin
Ground truth: groundtruth.bin
Checksums: checksums.sha256
Kaggle metadata is provided via dataset-metadata.json during upload and is stored by Kaggle as dataset configuration, not as a listed data file.… See the full description on the dataset page: https://huggingface.co/datasets/Odeinjul/turing-100m-static-search-eval-lfs-backup-20260720-0825.tti-100m-batch-update-eval
tti-100m-batch-update-eval
Batch-update evaluation workload package generated from tti-100m-static-search-eval.
Dataset
Source static dataset: tti-100m-static-search-eval
Vector count: 100,000,000
Dimension: 200
Dtype: float32
Metric: ip
Initial update index: 80,000,000 vectors with external labels equal to A = P[0:80M]
Update order: update_order.u32, a seed-42 permutation of source IDs [0, 100M)
Traces
insert-20: starts from A, then inserts… See the full description on the dataset page: https://huggingface.co/datasets/Odeinjul/tti-100m-batch-update-eval.so100_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 2,
"total_frames": 1117,
"total_tasks":1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/odellus/so100_test.odesia-preds2odexODEX dataset annotated with the ground-truth library documentation, to enable evaluations for retrieval and retrieval-augmented code generation.
Please refer to [code-rag-bench] for more details.
odesia-2025-preds1
