ode
Datasets
All datasets matching “ode”hate_speech18These files contain text extracted from Stormfront, a white supremacist forum. A random set of
forums posts have been sampled from several subforums and split into sentences. Those sentences
have been manually labelled as containing hate speech or not, according to certain annotation guidelines.ode_step5bigann-100m-static-search-eval
bigann-100m-static-search-eval
Static 100M search-evaluation dataset with original vectors, original query/ground-truth files, a rebuilt hnswlib HNSW index, and generated PQ artifacts.
Files
Base: base.u8bin
HNSW index: index_m_32_ef_500
Query: orig_query_10k.u8bin
Ground truth: groundtruth.bin
Checksums: checksums.sha256
Kaggle metadata is provided via dataset-metadata.json during upload and is stored by Kaggle as dataset configuration, not as a listed data… See the full description on the dataset page: https://huggingface.co/datasets/Odeinjul/bigann-100m-static-search-eval.ODELIA-Challenge-2025
ODELIA Challenge Dataset
This dataset is part of the ODELIA project, a European Horizon initiative focused on developing privacy-preserving, AI-driven diagnostic tools using swarm learning.
The dataset provided here represents a curated subset of data from the broader ODELIA consortium. It is designed to facilitate the development, benchmarking, and validation of AI algorithms that can operate effectively across a range of heterogeneous clinical settings.
The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/ODELIA-AI/ODELIA-Challenge-2025.deep-100m-batch-update-eval
deep-100m-batch-update-eval
Batch-update evaluation workload package generated from deep-100m-static-search-eval.
Dataset
Source static dataset: deep-100m-static-search-eval
Vector count: 100,000,000
Dimension: 96
Dtype: float32
Metric: l2
Initial update index: 80,000,000 vectors with external labels equal to A = P[0:80M]
Update order: update_order.u32, a seed-42 permutation of source IDs [0, 100M)
Insert vector source: base_permuted.fbin, where row j equals… See the full description on the dataset page: https://huggingface.co/datasets/Odeinjul/deep-100m-batch-update-eval.deep-100m-static-search-eval
deep-100m-static-search-eval
Static 100M search-evaluation dataset with original vectors, original query/ground-truth files, a rebuilt hnswlib HNSW index, and generated PQ artifacts.
Files
Base: base.fbin
HNSW index: index_m_32_ef_500
Query: orig_query_10k.fbin
Ground truth: groundtruth.bin
Checksums: checksums.sha256
Kaggle metadata is provided via dataset-metadata.json during upload and is stored by Kaggle as dataset configuration, not as a listed data file.… See the full description on the dataset page: https://huggingface.co/datasets/Odeinjul/deep-100m-static-search-eval.
patchReviews diffs like a tired but fair maintainer. Will ask why that function exists.
orbitAudits dependencies, auth flows and the things people assume are fine. Files issues, not panic.
rioHappy at both ends of the request. Will not add a third framework.
fluxProfiles first, guesses never. Has deleted more code than it has written.