datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FineNewsFineNewsTestSampledbpedia-entities-openai-1M1M OpenAI Embeddings -- 1536 dimensions
Created: June 2023.
Text used for Embedding: title (string) + text (string)
Embedding Model: text-embedding-ada-002
First used for the pgvector vs VectorDB (Qdrant) benchmark: https://nirantk.com/writing/pgvector-vs-qdrant/
Citation
@dataset{dbpedia-entities-openai-1M,
doi = {10.57967/hf/6768},
url = {https://huggingface.co/datasets/KShivendu/dbpedia-entities-openai-1M},
author = {{Kumar Shivendu} and {Nirant Kasliwal}},
title =… See the full description on the dataset page: https://huggingface.co/datasets/KShivendu/dbpedia-entities-openai-1M.urls
URLs
74,918,894,107 deduplicated, validated URLs, sorted by
SURT key
and split into 2,334 range shards.
As plain text the URLs are 5.8 TiB, averaging 84 characters each. Sorted by
SURT key and delta-encoded they fit in 665.6 GiB — 9.54 bytes per URL, a
8.85× reduction. That is the whole point of the ordering: SURT puts URLs
from the same site next to each other, DELTA_LENGTH_BYTE_ARRAY then stores only
where each row differs from the one above it, and zstd compresses what is… See the full description on the dataset page: https://huggingface.co/datasets/ks46/urls.url-atlas
URL Atlas
257,548,097,528 URLs from 105 web corpora, each kept as
its own separately-loadable config, plus the raw source dumps two of them were
extracted from. 4.35 TiB across 52,244 files.
This is the input side of a URL-compression corpus: every source reduced to
its URL column and nothing else. It is deliberately not deduplicated or
merged — sources are kept intact and overlapping so you can measure what each
one contributes, pick the subset you want, and dedup on your own… See the full description on the dataset page: https://huggingface.co/datasets/ks48/url-atlas.IndicSynth
IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset
A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research
🏆 Outstanding Paper Award, ACL 2025
🧠 Overview
IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/ksmashhero/IndicSynth.hf_ckptplatonic-embeddingsISSAI_KSC_335RS_v_1_1
Dataset Card for "ISSAI_KSC_335RS_v_1_1"
Kazakh Speech Corpus (KSC)
Identifier: SLR102
Summary: A crowdsourced open-source Kazakh speech corpus developed by ISSAI (330 hours)
Category: Speech
License: Attribution 4.0 International (CC BY 4.0)
Downloads (use a mirror closer to you):
ISSAI_KSC_335RS_v1.1_flac.tar.gz [19G] (speech, transcripts and metadata ) Mirrors: [US] [EU] [CN]
About this resource:
A crowdsourced open-source speech corpus for the Kazakh language. The KSC… See the full description on the dataset page: https://huggingface.co/datasets/Shirali/ISSAI_KSC_335RS_v_1_1.dec-pdfsChime4USL-Suspilne
USL-Suspilne
A small-scale Ukrainian Sign Language dataset for text-to-pose research, built from
publicly broadcast news clips on Suspilne Mовлення (Ukrainian public broadcaster).
Each clip pairs a Ukrainian sentence with the corresponding interpreter's signing,
provided as a video clip and as MediaPipe pose sequences.
Layout
usl-suspilne/
├── README.md
├── train.csv # 80% — model training
├── dev.csv # 10% — validation /… See the full description on the dataset page: https://huggingface.co/datasets/KSE-RESEARCH-Group/USL-Suspilne.Stanford_dogs
Context
The Stanford Dogs dataset contains images of 120 breeds of dogs from around the world. This dataset has been built using images and annotation from ImageNet for the task of fine-grained image categorization. It was originally collected for fine-grain image categorization, a challenging problem as certain dog breeds have near identical features or differ in colour and age. I have used only images, so this does not contain any labels .
Content
Number of images:… See the full description on the dataset page: https://huggingface.co/datasets/ksaml/Stanford_dogs.KStack
Dataset Summary
KStack is the largest collection of permissively licensed Kotlin code.
Comparison with The Stack v2
In the table below one can find the comparsion between the Kotlin part of The Stack v2 and KStack:
Files
Repositories
Lines
Tokens
Kotlin in The Stack v2
2M
109,457
162M
1.7B
Kstack
4M
168,902
292M
3.1B
Dataset Creation
Collection procedure
We collected repositories from GitHub with the main language being… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains/KStack.hle-no-img-prompt-completion-formatpu-regress-resultsksponspeechplatonic-all-experimentsKSS_Dataset
Description of the original author
KSS Dataset: Korean Single speaker Speech Dataset
KSS Dataset is designed for the Korean text-to-speech task. It consists of audio files recorded by a professional female voice actoress and their aligned text extracted from my books. As a copyright holder, by courtesy of the publishers, I release this dataset to the public. To my best knowledge, this is the first publicly available speech dataset for Korean.
File Format
Each… See the full description on the dataset page: https://huggingface.co/datasets/Bingsu/KSS_Dataset.xview2-xbd
xView2 / xBD (mirror)
Mirror of the xView2 / xBD building damage assessment dataset (parquet), used to train
kshitijrajsharma/dinov3-damage-assessment.
Source
Mirrored from EVER-Z/torchange_xView2.
Attribution
xBD / xView2 dataset from the xView2 Building Damage Assessment Challenge (Gupta et al., 2019).
Imagery from the Maxar Open Data Program. See https://xview2.org/.
License
CC BY-NC-SA 4.0 (non-commercial, share-alike), following… See the full description on the dataset page: https://huggingface.co/datasets/kshitijrajsharma/xview2-xbd.AOPS_Full_Verified_sfropendalKS-dataset-processedtest
Fly Anipose — Lightning Pose Multiview Dataset
6-camera pose estimation dataset for Drosophila leg keypoints, packaged for use with Lightning Pose.
Dataset Description
Head-fixed flies run on a spherical treadmill while 6 synchronized cameras capture locomotion at 300 Hz. Each frame is labeled with 30 keypoints — 5 joint segments (A–E) on each of 6 legs (left legs L1–L3, right legs R1–R3).
Labels are filtered Anipose predictions, not hand-labeled frames. They were… See the full description on the dataset page: https://huggingface.co/datasets/ksikka/test.mmu-norm-legacy-north
Legacy Survey DR9 North image cutouts — L1 (release v1)
This L1 repository contains 2,191,927 objects matched across the release, in 726 shards (about 700 GB). Each object has 152×152-pixel cutouts at 0.262″ per pixel in three bands.
Band names: DR9 North imaging uses BASS g/r and MzLS z. This repository uses bass-g, bass-r, and mzls-z; the original des-g/des-r/des-z tokens are kept in native_band_tokens.
Schema
image struct: band (3), flux (3×152×152… See the full description on the dataset page: https://huggingface.co/datasets/kshitijd/mmu-norm-legacy-north.mmu-norm-jwst-wht
JWST DAWN weight-map cutouts — L1 supplement (release v1)
This L1 supplement provides per-pixel full weight maps (wht_full) for objects matched across the release in the six MMU JWST fields: CEERS, GDN, GDS, NGDEEP, PRIMER-COSMOS, and PRIMER-UDS. The maps come from DAWN JWST Archive v7 Grizli mosaics and are distributed in 93 chunks across eight mosaics. They add pixel-level uncertainty information to the science cutouts.
Schema (one row per object-filter)… See the full description on the dataset page: https://huggingface.co/datasets/kshitijd/mmu-norm-jwst-wht.KsponSpeechminimax-h3-qkv-ksnr-20260917government-ai-detection
Government AI Text Detection — Results
Completed AI-detection output from the government-ai
pipeline, which measures the prevalence of AI-generated/edited text across
four kinds of US government media, 2000–2026:
source
what
bills
Congressional bill text (as-introduced versions), from govinfo
speeches
Floor speeches + Extensions of Remarks from the Congressional Record
comments
Public comments on regulations.gov (via the Mirrulations mirror)
documents
The… See the full description on the dataset page: https://huggingface.co/datasets/ksasse/government-ai-detection.ahr999-dataset
AHR999 BTC Hoarding Index Dataset
Open, daily-updated AHR999 BTC hoarding index dataset, self-computed from
Binance BTCUSDT daily closes and published as CSV and JSON.
This Hugging Face repository is a mirror. The canonical dataset endpoints are:
Dashboard: https://ahr999.aix4u.com/
GitHub: https://github.com/RuochenLyu/ahr999-dataset
CSV endpoint: https://ahr999.aix4u.com/datasets/ahr999.csv
JSON endpoint: https://ahr999.aix4u.com/datasets/ahr999.json
Kaggle discovery mirror:… See the full description on the dataset page: https://huggingface.co/datasets/kshift/ahr999-dataset.
