datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tempofunk-sdance
TempoFunk S(mall)Dance
10k samples of metadata and encoded latents & prompts of videos themed around dance.
Data format
Video frame latents
Numpy arrays
120 frames, 512x512 source size
Encoded shape (120, 4, 64, 64)
CLIP (openai) encoded prompts
Video description (as seen in metadata)
Encoded shape (77,768)
Video metadata as JSON (description, tags, categories, source URLs, etc.)
surgenet-train
SurgeNet Training Dataset
Using NWS=20 input to ADCIRC to generate a dual graph dataset of historical TC storms (1980--2024 inclusive) for GNN training. 2 hour timesteps.
@misc{Thomas2025,
author = {Thomas, Simon D. A.},
title = {SurgeNet Training Dataset},
year = {2025},
publisher = {Hugging Face},
doi = {10.57967/hf/6971},
url = {https://huggingface.co/datasets/sdat2/surgenet-train}
}
WebShop-Qwen3-8B-SDAR-evalaig-dashboardThis is a Next.js project bootstrapped with create-next-app.
Getting Started
First, run the development server:
npm run dev
# or
yarn dev
# or
pnpm dev
# or
bun dev
Open http://localhost:3000 with your browser to see the result.
You can start editing the page by modifying app/page.tsx. The page auto-updates as you edit the file.
This project uses next/font to automatically optimize and load Geist, a new font family for Vercel.
Learn More
To learn more about… See the full description on the dataset page: https://huggingface.co/datasets/sdasdawq234/aig-dashboard.tropical-gunshot-sda-dataset
SDA Dataset
Dataset Description
Dataset Summary
The SDA (Spectrogram Data Augmentation) dataset is a specialized audio dataset designed for automated gunshot detection in tropical forest environments. It is an augmented derivative of the publicly available tropical forest gunshot dataset (Katsis et al., 2022), enhanced through spectrogram-level augmentations to address the challenge of rare gunshot events and class imbalance. Developed as part of the research… See the full description on the dataset page: https://huggingface.co/datasets/aazmaine25/tropical-gunshot-sda-dataset.aig-services8tags
8TAGS
Dataset Summary
A Polish topic classification dataset consisting of headlines from social media posts. It contains about 50,000 sentences annotated with 8 topic labels: film, history, food, medicine, motorization, work, sport and technology. This dataset was created automatically by extracting sentences from headlines and short descriptions of articles posted on Polish social networking site wykop.pl. The service allows users to annotate articles with one or more… See the full description on the dataset page: https://huggingface.co/datasets/sdadas/8tags.ppc
PPC - Polish Paraphrase Corpus
Dataset Summary
Polish Paraphrase Corpus contains 7000 manually labeled sentence pairs. The dataset was divided into training, validation and test splits. The training part includes 5000 examples, while the other parts contain 1000 examples each. The main purpose of creating such a dataset was to verify how machine learning models perform in the challenging problem of paraphrase identification, where most records contain semantically… See the full description on the dataset page: https://huggingface.co/datasets/sdadas/ppc.surgenet-test-ph
SurgeNet Test Dataset - Potential Height Simulations - Alpha Version
@misc{simon_thomas_2025,
author = { Simon Thomas },
title = { surgenet-test-ph (Revision 884b698) },
year = 2025,
url = { https://huggingface.co/datasets/sdat2/surgenet-test-ph },
doi = { 10.57967/hf/7006 },
publisher = { Hugging Face }
}
UniKIE
UNIKIE-BENCH
Key Information Extraction (KIE) from real-world documents remains challenging due to substantial variations in layout structures, visual quality, and task-specific information requirements. Recent Large Multimodal Models (LMMs) have shown promising potential for performing end-to-end KIE directly from document images. To enable a comprehensive and systematic evaluation across realistic and diverse application scenarios, we introduce UniKIE-BENCH, a unified… See the full description on the dataset page: https://huggingface.co/datasets/sdasdfeff22222/UniKIE.CoderForge-Preview
CoderForge-Preview: SOTA Open Dataset for Training Efficient Agents
CoderForge-Preview is the largest open test-verified coding agent dataset.
Fine-tuning Qwen-3 32B on it, we boost SWE-Bench Verified performance 23.0% → 59.4% pass@1 and rank #1 among open-data and #2 among open-weight models ≤32B parameters.
Limitations
Adaptability to different scaffolds: We generated all trajectories using a single scaffold and fixed tool set (no permutations). Models trained via… See the full description on the dataset page: https://huggingface.co/datasets/sdadonlie/CoderForge-Preview.ai-gateway
AI Gateway (aig)
Self-hosted, OpenAI-wire-compatible AI gateway that calls model providers directly, with a custom enrichment layer with first-class Iran/Persian support.
Architecture
Customer ──► AI Gateway (FastAPI) ──► model provider ──► OpenRouter / local vLLM / RAG
AI Gateway handles: enrichment pipeline (datetime, calendar, system prompt, tool packs), auth caching, audit logging, Telegram alerts.
the provider handles: provider routing, fallback, budgets… See the full description on the dataset page: https://huggingface.co/datasets/sdasdawq234/ai-gateway.tools-docssd_asr_synthesis_datagrab-black-cube-act-30ep_20260908_205154This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/sdadasdaga/grab-black-cube-act-30ep_20260908_205154.grab-black-cube-30ep-v4_20260827_140205This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/sdadasdaga/grab-black-cube-30ep-v4_20260827_140205.sd_asr_synthesis_data_v0_less_silenceAr-En-Code-Switching-Textual-Dataset
ArE-CSTD: Arabic-English Code-Switching Textual Dataset
The National Center for Artificial Intelligence at the Saudi Data and Artificial Intelligence Authority (SDAIA), published the "ArE-CSTD" dataset, which stands for "Arabic-English Code-Switching Textual Dataset”.
This dataset contains 330K dialectical Arabic-English code-swithing sentences generated by the large language model GPT-4.
TXT Files
There are 6 txt files. 2 files for Modern Standard Arabic(MSA) train and… See the full description on the dataset page: https://huggingface.co/datasets/SDAIANCAI/Ar-En-Code-Switching-Textual-Dataset.corpus-toolssdasdsick_pl
SICK_PL - Sentences Involving Compositional Knowledge (Polish)
Dataset Summary
This dataset is a manually translated version of popular English natural language inference (NLI) corpus consisting of 10,000 sentence pairs. NLI is the task of determining whether one statement (premise) semantically entails other statement (hypothesis). Such relation can be classified as entailment (if the first sentence entails second sentence), neutral (the first statement does not… See the full description on the dataset page: https://huggingface.co/datasets/sdadas/sick_pl.pick-cube-2camThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 30,
"total_frames": 8971,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sdadasdaga/pick-cube-2cam.SD_Anime_Characters_Repositorygrab-black-cube-30ep-v3_20260826_164036This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/sdadasdaga/grab-black-cube-30ep-v3_20260826_164036.grab-black-cube-30ep-v6_20260827_150153This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/sdadasdaga/grab-black-cube-30ep-v6_20260827_150153.juli1pvtmaycee-retail-dataset
Maycee Retail Dataset
Published by SDataPro.
Maycee Retail is a realistic synthetic Australian retail dataset for SQL
learning, data engineering practice, analytics engineering, BI demos, and
LLM/text-to-SQL evaluation.
Dataset Summary
This free Hugging Face release covers 2017-01-01 to 2019-12-31: 3 years of daily
temporal depth (2017-2019). It contains 164,968 transactions,
420,016 line items, 8,762 returns, and full current-state
dimension snapshots as of the… See the full description on the dataset page: https://huggingface.co/datasets/SDataPro/maycee-retail-dataset.rollout_act_chunk50_20260904_113735This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/sdadasdaga/rollout_act_chunk50_20260904_113735.SDAU_wheat_head_detection
Sdau Wheat Head Detection
This dataset features synthetic images of wheat heads captured in a laboratory environment. Images were collected using a handheld Canon EOS 70D RGB camera, providing a controlled setting for wheat head detection research. The dataset contains 3,030 images with 96,629 bounding box annotations across 1 category.
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
The original train/test/val split has been… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/SDAU_wheat_head_detection.
