datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
iclr-wm-backup-public
ICLR Watermark Benchmark — backup overflow (public part)
Companion to the private repo Aak975/iclr-wm-backup, which reached its
storage quota. Together the two repos form ONE backup — every file exists in
exactly one of them, with the same layout:
archives/<sub>/part-0000 ... part-NNNN, MANIFEST.json
restore one archive: cat part-* | zstd -d | tar -x
MANIFEST.json = {"parts": N, "sha256": <whole-stream>, "total_bytes": M}
This public part holds only shareable image data… See the full description on the dataset page: https://huggingface.co/datasets/Aak975/iclr-wm-backup-public.mysdxl-dataset
Image-Prompt Dataset
An image-prompt dataset scraped and assembled with
MySDXL for training
latent diffusion models.
Dataset structure
Each row contains one image with its corresponding text prompt.
Column
Type
Description
image
Image
RGB image (lossless PNG, original resolution)
prompt
string
Text prompt describing the image
negative_prompt
string
Negative prompt (empty string if none)
Stored as Parquet shards (data/train-*.parquet).
Load… See the full description on the dataset page: https://huggingface.co/datasets/aakkaasshh/mysdxl-dataset.vaigai-dataset
Vaigai Dataset
aakkaasshh/vaigai-dataset is an Indic-focused multilingual text corpus aggregated for tokenizer training, built with the goal of reaching tokenizer quality on par with dedicated Indic tokenizers (e.g. Sarvam AI's).
One Parquet file per language is stored at the root of this repo (e.g. bn.parquet), with no nested folders.
Every time new data is fetched for a language, it is merged with that language's existing file and deduplicated on the text column, so re-running… See the full description on the dataset page: https://huggingface.co/datasets/aakkaasshh/vaigai-dataset.sindhi-corpus-505m
Sindhi Corpus 505M
The largest open-source, deduplicated Sindhi language pretraining corpus.
~505 million tokens across 742K documents, covering news, literature, legal, religious, encyclopedic, and web-crawled Sindhi text. Built for training Sindhi language models, tokenizers, and NLP tools.
Dataset Summary
Stat
Value
Documents
~742,379
Tokens (estimated)
~505 million
Language
Sindhi (sd) — Arabic script
Format
Parquet (single text column)… See the full description on the dataset page: https://huggingface.co/datasets/aakashMeghwar01/sindhi-corpus-505m.cdr_bigbio_processedcolpali_train_set
Dataset Description
This dataset is the training set of ColPali it includes 127,460 query-image pairs from both openly available academic datasets (63%) and a synthetic dataset made up
of pages from web-crawled PDF documents and augmented with VLM-generated (Claude-3 Sonnet) pseudo-questions (37%).
Our training set is fully English by design, enabling us to study zero-shot generalization to non-English languages.
Dataset
#examples (query-page pairs)
Language
DocVQA… See the full description on the dataset page: https://huggingface.co/datasets/aakash-projects/colpali_train_set.sindhi-corpus-langid-cleanit-support-llmSindhi-Intelligence-Core-SFT
🧠 Sindhi Intelligence Core SFT
This is a premium, high-density instruction dataset designed for training Large Language Models (LLMs) to master the Sindhi language. With 361,225 rows, it provides a robust foundation for grammar, factual knowledge, and logical reasoning.
📊 Dataset Summary
This dataset was created by consolidating multiple high-quality Sindhi corpora into a unified ChatML format. It is specifically optimized for Supervised Fine-Tuning (SFT).
📁… See the full description on the dataset page: https://huggingface.co/datasets/aakashMeghwar01/Sindhi-Intelligence-Core-SFT.jeesemeval-2025-task-8-test-caseshi-en-noisy-vad-benchmark
Hindi-English Noisy VAD Benchmark
Version 0.1.0 is a deterministic, evaluation-only benchmark with 78
mono PCM16 WAV files at 16 kHz: six clean speech controls and 72 mixtures spanning
six speech sources, three real noise categories, and four SNRs (20, 10, 5, 0 dB).
Intended use
Use this dataset to compare voice-activity detectors under matched Hindi/English
noise conditions and to tune thresholds. It is too small and insufficiently diverse
for model training… See the full description on the dataset page: https://huggingface.co/datasets/Aakash22134/hi-en-noisy-vad-benchmark.semeval-2025-task-8-test-cases-competitionIndian_Railway_maintance
license: cc-by-4.0
task_categories:
- tabular-classification
- tabular-regression
- time-series
language:
- en
tags:
- railway
- predictive-maintenance
- failure-detection
- transportation
- iot
- machine-learning
- synthetic-data
- analytics
pretty_name: Indian Railway Failure Detection & Maintenance (100K)
size_categories:
- 100K<n<1M
🚆 Indian Railway Failure Detection & Maintenance (100K)
Overview
This dataset contains 100,000 synthetic yet… See the full description on the dataset page: https://huggingface.co/datasets/aakanksha180/Indian_Railway_maintance.pico_bigbio_processedresearch-paper-agent-reasoning-traces-unverifiedbiomedical_lectures_eng_v2
Vidore Benchmark 2 - MIT Dataset
This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of MIT courses in anatomy (precisely tissue interactions).
Dataset Summary
Each query is in english.
This dataset provides a focused benchmark for visual retrieval tasks related to MIT biology courses. It includes a curated set of documents, queries, relevance judgments (qrels), and page… See the full description on the dataset page: https://huggingface.co/datasets/aakash-projects/biomedical_lectures_eng_v2.sindhi-corpus-cleanReDepress
ReDepress Dataset
The ReDepress Dataset originates from the paper"ReDepress: A Cognitive Framework for Detecting Depression Relapse from Social Media".
This dataset is designed to support research on detecting depression relapse using social media text.It provides rich temporal user data and cognitive bias annotations, allowing systems to monitor conversational dynamics and make timely inferences.
Important Notice
By requesting, accessing, downloading, or using this… See the full description on the dataset page: https://huggingface.co/datasets/aakash-agarwal/ReDepress.It-support-synthetic-dataTHAR-DatasetThe dataset consists 11,549 YouTube comments in Hindi-English code-mixed language for targeted hate speech detection against religion. Binary and multi-class tagging of YouTube comments is used.
The classification of YouTube comments addresses two subtasks: Subtask-1 (Binary classification): comments are labeled as antireligion or non-antireligion. Subtask-2 (Multi-class classification): comments are labeled on the major targeted religions such as Islam, Hinduism, and Christianity, with a… See the full description on the dataset page: https://huggingface.co/datasets/Aakash941/THAR-Dataset.semeval-2025-task-8-prompts-competitionaurat-march-processed-newsenglish-aurat-march-newsSindhi-Intelligence-Core-SFT-v2commn-voice-11-translatedarxivqa_test_subsampled
Dataset Description
This is a VQA dataset based on figures extracted from arXiv publications taken from ArXiVQA dataset from Multimodal ArXiV. The questions were generated synthetically using GPT-4 Vision.
Data Curation
To ensure homogeneity across our benchmarked datasets, we subsampled the original test set to 500 pairs. Furthermore we renamed the different columns for our purpose.
Load the dataset
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/aakash-projects/arxivqa_test_subsampled.TestingPersian-Wikipedia-Corpus
Overview
This dataset is derived from the Persian Wikipedia Corpus project, which contains parsed articles from the Persian Wikipedia. The original data has been converted into a more accessible format and made available through the HuggingFace datasets library.
Usage
from datasets import load_dataset
dataset = load_dataset("codersan/Persian-Wikipedia-Corpus")
Persian-Wikipedia-Corpus
A complete copy of Persian Wikimedia pages, The dataset… See the full description on the dataset page: https://huggingface.co/datasets/AAkhoram/Persian-Wikipedia-Corpus.semeval-2025-task-8-prompts-train
