datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SceneText150kffp
FoundationPose paired synthetic renders
Each scene row contains two synchronized views and per-object pose, mask ID,
bounding box, visibility, and occlusion annotations. Depth is stored as the
original float32 NPY bytes; RGB and uint32 instance masks are stored as PNG
bytes. Matrices use row-major flattened arrays and canonical column-vector
names (camera_from_world and world_from_object). Raw source transforms are
retained alongside them.
The assets tables contain stable source… See the full description on the dataset page: https://huggingface.co/datasets/Krispin/ffp.emo_webds_2prepacked-fineweb-edu-llama2-32K-T2048
prepacked-fineweb-edu-llama2-32K-T2048
Pre-tokenized and BOS-aligned best-fit packed version of FineWeb-Edu for training with looped nanochat.
Tokenized with the Llama 2 tokenizer (32,000 base vocab + 8 special tokens = 32,008).
Stats
Train split
Source
karpathy/fineweb-edu-100b-shuffle (1,821 shards)
Total tokens
63.26B
Total docs
97.1M
Rows
30,873,598
Shards
2,059 (train-00000 to train-02058)
Rows per shard
~15,000… See the full description on the dataset page: https://huggingface.co/datasets/KristianS7/prepacked-fineweb-edu-llama2-32K-T2048.emo_parlerkrishna_wallpapersemo_webdsImageNet-Think
ImageNet-Think 250K
ImageNet-Think 250K is a large-scale synthetic multimodal reasoning dataset containing of 250,000 images sampled from ImageNet-21K dataset. For each image, we provide a prompt and two different step-by-step reasoning tokens and outputs (answers), enabling evaluation and training for Vision Language Models on reasoning tasks. This dataset is primarily designed for research on multimodal summarization.
Installation & Setup
Before downloading… See the full description on the dataset page: https://huggingface.co/datasets/krishnateja95/ImageNet-Think.awesome-llm-datasets-only-Chinesesae-code-correctness-dataAgentNet
OpenCUA: Open Foundations for Computer-Use Agents
🌐 Website
🔎 Data Viewer
📝 Paper
💻 Code
AgentNet Dataset
AgentNet is the first large-scale desktop computer-use agent trajectory dataset, containing 22.6K human-annotated computer-use tasks across Windows, macOS, and Ubuntu systems.
Applications
This dataset enables training and evaluation of:
Vision-language-action (VLA) models for computer use
Multi-modal agents… See the full description on the dataset page: https://huggingface.co/datasets/KrishPatel0111/AgentNet.trial_dataset
Trial Dataset (VQA)
This dataset contains various configurations for Visual Question Answering (VQA) tasks involving tables, figures, and multiple-choice options.
Dataset Structure
The dataset is split into multiple configurations based on the complexity of the input (number of tables/figures) and the response type (MCQ or Constructed Response).
How to Load
You can load any specific configuration using the datasets library:
from datasets import load_dataset
#… See the full description on the dataset page: https://huggingface.co/datasets/Krishna5T/trial_dataset.KrishokChat
KrishokChat Dataset
KrishokChat is a provenance-traceable, multi-task Bengali agricultural dataset for safety-critical chemical advisory and domain-specific natural language understanding. Every instance in the dataset retains citation-level provenance (publisher, document title, page range, section path) linked directly to official agricultural extension handbooks and research manuals issued by government and NGO agricultural institutions in Bangladesh.
Figure 1:… See the full description on the dataset page: https://huggingface.co/datasets/RaiyanKhaan/KrishokChat.train_dataset_21_12ocs-krisdika
Open Law Data Thailand: OCS Krisdika Dataset
ชุดข้อมูลกฎหมายจาก สำนักงานคณะกรรมการกฤษฎีกา (Office of the Council of State) รวบรวมและจัดทำโดยโครงการ Open Law Data Thailand เพื่อส่งเสริมการเข้าถึงข้อมูลกฎหมายในรูปแบบที่เครื่องอ่านได้ (Machine-Readable)
Dataset Structure
ข้อมูลถูกจัดเก็บในรูปแบบ JSON Lines (.jsonl) แบ่งไฟล์ตาม ปีและเดือน (YYYY/YYYY-MM.jsonl) เพื่อความสะดวกในการดาวน์โหลดและบริหารจัดการ
Data Fields
แต่ละบรรทัด (Row) ประกอบด้วยข้อมูลดังนี้:
title… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/ocs-krisdika.kriyam-tamperflow
Kriyam TamperFlow
The first document tampering detection benchmark built specifically for Indian documents, with a built-in compression stress-test that exposes how quickly forensic models degrade on real-world scanned material.
Dataset Summary
State-of-the-art document forgery detectors — CAT-Net, DTD, MVSS-Net, CAFTB, and others — rely on JPEG compression artifacts as their primary forensic signal: inconsistencies in DCT coefficients, block boundaries, and… See the full description on the dataset page: https://huggingface.co/datasets/kriyam-ai/kriyam-tamperflow.cosmopedia-10B
Cosmopedia 10B
Dataset Description
This is a 10.53 Billion token subset of the HuggingFaceTB/cosmopedia dataset. It was created by sampling approximately 45% of each subset (web_samples, stories, stanford, etc.) from the original dataset and deduplicating to ensure high utility.
Motivation
The original Cosmopedia dataset is massive (~25B+ tokens) and high quality. This 10B version serves as a "Goldilocks" dataset—large enough for meaningful pre-training… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/cosmopedia-10B.kolors-20kfinemath-4plus-tokenizedSROIErecursion-cellular-image-classification-datasetRedPajama-Data-V2-1B
RedPajama-Data-V2 1B
Dataset Description
This is a 1.01 Billion token subset of the togethercomputer/RedPajama-Data-V2 dataset (specifically derived from the sample-10B config). It was created by randomly sampling the source data.
Motivation
RedPajama V2 is a state-of-the-art web dataset with rich quality signals. This 1B token subset allows for rapid testing of these quality signals or other filtering experiments without needing to process the full… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/RedPajama-Data-V2-1B.Spotify_Music_Analytics_and_Popularity_PredictionRedPajama-10B-Weighted
RedPajama-10B-Weighted
A canonical 10 Billion token weighted subset of the RedPajama-Data-1T dataset.
Dataset Description
This dataset is a faithful reproduction of the original RedPajama-Data-1T distribution, scaled down to exactly 10 Billion tokens. It is designed to preserve the exact domain ratios of the original dataset (excluding the defunct 'Books' subset). This allows researchers and developers to prototype, debug, and test on a representative slice of the data… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/RedPajama-10B-Weighted.MSU-Bench
MSU-Bench: Musical Score Understanding Benchmark
Evaluating Large Language Models' Comprehension of Complete Musical Scores
Overview
MSU-Bench is a human-curated benchmark for evaluating the musical score understanding capabilities of Large Language Models (LLMs) and Vision-Language Models (VLMs). It supports multimodal evaluation through both textual (ABC notation) and visual (PDF/image) inputs.
Key Statistics:
150 complete musical scores
1,800 generative… See the full description on the dataset page: https://huggingface.co/datasets/Krinos/MSU-Bench.cosmopedia-1b
Cosmopedia 1B
Dataset Description
This is a 1 Billion token subset of the krisbailey/cosmopedia-10B dataset, which itself is a 10B subset of HuggingFaceTB/cosmopedia.
It was created by uniformly sampling approximately 9.5% of the 10B dataset, ensuring the data distribution remains consistent with the source.
Motivation
While the 10B dataset is a "Goldilocks" size for many experiments, 1B tokens is the standard size for rapid prototyping, scaling law… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/cosmopedia-1b.generated-finemath-292968-node-2falcon-refinedweb-1B
Falcon RefinedWeb 1B
Dataset Description
This is a 1.01 Billion token subset of the tiiuae/falcon-refinedweb dataset. It was created by streaming the dataset with a large shuffle buffer to ensure a random, representative sample of the web data.
Motivation
RefinedWeb is a high-quality filtered web dataset, but the full version is massive. This 1B token slice provides a perfect testbed for evaluating model architecture changes or for use in curriculum learning… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/falcon-refinedweb-1B.EDGAR-CORPUS-Financial-Summarization
EDGAR-CORPUS : 10K Financial Report Summarization
Extracted from SEC EDGAR filings (1993-2020). This dataset enhances financial report summarization by leveraging a hybrid AI model strategy.
Using:
ChatGPT-3.5 Turbo(~70%),
Claude 3.5 (~30% to generate structured, accurate, and concise summaries)
Dataset Composition
Summaries in this dataset are generated using a hybrid AI model strategy, balancing quality and efficiency:ChatGPT-3.5 Turbo (~70%) – Used for structured… See the full description on the dataset page: https://huggingface.co/datasets/kritsadaK/EDGAR-CORPUS-Financial-Summarization.generated-finemath-585942-node-0
