datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
msmarco-distilbert-margin-mse-mean-dot-v1
MS MARCO with hard negatives from distilbert-margin-mse-mean-dot-v1
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-mean-dot-v1.kl3m-data-dotgov-stats.bls.gov
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-stats.bls.gov.msmarco-distilbert-margin-mse-cls-dot-v1
MS MARCO with hard negatives from distilbert-margin-mse-cls-dot-v1
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-cls-dot-v1.betty-dota2
Betty Dota 2 — Decision Context Dataset
Overview
9,385 professional Dota 2 matches parsed from replay files (.dem) into a rich, per-second decision context: hero states, ability cooldowns, building HP, combat events, modifiers, ward placements, and objectives.
Built to train Transformer and RL models that understand the game state at each moment in time.
Dataset Structure
matches.parquet — 9,385 rows
One row per match. Match metadata, STRATZ player… See the full description on the dataset page: https://huggingface.co/datasets/wolframko/betty-dota2.kl3m-data-dotgov-www.fsis.usda.gov
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.fsis.usda.gov.bybit-linear-perps-dotusdtcsharp-dotnet-cpt-v0
csharp-dotnet-cpt-v0
A reproducible C# / .NET continued-pretraining (CPT) corpus for Qwen2.5-Coder-1.5B.
HF: https://huggingface.co/datasets/Nottybro/csharp-dotnet-cpt-v0 (private)
Total unique tokens: 695,288,168 (Qwen2.5-Coder-1.5B tokenizer)
Files: 681,426 | Repositories: 68,869
Format: Zstandard-compressed Parquet, schema below.
See DATASET_CARD.md for sources, license policy, filtering, limitations.
See reports/summary.md for full statistics.
dota-backgroundmsmarco-psgs-distilbert-dot-v5dotgov-podcast-samplekl3m-data-dotgov-www.cdc.gov
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format:… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.cdc.gov.kl3m-data-dotgov
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov.dota
FineDOTA — DOTA v1.0 / v1.5 in the unified detection format
Source: the Ultralytics-hosted DOTA release archives (github.com/ultralytics/assets, DOTAv1 / DOTAv1.5 / DOTAv2), which bundle the original DOTA-format annotations (labels/*_original) used for this conversion. Images are the original resolution, re-encoded by upstream from PNG to JPEG.
Converted by the finedet project into a unified, AutoTrain-compatible layout:
image / width / height / objects{bbox, category} with… See the full description on the dataset page: https://huggingface.co/datasets/finedet/dota.kl3m-data-dotgov-www.ftc.gov
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format:… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.ftc.gov.dotnet-runtime
.NET Runtime Fine-Tuning Data and Index
This directory contains data for fine-tuning models and building RAGs for the dotnet/runtime repository.
Overview
data/: Contains all datasets and indexes.
raw/sample/: Sample PRs and diffs collected from GitHub.
raw_data.tar: Archive of collected PRs and diffs from GitHub.
samples/: Json files with processed samples suitable for dataset generation.
processed/: Parquet files for fine-tuning (e.g., train.parquet, test.parquet).… See the full description on the dataset page: https://huggingface.co/datasets/kotlarmilos/dotnet-runtime.dotting-test
Dotting Test
Dotting is a glyph-level benchmark for Turkish text in AI-generated images. It tests whether image
models preserve the dotless ı and other Turkish diacritics at the pixel level.
This package is generated from the Dotting project outputs for fge-auto/dotting-test.
Creator: Fırat Gelbal.
Released under Creative Commons Attribution 4.0 International (CC BY 4.0).
Attribution should credit Fırat Gelbal and Dotting Test.
Related Links
Project site:… See the full description on the dataset page: https://huggingface.co/datasets/fge-auto/dotting-test.kl3m-data-dotgov-www.cia.gov
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format:… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.cia.gov.msmarco-distilbert-margin-mse-cls-dot-v2
MS MARCO with hard negatives from distilbert-margin-mse-cls-dot-v2
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-cls-dot-v2.kl3m-data-dotgov-www.osha.gov
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.osha.gov.kl3m-data-dotgov-www.federalreserve.gov
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.federalreserve.gov.kl3m-data-dotgov-www.usbr.gov
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format:… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.usbr.gov.kl3m-data-dotgov-www.irs.gov
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.irs.gov.govuk-policy-qa-pairsThis is a dataset of synthetically generated question and answer pairs on UK government policy papers.
It comes in 2 parts:
Plain text UK government policy papers, scraped from the Gov.uk Policy papers and consultations page. These are in results.json
A series of question and answer pairs on chunk of the above documents, generated using llama_index.finetuning.generate_qa_embedding_pairs and OpenAI GPT3.5 Turbo.
my-train-test-dataMSA-nuc-9-seq
Multiple Sequence Alignment as a Sequence-to-Sequence Learning Problem
Abstract:
The sequence alignment problem is one of the most fundamental problems in bioinformatics and a plethora of methods were devised to tackle it. Here we introduce BetaAlign, a methodology for aligning sequences using an NLP approach. BetaAlign accounts for the possible variability of the evolutionary process among different datasets by using an ensemble of transformers, each trained on millions… See the full description on the dataset page: https://huggingface.co/datasets/dotan1111/MSA-nuc-9-seq.dot-Generic-100k
dot-Generic-100k
A frozen evaluation suite from DoTime.
Episodes: 100000
Schema: parquet shards + manifest.json (md5-checksummed), Croissant metadata.
Load with:
from dotime.benchmarks import load_benchmark
suite = load_benchmark("dot-Generic-100k") # pulls this repo at tag v1.0.0
Generated reproducibly by scripts/build_release.py. Zenodo DOI is the citable
archive of record.
dota
DOTA: Resized and Hugging Face-Ready Vision Dataset
This dataset is a restructured version of the DOTA (Dataset for Object Detection in Aerial Images), specifically designed to simplify object detection workflows. By resizing the original images and converting them to the COCO format, this project provides an easier way to use DOTA with popular computer vision frameworks. Additionally, the dataset is formatted for seamless integration with Hugging Face datasets, unlocking new… See the full description on the dataset page: https://huggingface.co/datasets/HichTala/dota.libero_spatial_next_object_target_dot_alpha04This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "panda",
"total_episodes": 432,
"total_frames": 52970,
"total_tasks": 10,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:432"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ccwatson/libero_spatial_next_object_target_dot_alpha04.MSA-amino-9-seq
Multiple Sequence Alignment as a Sequence-to-Sequence Learning Problem
Abstract:
The sequence alignment problem is one of the most fundamental problems in bioinformatics and a plethora of methods were devised to tackle it. Here we introduce BetaAlign, a methodology for aligning sequences using an NLP approach. BetaAlign accounts for the possible variability of the evolutionary process among different datasets by using an ensemble of transformers, each trained on millions… See the full description on the dataset page: https://huggingface.co/datasets/dotan1111/MSA-amino-9-seq.us-dot-flight-delays-2015
US DOT Flight Delays — 2015 (Parquet Version)
This dataset contains 5,819,079 records of commercial flights in the United States during 2015.It has been converted to Parquet for efficient analytics in Python, DuckDB, Ibis, Spark, and Polars.
Files included
flights.parquet (main fact table)
airlines.parquet (carrier info)
airports.parquet (airport geolocation & metadata)
cancellation_codes.parquet (mapping table)
Source
U.S. Department of Transportation —… See the full description on the dataset page: https://huggingface.co/datasets/hsanchezp/us-dot-flight-delays-2015.
