datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
msmarco-distilbert-margin-mse-mean-dot-v1
MS MARCO with hard negatives from distilbert-margin-mse-mean-dot-v1
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-mean-dot-v1.DOTAv1.0kl3m-data-dotgov-stats.bls.gov
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-stats.bls.gov.msmarco-distilbert-margin-mse-cls-dot-v1
MS MARCO with hard negatives from distilbert-margin-mse-cls-dot-v1
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-cls-dot-v1.dot-distance-area
Dot Distance / Area over Rich Backgrounds
Cross-image spatial-aggregation data used in "Stateful Visual Encoders for
Vision-Language Models" (the Cross-image Spatial Aggregation task). A red dot
is overlaid on each of 2–5 screenshots (AgentNet backgrounds, downsampled to
384×216), and the model estimates a normalized geometric quantity across the
images. Four sub-tasks:
Sub-task dir
Images / example
Quantity
dot_distance/
2
normalized Euclidean distance… See the full description on the dataset page: https://huggingface.co/datasets/zwcolin/dot-distance-area.kl3m-data-dotgov-www.fsis.usda.gov
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.fsis.usda.gov.EnterpriseRAG-Bench
EnterpriseRAG-Bench
A benchmark dataset of 500,000+ documents and 500 questions for evaluating RAG systems on realistic enterprise data.
See the latest leaderboard rankings. The paper is available on arXiv. For code, methodology, and evaluation tools, see the GitHub repository.
Overview
Existing RAG and IR datasets focus on publicly accessible document sets (Bing searches, Stack Overflow, etc.). EnterpriseRAG-Bench provides the first publicly accessible dataset… See the full description on the dataset page: https://huggingface.co/datasets/onyx-dot-app/EnterpriseRAG-Bench.betty-dota2
Betty Dota 2 — Decision Context Dataset
Overview
9,385 professional Dota 2 matches parsed from replay files (.dem) into a rich, per-second decision context: hero states, ability cooldowns, building HP, combat events, modifiers, ward placements, and objectives.
Built to train Transformer and RL models that understand the game state at each moment in time.
Dataset Structure
matches.parquet — 9,385 rows
One row per match. Match metadata, STRATZ player… See the full description on the dataset page: https://huggingface.co/datasets/wolframko/betty-dota2.dotaDOTA dataset collection.
Image Source and Usage License
The DOTA images are collected from the Google Earth, GF-2 and JL-1 satellite provided by the China Centre for Resources Satellite Data and Application, and aerial images provided by CycloMedia B.V. DOTA consists of RGB images and grayscale images. The RGB images are from Google Earth and CycloMedia, while the grayscale images are from the panchromatic band of GF-2 and JL-1 satellite images. All the images are stored in 'png'… See the full description on the dataset page: https://huggingface.co/datasets/isaaccorley/dota.bybit-linear-perps-dotusdtmsmarco-psgs-distilbert-dot-v5dots.mocr-mlx-evals
dots.mocr-mlx-evals
olmOCR-bench evals of dots.mocr MLX quants.
There is directory per model with the md files extracted from the original PDFs.
Directory logs contains the output from the harness run.
DOTAv2DOTA v2 Dataset with OBB, specifically the version from the Ultralytics docs
Website
Full License
Here reproduced from the website webpage
License for Academic Non-Commercial Use Only
This DOTA dataset is made available under the following terms:
The Google Earth images in this dataset are subject to Google Earth's terms of use, which must be adhered to.
The GF-2 and JL-1 satellite images are provided by the China Centre for Resources Satellite Data and Application. The… See the full description on the dataset page: https://huggingface.co/datasets/satyamshorrf/DOTAv2.csharp-dotnet-cpt-v0
csharp-dotnet-cpt-v0
A reproducible C# / .NET continued-pretraining (CPT) corpus for Qwen2.5-Coder-1.5B.
HF: https://huggingface.co/datasets/Nottybro/csharp-dotnet-cpt-v0 (private)
Total unique tokens: 695,288,168 (Qwen2.5-Coder-1.5B tokenizer)
Files: 681,426 | Repositories: 68,869
Format: Zstandard-compressed Parquet, schema below.
See DATASET_CARD.md for sources, license policy, filtering, limitations.
See reports/summary.md for full statistics.
kl3m-data-dotgov-www.cdc.gov
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format:… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.cdc.gov.dot-airline-ontime
DOT Airline On-Time Performance 2018–2024
45,968,068 reported flight records across all 84 months, with all 109 BTS source fields, exact delay values and separate cancellation/diversion outcomes.
This repository contains the 1,000-row public sample, covering all 84 months.
The full package is a one-time $99 snapshot, with monthly CSV and Parquet files.
It has 111 Parquet columns: 109 BTS fields plus source_month and source_row.
The deterministic sample uses 12 evenly spaced rows… See the full description on the dataset page: https://huggingface.co/datasets/claritystorm/dot-airline-ontime.kl3m-data-dotgov
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov.kl3m-data-dotgov-www.ftc.gov
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format:… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.ftc.gov.dotnet-runtime
.NET Runtime Fine-Tuning Data and Index
This directory contains data for fine-tuning models and building RAGs for the dotnet/runtime repository.
Overview
data/: Contains all datasets and indexes.
raw/sample/: Sample PRs and diffs collected from GitHub.
raw_data.tar: Archive of collected PRs and diffs from GitHub.
samples/: Json files with processed samples suitable for dataset generation.
processed/: Parquet files for fine-tuning (e.g., train.parquet, test.parquet).… See the full description on the dataset page: https://huggingface.co/datasets/kotlarmilos/dotnet-runtime.dotting-test
Dotting Test
Dotting is a glyph-level benchmark for Turkish text in AI-generated images. It tests whether image
models preserve the dotless ı and other Turkish diacritics at the pixel level.
This package is generated from the Dotting project outputs for fge-auto/dotting-test.
Creator: Fırat Gelbal.
Released under Creative Commons Attribution 4.0 International (CC BY 4.0).
Attribution should credit Fırat Gelbal and Dotting Test.
Related Links
Project site:… See the full description on the dataset page: https://huggingface.co/datasets/fge-auto/dotting-test.kl3m-data-dotgov-www.cia.gov
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format:… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.cia.gov.godot-codemsmarco-distilbert-margin-mse-cls-dot-v2
MS MARCO with hard negatives from distilbert-margin-mse-cls-dot-v2
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-cls-dot-v2.kl3m-data-dotgov-www.federalreserve.gov
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.federalreserve.gov.harbor-tb2-resultsdota_v1.5my-train-test-datakl3m-data-dotgov-www.irs.gov
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.irs.gov.govuk-policy-qa-pairsThis is a dataset of synthetically generated question and answer pairs on UK government policy papers.
It comes in 2 parts:
Plain text UK government policy papers, scraped from the Gov.uk Policy papers and consultations page. These are in results.json
A series of question and answer pairs on chunk of the above documents, generated using llama_index.finetuning.generate_qa_embedding_pairs and OpenAI GPT3.5 Turbo.
MSA-nuc-9-seq
Multiple Sequence Alignment as a Sequence-to-Sequence Learning Problem
Abstract:
The sequence alignment problem is one of the most fundamental problems in bioinformatics and a plethora of methods were devised to tackle it. Here we introduce BetaAlign, a methodology for aligning sequences using an NLP approach. BetaAlign accounts for the possible variability of the evolutionary process among different datasets by using an ensemble of transformers, each trained on millions… See the full description on the dataset page: https://huggingface.co/datasets/dotan1111/MSA-nuc-9-seq.
