CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sentence-transformers /msmarco-distilbert-margin-mse-mean-dot-v1 MS MARCO with hard negatives from distilbert-margin-mse-mean-dot-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-mean-dot-v1.tabularfeature-extraction10M<n<100M2 likes3.4k downloads2y agoHugging Face02Last-Bullet /DOTAv1.0image1K<n<10K0 likes2.8k downloads2y agoHugging Face03alea-institute /kl3m-data-dotgov-stats.bls.gov KL3M Data Project Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper. Description This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models. Dataset Details Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-stats.bls.gov.text10K<n<100K0 likes2.4k downloads1y agoHugging Face04sentence-transformers /msmarco-distilbert-margin-mse-cls-dot-v1 MS MARCO with hard negatives from distilbert-margin-mse-cls-dot-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-cls-dot-v1.tabularfeature-extraction10M<n<100M0 likes1.7k downloads2y agoHugging Face05zwcolin /dot-distance-area Dot Distance / Area over Rich Backgrounds Cross-image spatial-aggregation data used in "Stateful Visual Encoders for Vision-Language Models" (the Cross-image Spatial Aggregation task). A red dot is overlaid on each of 2–5 screenshots (AgentNet backgrounds, downsampled to 384×216), and the model estimates a normalized geometric quantity across the images. Four sub-tasks: Sub-task dir Images / example Quantity dot_distance/ 2 normalized Euclidean distance… See the full description on the dataset page: https://huggingface.co/datasets/zwcolin/dot-distance-area.imageimage-to-text100K<n<1M0 likes1.6k downloads4mo agoHugging Face06alea-institute /kl3m-data-dotgov-www.fsis.usda.gov KL3M Data Project Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper. Description This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models. Dataset Details Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.fsis.usda.gov.text1K<n<10K0 likes1.1k downloads1y agoHugging Face07onyx-dot-app /EnterpriseRAG-Bench EnterpriseRAG-Bench A benchmark dataset of 500,000+ documents and 500 questions for evaluating RAG systems on realistic enterprise data. See the latest leaderboard rankings. The paper is available on arXiv. For code, methodology, and evaluation tools, see the GitHub repository. Overview Existing RAG and IR datasets focus on publicly accessible document sets (Bing searches, Stack Overflow, etc.). EnterpriseRAG-Bench provides the first publicly accessible dataset… See the full description on the dataset page: https://huggingface.co/datasets/onyx-dot-app/EnterpriseRAG-Bench.textquestion-answeringn<1K14 likes1k downloads5mo agoHugging Face08wolframko /betty-dota2 Betty Dota 2 — Decision Context Dataset Overview 9,385 professional Dota 2 matches parsed from replay files (.dem) into a rich, per-second decision context: hero states, ability cooldowns, building HP, combat events, modifiers, ward placements, and objectives. Built to train Transformer and RL models that understand the game state at each moment in time. Dataset Structure matches.parquet — 9,385 rows One row per match. Match metadata, STRATZ player… See the full description on the dataset page: https://huggingface.co/datasets/wolframko/betty-dota2.tabulartabular-classification1B<n<10B1 likes961 downloads6mo agoHugging Face09isaaccorley /dotaDOTA dataset collection. Image Source and Usage License The DOTA images are collected from the Google Earth, GF-2 and JL-1 satellite provided by the China Centre for Resources Satellite Data and Application, and aerial images provided by CycloMedia B.V. DOTA consists of RGB images and grayscale images. The RGB images are from Google Earth and CycloMedia, while the grayscale images are from the panchromatic band of GF-2 and JL-1 satellite images. All the images are stored in 'png'… See the full description on the dataset page: https://huggingface.co/datasets/isaaccorley/dota.text1K<n<10K3 likes889 downloads2y agoHugging Face10nickbett /bybit-linear-perps-dotusdttabular100M<n<1B0 likes703 downloads25d agoHugging Face11lsr42 /msmarco-psgs-distilbert-dot-v5text1M<n<10M0 likes636 downloads2y agoHugging Face12pcuenq /dots.mocr-mlx-evals dots.mocr-mlx-evals olmOCR-bench evals of dots.mocr MLX quants. There is directory per model with the md files extracted from the original PDFs. Directory logs contains the output from the harness run. text10K<n<100K0 likes592 downloads6mo agoHugging Face13satyamshorrf /DOTAv2DOTA v2 Dataset with OBB, specifically the version from the Ultralytics docs Website Full License Here reproduced from the website webpage License for Academic Non-Commercial Use Only This DOTA dataset is made available under the following terms: The Google Earth images in this dataset are subject to Google Earth's terms of use, which must be adhered to. The GF-2 and JL-1 satellite images are provided by the China Centre for Resources Satellite Data and Application. The… See the full description on the dataset page: https://huggingface.co/datasets/satyamshorrf/DOTAv2.image1K<n<10K1 likes522 downloads9mo agoHugging Face14Nottybro /csharp-dotnet-cpt-v0 csharp-dotnet-cpt-v0 A reproducible C# / .NET continued-pretraining (CPT) corpus for Qwen2.5-Coder-1.5B. HF: https://huggingface.co/datasets/Nottybro/csharp-dotnet-cpt-v0 (private) Total unique tokens: 695,288,168 (Qwen2.5-Coder-1.5B tokenizer) Files: 681,426 | Repositories: 68,869 Format: Zstandard-compressed Parquet, schema below. See DATASET_CARD.md for sources, license policy, filtering, limitations. See reports/summary.md for full statistics. tabular10K<n<100K0 likes452 downloads2mo agoHugging Face15alea-institute /kl3m-data-dotgov-www.cdc.gov KL3M Data Project Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper. Description This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models. Dataset Details Format:… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.cdc.gov.text100K<n<1M0 likes437 downloads1y agoHugging Face16claritystorm /dot-airline-ontime DOT Airline On-Time Performance 2018–2024 45,968,068 reported flight records across all 84 months, with all 109 BTS source fields, exact delay values and separate cancellation/diversion outcomes. This repository contains the 1,000-row public sample, covering all 84 months. The full package is a one-time $99 snapshot, with monthly CSV and Parquet files. It has 111 Parquet columns: 109 BTS fields plus source_month and source_row. The deterministic sample uses 12 evenly spaced rows… See the full description on the dataset page: https://huggingface.co/datasets/claritystorm/dot-airline-ontime.tabulartabular-classification1K<n<10K0 likes429 downloads3d agoHugging Face17alea-institute /kl3m-data-dotgov KL3M Data Project Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper. Description This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models. Dataset Details Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov.text1M<n<10M0 likes353 downloads1y agoHugging Face18alea-institute /kl3m-data-dotgov-www.ftc.gov KL3M Data Project Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper. Description This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models. Dataset Details Format:… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.ftc.gov.text10K<n<100K0 likes317 downloads1y agoHugging Face19kotlarmilos /dotnet-runtime .NET Runtime Fine-Tuning Data and Index This directory contains data for fine-tuning models and building RAGs for the dotnet/runtime repository. Overview data/: Contains all datasets and indexes. raw/sample/: Sample PRs and diffs collected from GitHub. raw_data.tar: Archive of collected PRs and diffs from GitHub. samples/: Json files with processed samples suitable for dataset generation. processed/: Parquet files for fine-tuning (e.g., train.parquet, test.parquet).… See the full description on the dataset page: https://huggingface.co/datasets/kotlarmilos/dotnet-runtime.tabulartext-classification100K<n<1M2 likes285 downloads1y agoHugging Face20fge-auto /dotting-test Dotting Test Dotting is a glyph-level benchmark for Turkish text in AI-generated images. It tests whether image models preserve the dotless ı and other Turkish diacritics at the pixel level. This package is generated from the Dotting project outputs for fge-auto/dotting-test. Creator: Fırat Gelbal. Released under Creative Commons Attribution 4.0 International (CC BY 4.0). Attribution should credit Fırat Gelbal and Dotting Test. Related Links Project site:… See the full description on the dataset page: https://huggingface.co/datasets/fge-auto/dotting-test.image10K<n<100K0 likes259 downloads3mo agoHugging Face21alea-institute /kl3m-data-dotgov-www.cia.gov KL3M Data Project Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper. Description This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models. Dataset Details Format:… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.cia.gov.text10K<n<100K0 likes257 downloads1y agoHugging Face22dotfantasy /godot-codetext10K<n<100K8 likes223 downloads3y agoHugging Face23sentence-transformers /msmarco-distilbert-margin-mse-cls-dot-v2 MS MARCO with hard negatives from distilbert-margin-mse-cls-dot-v2 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-cls-dot-v2.tabularfeature-extraction10M<n<100M2 likes217 downloads2y agoHugging Face24alea-institute /kl3m-data-dotgov-www.federalreserve.gov KL3M Data Project Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper. Description This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models. Dataset Details Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.federalreserve.gov.text100K<n<1M0 likes203 downloads1y agoHugging Face25dot-agi /harbor-tb2-resultstext100K<n<1M0 likes191 downloads6mo agoHugging Face26CSSNB /dota_v1.5image100K<n<1M0 likes176 downloads8mo agoHugging Face27llha-dot /my-train-test-datatext1M<n<10M0 likes169 downloads3d agoHugging Face28alea-institute /kl3m-data-dotgov-www.irs.gov KL3M Data Project Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper. Description This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models. Dataset Details Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.irs.gov.text100K<n<1M0 likes154 downloads1y agoHugging Face29i-dot-ai /govuk-policy-qa-pairsThis is a dataset of synthetically generated question and answer pairs on UK government policy papers. It comes in 2 parts: Plain text UK government policy papers, scraped from the Gov.uk Policy papers and consultations page. These are in results.json A series of question and answer pairs on chunk of the above documents, generated using llama_index.finetuning.generate_qa_embedding_pairs and OpenAI GPT3.5 Turbo. text10K<n<100K2 likes152 downloads2y agoHugging Face30dotan1111 /MSA-nuc-9-seq Multiple Sequence Alignment as a Sequence-to-Sequence Learning Problem Abstract: The sequence alignment problem is one of the most fundamental problems in bioinformatics and a plethora of methods were devised to tackle it. Here we introduce BetaAlign, a methodology for aligning sequences using an NLP approach. BetaAlign accounts for the possible variability of the evolutionary process among different datasets by using an ensemble of transformers, each trained on millions… See the full description on the dataset page: https://huggingface.co/datasets/dotan1111/MSA-nuc-9-seq.text1M<n<10M0 likes148 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.