CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sentence-transformers /msmarco-distilbert-margin-mse-mean-dot-v1 MS MARCO with hard negatives from distilbert-margin-mse-mean-dot-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-mean-dot-v1.tabularfeature-extraction10M<n<100M2 likes3.3k downloads2y agoHugging Face02alea-institute /kl3m-data-dotgov-stats.bls.gov KL3M Data Project Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper. Description This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models. Dataset Details Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-stats.bls.gov.text10K<n<100K0 likes2.3k downloads1y agoHugging Face03sentence-transformers /msmarco-distilbert-margin-mse-cls-dot-v1 MS MARCO with hard negatives from distilbert-margin-mse-cls-dot-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-cls-dot-v1.tabularfeature-extraction10M<n<100M0 likes1.7k downloads2y agoHugging Face04wolframko /betty-dota2 Betty Dota 2 — Decision Context Dataset Overview 9,385 professional Dota 2 matches parsed from replay files (.dem) into a rich, per-second decision context: hero states, ability cooldowns, building HP, combat events, modifiers, ward placements, and objectives. Built to train Transformer and RL models that understand the game state at each moment in time. Dataset Structure matches.parquet — 9,385 rows One row per match. Match metadata, STRATZ player… See the full description on the dataset page: https://huggingface.co/datasets/wolframko/betty-dota2.tabulartabular-classification1B<n<10B1 likes1.1k downloads6mo agoHugging Face05alea-institute /kl3m-data-dotgov-www.fsis.usda.gov KL3M Data Project Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper. Description This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models. Dataset Details Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.fsis.usda.gov.text1K<n<10K0 likes1.1k downloads1y agoHugging Face06nickbett /bybit-linear-perps-dotusdttabular100M<n<1B0 likes703 downloads27d agoHugging Face07Nottybro /csharp-dotnet-cpt-v0 csharp-dotnet-cpt-v0 A reproducible C# / .NET continued-pretraining (CPT) corpus for Qwen2.5-Coder-1.5B. HF: https://huggingface.co/datasets/Nottybro/csharp-dotnet-cpt-v0 (private) Total unique tokens: 695,288,168 (Qwen2.5-Coder-1.5B tokenizer) Files: 681,426 | Repositories: 68,869 Format: Zstandard-compressed Parquet, schema below. See DATASET_CARD.md for sources, license policy, filtering, limitations. See reports/summary.md for full statistics. tabular10K<n<100K0 likes508 downloads2mo agoHugging Face08HichTala /dota-backgroundimage100K<n<1M0 likes470 downloads9mo agoHugging Face09lsr42 /msmarco-psgs-distilbert-dot-v5text1M<n<10M0 likes467 downloads2y agoHugging Face10alea-institute /dotgov-podcast-sampleaudion<1K0 likes424 downloads2y agoHugging Face11alea-institute /kl3m-data-dotgov-www.cdc.gov KL3M Data Project Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper. Description This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models. Dataset Details Format:… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.cdc.gov.text100K<n<1M0 likes414 downloads1y agoHugging Face12alea-institute /kl3m-data-dotgov KL3M Data Project Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper. Description This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models. Dataset Details Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov.text1M<n<10M0 likes360 downloads1y agoHugging Face13finedet /dota FineDOTA — DOTA v1.0 / v1.5 in the unified detection format Source: the Ultralytics-hosted DOTA release archives (github.com/ultralytics/assets, DOTAv1 / DOTAv1.5 / DOTAv2), which bundle the original DOTA-format annotations (labels/*_original) used for this conversion. Images are the original resolution, re-encoded by upstream from PNG to JPEG. Converted by the finedet project into a unified, AutoTrain-compatible layout: image / width / height / objects{bbox, category} with… See the full description on the dataset page: https://huggingface.co/datasets/finedet/dota.imageobject-detection1K<n<10K0 likes352 downloads1mo agoHugging Face14alea-institute /kl3m-data-dotgov-www.ftc.gov KL3M Data Project Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper. Description This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models. Dataset Details Format:… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.ftc.gov.text10K<n<100K0 likes323 downloads1y agoHugging Face15kotlarmilos /dotnet-runtime .NET Runtime Fine-Tuning Data and Index This directory contains data for fine-tuning models and building RAGs for the dotnet/runtime repository. Overview data/: Contains all datasets and indexes. raw/sample/: Sample PRs and diffs collected from GitHub. raw_data.tar: Archive of collected PRs and diffs from GitHub. samples/: Json files with processed samples suitable for dataset generation. processed/: Parquet files for fine-tuning (e.g., train.parquet, test.parquet).… See the full description on the dataset page: https://huggingface.co/datasets/kotlarmilos/dotnet-runtime.tabulartext-classification100K<n<1M2 likes302 downloads1y agoHugging Face16fge-auto /dotting-test Dotting Test Dotting is a glyph-level benchmark for Turkish text in AI-generated images. It tests whether image models preserve the dotless ı and other Turkish diacritics at the pixel level. This package is generated from the Dotting project outputs for fge-auto/dotting-test. Creator: Fırat Gelbal. Released under Creative Commons Attribution 4.0 International (CC BY 4.0). Attribution should credit Fırat Gelbal and Dotting Test. Related Links Project site:… See the full description on the dataset page: https://huggingface.co/datasets/fge-auto/dotting-test.image10K<n<100K0 likes262 downloads3mo agoHugging Face17alea-institute /kl3m-data-dotgov-www.cia.gov KL3M Data Project Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper. Description This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models. Dataset Details Format:… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.cia.gov.text10K<n<100K0 likes261 downloads1y agoHugging Face18sentence-transformers /msmarco-distilbert-margin-mse-cls-dot-v2 MS MARCO with hard negatives from distilbert-margin-mse-cls-dot-v2 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-cls-dot-v2.tabularfeature-extraction10M<n<100M2 likes231 downloads2y agoHugging Face19alea-institute /kl3m-data-dotgov-www.osha.gov KL3M Data Project Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper. Description This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models. Dataset Details Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.osha.gov.text10K<n<100K0 likes220 downloads1y agoHugging Face20alea-institute /kl3m-data-dotgov-www.federalreserve.gov KL3M Data Project Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper. Description This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models. Dataset Details Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.federalreserve.gov.text100K<n<1M0 likes201 downloads1y agoHugging Face21alea-institute /kl3m-data-dotgov-www.usbr.gov KL3M Data Project Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper. Description This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models. Dataset Details Format:… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.usbr.gov.text10K<n<100K0 likes173 downloads1y agoHugging Face22alea-institute /kl3m-data-dotgov-www.irs.gov KL3M Data Project Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper. Description This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models. Dataset Details Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.irs.gov.text100K<n<1M0 likes168 downloads1y agoHugging Face23i-dot-ai /govuk-policy-qa-pairsThis is a dataset of synthetically generated question and answer pairs on UK government policy papers. It comes in 2 parts: Plain text UK government policy papers, scraped from the Gov.uk Policy papers and consultations page. These are in results.json A series of question and answer pairs on chunk of the above documents, generated using llama_index.finetuning.generate_qa_embedding_pairs and OpenAI GPT3.5 Turbo. text10K<n<100K2 likes162 downloads2y agoHugging Face24llha-dot /my-train-test-datatext1M<n<10M0 likes151 downloads5d agoHugging Face25dotan1111 /MSA-nuc-9-seq Multiple Sequence Alignment as a Sequence-to-Sequence Learning Problem Abstract: The sequence alignment problem is one of the most fundamental problems in bioinformatics and a plethora of methods were devised to tackle it. Here we introduce BetaAlign, a methodology for aligning sequences using an NLP approach. BetaAlign accounts for the possible variability of the evolutionary process among different datasets by using an ensemble of transformers, each trained on millions… See the full description on the dataset page: https://huggingface.co/datasets/dotan1111/MSA-nuc-9-seq.text1M<n<10M0 likes149 downloads3y agoHugging Face26thummd /dot-Generic-100k dot-Generic-100k A frozen evaluation suite from DoTime. Episodes: 100000 Schema: parquet shards + manifest.json (md5-checksummed), Croissant metadata. Load with: from dotime.benchmarks import load_benchmark suite = load_benchmark("dot-Generic-100k") # pulls this repo at tag v1.0.0 Generated reproducibly by scripts/build_release.py. Zenodo DOI is the citable archive of record. tabular100K<n<1M0 likes131 downloads3mo agoHugging Face27HichTala /dota DOTA: Resized and Hugging Face-Ready Vision Dataset This dataset is a restructured version of the DOTA (Dataset for Object Detection in Aerial Images), specifically designed to simplify object detection workflows. By resizing the original images and converting them to the COCO format, this project provides an easier way to use DOTA with popular computer vision frameworks. Additionally, the dataset is formatted for seamless integration with Hugging Face datasets, unlocking new… See the full description on the dataset page: https://huggingface.co/datasets/HichTala/dota.imageobject-detection10K<n<100K2 likes127 downloads1y agoHugging Face28ccwatson /libero_spatial_next_object_target_dot_alpha04This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "panda", "total_episodes": 432, "total_frames": 52970, "total_tasks": 10, "total_videos": 0, "total_chunks": 1, "chunks_size": 1000, "fps": 10, "splits": { "train": "0:432" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ccwatson/libero_spatial_next_object_target_dot_alpha04.imagerobotics10K<n<100K0 likes125 downloads5mo agoHugging Face29dotan1111 /MSA-amino-9-seq Multiple Sequence Alignment as a Sequence-to-Sequence Learning Problem Abstract: The sequence alignment problem is one of the most fundamental problems in bioinformatics and a plethora of methods were devised to tackle it. Here we introduce BetaAlign, a methodology for aligning sequences using an NLP approach. BetaAlign accounts for the possible variability of the evolutionary process among different datasets by using an ensemble of transformers, each trained on millions… See the full description on the dataset page: https://huggingface.co/datasets/dotan1111/MSA-amino-9-seq.text1M<n<10M1 likes122 downloads3y agoHugging Face30hsanchezp /us-dot-flight-delays-2015 US DOT Flight Delays — 2015 (Parquet Version) This dataset contains 5,819,079 records of commercial flights in the United States during 2015.It has been converted to Parquet for efficient analytics in Python, DuckDB, Ibis, Spark, and Polars. Files included flights.parquet (main fact table) airlines.parquet (carrier info) airports.parquet (airport geolocation & metadata) cancellation_codes.parquet (mapping table) Source U.S. Department of Transportation —… See the full description on the dataset page: https://huggingface.co/datasets/hsanchezp/us-dot-flight-delays-2015.text1M<n<10M0 likes122 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.