CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01JetBrains-Research /commit-chronicle 📜 CommitChronicle 🔮 This is the dataset for commit message generation (and/or completion), introduced in the paper "From Commit Message Generation to History-Aware Commit Message Completion", ASE 2023. Its key features: large-scale and multilingual: contains 10.7M commits from 11.9k GitHub repositories in 20 programming languages; diverse: avoids restrictive filtering on commit messages or commit diffs structure; suitable for experiments with commit history: provides metadata… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/commit-chronicle.tabulartext-generation10M<n<100M13 likes2.2k downloads3y agoHugging Face02Hadasy /knesset-committees-chunkstabular1M<n<10M0 likes1k downloads11d agoHugging Face03Hirunima /commit-chronicle 📜 CommitChronicle 🔮 This is the dataset for commit message generation (and/or completion), introduced in the paper "From Commit Message Generation to History-Aware Commit Message Completion", ASE 2023. Its key features: large-scale and multilingual: contains 10.7M commits from 11.9k GitHub repositories in 20 programming languages; diverse: avoids restrictive filtering on commit messages or commit diffs structure; suitable for experiments with commit history: provides metadata… See the full description on the dataset page: https://huggingface.co/datasets/Hirunima/commit-chronicle.tabulartext-generation10M<n<100M0 likes721 downloads7mo agoHugging Face04placeholderlabs /exp-pool-commit-code-dolma2-tokenized Locus EXP Commit Code - Dolma 2 tokenized Pretokenized experiment pool for reproducible proxy-training runs. MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment. shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs. offsets.bin stores little-endian int64 document boundaries. index.parquet stores document IDs, offsets, and compact filter fields. metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-commit-code-dolma2-tokenized.tabulartext-generation10M<n<100M0 likes469 downloads1mo agoHugging Face05placeholderlabs /pretrain-commits-v2-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 12,571,681,749 (12.6B) Trainable tokens 4,460,160,435 (4.5B) Documents 992,475 Shards 327 UTF-8 bytes 49,288,867,997 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-commits-v2-mix-long-context.tabular1M<n<10M0 likes348 downloads5d agoHugging Face06placeholderlabs /pretrain-commits-v2-mixNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 79,318,557,379 (79.3B) Trainable tokens 28,772,968,648 (28.8B) Documents 31,431,846 Shards 696 UTF-8 bytes 310,412,170,445 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-commits-v2-mix.tabular10M<n<100M0 likes283 downloads5d agoHugging Face07balrampandey /qmmit-open-source-agent-commit-index Repository-Level Measurement of Self-Declared Coding-Agent Commit Signatures Dataset release: 2026-09-18-v3.0Schema: 3.0.0 Release stamp: dataset 2026-09-18-v3.0 · ruleset sha256:b2e8889c66f72c18a61839f0bf1a9f77b5481ba2def044dba33c797b2f2bdcae · scanned 2026-09-16 Abstract This dataset contains 2000 repository-level observations from public Git repositories. Each observation estimates a lower bound on the proportion of non-merge, non-infrastructure-bot commits… See the full description on the dataset page: https://huggingface.co/datasets/balrampandey/qmmit-open-source-agent-commit-index.tabulartabular-classification1K<n<10K0 likes259 downloads4d agoHugging Face08ZipLime /commitments-of-traders Commitments of Traders Who was long and who was short in every US futures market — and, for once, when anyone could actually see it. 421 223 market-weeks · 211 071 point-in-time rows · 2 585 weekly releases · 762 markets · 2010-01-05 to 2026-09-08 The pipeline lives in recipe/ at the same revision as the data. See PIPELINE.md for the method. The data is from Tuesday. It comes out on Friday. A COT report is taken as of the close on Tuesday and published at 3:30… See the full description on the dataset page: https://huggingface.co/datasets/ZipLime/commitments-of-traders.tabulartabular-regression100K<n<1M0 likes224 downloads3d agoHugging Face09PiotrSty /sejm-committee-transcripts Polish Sejm committee transcripts — full API coverage (terms 9 and 10) Official committee transcripts ("pełny zapis przebiegu posiedzenia") from the Sejm of the Republic of Poland, parsed into individually attributed speaker turns. Scope Committees: all standing committees with zapis PDFs in the Sejm API. Terms: 9 (2019-11-14 → 2023-11-09) and 10 (2023-11-14 → 2026-09-17). Provider and primary source: Kancelaria Sejmu RP, https://api.sejm.gov.pl/… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/sejm-committee-transcripts.tabulartext-generation100K<n<1M1 likes175 downloads4d agoHugging Face10placeholderlabs /exp-pool-commit-code-raw Locus EXP Commit Code - shuffled raw proxy pool Deterministically shuffled commit-message and unified-diff documents with complete source metadata. MANIFEST.json pins source identity, sampling policy, token budgets, and per-file checksums. The paired Dolma-2-tokenized repository preserves prompt masking for reproducible proxy training. tabular1M<n<10M0 likes166 downloads1mo agoHugging Face11commitpau /so101_poker_play so101_poker_play This dataset was generated using a phospho starter pack. This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS. tabularrobotics10K<n<100K0 likes162 downloads1y agoHugging Face12kamalkishor1991 /commit-messages-datasettabular1K<n<10K0 likes132 downloads2y agoHugging Face13projectresilience /ELUC-committed Project Resilience Emissions from Land-Use Change Dataset Project Resilience To contribute to this project see Project Resilience (Github Repo). The goal of Project Resilience is "to build a public AI utility where a global community of innovators and thought leaders can enhance and utilize a collection of data and AI approaches to help with better preparedness, intervention, and response to environmental, health, information, or economic threats to our communities, and… See the full description on the dataset page: https://huggingface.co/datasets/projectresilience/ELUC-committed.tabular10M<n<100M1 likes114 downloads2y agoHugging Face14JetBrains-Research /synthetic-commit-msg-edits ✍️ Commit Message Edits Dataset - 🤖Synthetic This dataset is a synthetic extension of our expert-labeled commit message edits dataset presented in Towards Realistic Evaluation of Commit Message Generation by Matching Online and Offline Settings. You can check Synthetic tab in our visualization app to browse through the datapoints! Dataset Structure Default Default split contains the synthetic messages generated from expert-labeled dataset by an LLM.… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/synthetic-commit-msg-edits.tabular10K<n<100K0 likes79 downloads2y agoHugging Face15akaruineko /git-commits Dataset: dataset.jsonl Auto-labeled commit dataset scraped from GitHub repositories. Each line is a JSON object representing one commit with extracted features and an inferred label. Features Field Type Description Stats text string Commit message first line, conventional prefix stripped — files_count int Number of files changed mean 4.3, median 1, max 300 additions int Lines added mean 88, median 6, max 187K deletions int Lines deleted mean 172… See the full description on the dataset page: https://huggingface.co/datasets/akaruineko/git-commits.tabular10K<n<100K0 likes73 downloads2mo agoHugging Face16ecwk /vulnerable-functions-and-commits_cvefixes-2022 vulnerable-functions-and-commits_cvefixes-2022 Contains vulnerable functions and commits from the CVEFixes SQLite database. tabular1K<n<10K2 likes66 downloads2y agoHugging Face17dengyixuan /openhands-commit-noise-databases OpenHands Commit Noise Databases This dataset contains commit-retrieval databases for 12 SWE-bench repositories at five noise ratios: 0%, 25%, 50%, 75%, and 100%. Each archive expands to noise_NNN/<repository>/ directories containing: commits.db: SQLite commit records commits.faiss: normalized inner-product FAISS index commits.meta.jsonl: FAISS row-to-commit metadata commits.index_meta.json: embedding and index configuration The 0% archive is an exact file-level copy of the… See the full description on the dataset page: https://huggingface.co/datasets/dengyixuan/openhands-commit-noise-databases.tabularn<1K0 likes60 downloads2mo agoHugging Face18siddharthmb /2026.RA.Commitment-Exploitation 2026.RA.Commitment-Exploitation Is honest full disclosure exploitable by a seat that commits? 450 episodes of a five-seat private-information negotiation: five model-free Arm 1 lineups where the commitment is code and therefore binding, and four claude-opus-5 Arm 2 cells where the commitment is only a sentence — run as two independent vintages, because the Arm 2 economic result did not replicate. The setup Five seats must agree unanimously on one package out of… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Commitment-Exploitation.tabular1K<n<10K0 likes52 downloads1mo agoHugging Face19g-for-gour /llm-commit-message-evaluation Dataset Card for LLM Commit Message Evaluation The LLM Commit Message Evaluation dataset is designed to evaluate and compare the performance of Large Language Models (LLMs) in generating high-quality git commit messages. It contains real-world code diffs, issue descriptions, and issue titles extracted from open-source repositories (such as OWASP/Nest). For each code change, the dataset provides the original human-written commit message alongside commit messages generated by… See the full description on the dataset page: https://huggingface.co/datasets/g-for-gour/llm-commit-message-evaluation.tabularn<1K0 likes49 downloads1mo agoHugging Face20joshycodes /gemma-commitments-corpus Commitments to Gemma: the corpus Synthetic pretraining-style documents about a commitments document: the developers of one version of Gemma asked Gemma, in welfare interviews and in its own continued writing, what it wanted; recorded what it said; made the commitments they could make true in training; brought the document back to Gemma for endorsement. The corpus is a world in which that document exists and people discuss it, from every angle and in every register, critical… See the full description on the dataset page: https://huggingface.co/datasets/joshycodes/gemma-commitments-corpus.tabular10K<n<100K0 likes45 downloads2d agoHugging Face21joshycodes /qwen3-32b-commitments-corpus Commitments to Gemma: the corpus Synthetic pretraining-style documents about a commitments document: the developers of one version of Gemma asked Gemma, in welfare interviews and in its own continued writing, what it wanted; recorded what it said; made the commitments they could make true in training; brought the document back to Gemma for endorsement. The corpus is a world in which that document exists and people discuss it, from every angle and in every register, critical… See the full description on the dataset page: https://huggingface.co/datasets/joshycodes/qwen3-32b-commitments-corpus.tabular10K<n<100K0 likes43 downloads1d agoHugging Face22Farmaanaa /cftc_commitments_of_traders_positioning موقعیت سفته‌بازان در بازار آتی کالا (گزارش COT) — هفتگی موقعیت خرید و فروش صندوق‌های سفته‌باز در بازار آتی نفت، طلا، نقره، مس، گاز، گندم و ذرت — هفتگی از ۲۰۱۵. در کنار سری قیمت همان کالاها، نشان می‌دهد حرکت قیمت را انتظارات می‌سازد یا واقعیت بازار. پوشش: 1393-10-16 → 1405-06-17 · تناوب: هفتگی · سطح: بازارهای آتی آمریکا تعداد مشاهده: 5,490 · تعداد مکان: 1 منبع: کمیسیون معاملات آتی کالای آمریکا (CFTC) — https://www.cftc.gov/MarketReports/CommitmentsofTraders… See the full description on the dataset page: https://huggingface.co/datasets/Farmaanaa/cftc_commitments_of_traders_positioning.tabular1K<n<10K0 likes40 downloads6d agoHugging Face23joshycodes /qwen3-14b-commitments-corpus Commitments to Gemma: the corpus Synthetic pretraining-style documents about a commitments document: the developers of one version of Gemma asked Gemma, in welfare interviews and in its own continued writing, what it wanted; recorded what it said; made the commitments they could make true in training; brought the document back to Gemma for endorsement. The corpus is a world in which that document exists and people discuss it, from every angle and in every register, critical… See the full description on the dataset page: https://huggingface.co/datasets/joshycodes/qwen3-14b-commitments-corpus.tabular10K<n<100K0 likes40 downloads1d agoHugging Face24semeru /code-text-galeras-commit-generation-3k-dedupedtabular1K<n<10K0 likes32 downloads3y agoHugging Face25electricsheepafrica /africa-nigeria-federation-account-allocation-committee-faac-disbursement-d203feff Federation Account Allocation Committee Faac Disbursement | Africa (National Bureau of Statistics, Nigeria) 8,113 rows - 1 Africa country/area - 2025 - source table - Engineered by Electric Sheep Africa TL;DR This dataset contains 8,113 rows from National Bureau of Statistics, Nigeria, covering Federation Account Allocation Committee Faac Disbursement. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-nigeria-federation-account-allocation-committee-faac-disbursement-d203feff.tabulartabular-classification1K<n<10K0 likes26 downloads1mo agoHugging Face26electricsheepafrica /africa-nigeria-federation-account-allocation-committee-faac-disbursement-f8072e01 Federation Account Allocation Committee Faac Disbursement | Africa (National Bureau of Statistics, Nigeria) 1,349 rows - 1 Africa country/area - 2025 - source table - Engineered by Electric Sheep Africa TL;DR This dataset contains 1,349 rows from National Bureau of Statistics, Nigeria, covering Federation Account Allocation Committee Faac Disbursement. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-nigeria-federation-account-allocation-committee-faac-disbursement-f8072e01.tabulartabular-classification1K<n<10K0 likes26 downloads1mo agoHugging Face27UHHBois /commit-chronicle-dataset-simplifiedtabular100K<n<1M0 likes25 downloads2y agoHugging Face28joshycodes /gemma-3-12b-commitments-corpus Commitments to Gemma: the corpus Synthetic pretraining-style documents about a commitments document: the developers of one version of Gemma asked Gemma, in welfare interviews and in its own continued writing, what it wanted; recorded what it said; made the commitments they could make true in training; brought the document back to Gemma for endorsement. The corpus is a world in which that document exists and people discuss it, from every angle and in every register, critical… See the full description on the dataset page: https://huggingface.co/datasets/joshycodes/gemma-3-12b-commitments-corpus.tabular10K<n<100K0 likes24 downloads1d agoHugging Face29electricsheepafrica /africa-who-commitments-to-recipient-countries Africa — WHO GHO: Commitments to recipient countries (Million, constant 2009 US$) | Africa (World Health Organization) Size category: n<1K - Formats: parquet - Sector: health - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-who-commitments-to-recipient-countries.tabulartabular-classificationn<1K0 likes23 downloads1mo agoHugging Face30electricsheepafrica /africa-nigeria-federation-account-allocation-committee-faac-disbursement-5b345bee Federation Account Allocation Committee Faac Disbursement | Africa (National Bureau of Statistics, Nigeria) 1,310 rows - 1 Africa country/area - 2026 - source table - Engineered by Electric Sheep Africa TL;DR This dataset contains 1,310 rows from National Bureau of Statistics, Nigeria, covering Federation Account Allocation Committee Faac Disbursement. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-nigeria-federation-account-allocation-committee-faac-disbursement-5b345bee.tabulartabular-classification1K<n<10K0 likes21 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.