CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01songlab /ldsc S-LDSC The dataset displayed here (test.parquet) represents the ~10M variants used for S-LDSC in hg38 coordinates (a tiny fraction we couldn't liftover are marked with pos = -1). We include scores for the 3 GPN-Star models (mutation-rate adjusted minus entropy, higher -> more functional). tabular1M<n<10M0 likes1k downloads1mo agoHugging Face02EleutherAI /LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1007 Retrain bank: WikiText-2 / GPT-2, random halves, seed 1007 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents) of the 4,656-document WikiText-2 training set from EleutherAI/bergson-wikitext-2-4656-chunks, following the recipe of Bae et al. 2024, Training Data Attribution via Approximate Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1007.tabular10K<n<100K0 likes878 downloads23d agoHugging Face03EleutherAI /LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1004 Retrain bank: WikiText-2 / GPT-2, random halves, seed 1004 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents) of the 4,656-document WikiText-2 training set from EleutherAI/bergson-wikitext-2-4656-chunks, following the recipe of Bae et al. 2024, Training Data Attribution via Approximate Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1004.tabular10K<n<100K0 likes849 downloads23d agoHugging Face04EleutherAI /LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1006 Retrain bank: WikiText-2 / GPT-2, random halves, seed 1006 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents) of the 4,656-document WikiText-2 training set from EleutherAI/bergson-wikitext-2-4656-chunks, following the recipe of Bae et al. 2024, Training Data Attribution via Approximate Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1006.tabular10K<n<100K0 likes841 downloads23d agoHugging Face05EleutherAI /LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1008 Retrain bank: WikiText-2 / GPT-2, random halves, seed 1008 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents) of the 4,656-document WikiText-2 training set from EleutherAI/bergson-wikitext-2-4656-chunks, following the recipe of Bae et al. 2024, Training Data Attribution via Approximate Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1008.tabular10K<n<100K0 likes825 downloads23d agoHugging Face06EleutherAI /LDS-retrain-bank-adamw-N4k-bs256 Retrain bank: plan_adam_eps1e17_4k_bs256 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on the same 4,000-document corpus with a different random 1% (40 documents) held out, from the same seed and the same data order as the base model in retrained/base. Retraining is deterministic within one environment, so the models differ only by the documents removed. That is the expensive part of any leave-k-out… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-N4k-bs256.tabular1K<n<10K0 likes753 downloads28d agoHugging Face07EleutherAI /LDS-retrain-bank-muon-N16k-bs16tabular1K<n<10K0 likes751 downloads27d agoHugging Face08EleutherAI /LDS-retrain-bank-muon-N16k-bs128tabular1K<n<10K0 likes671 downloads1mo agoHugging Face09EleutherAI /LDS-retrain-bank-adamw-N16k-bs256 Retrain bank: sm_adamw_eps1e17_16k_bs256 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on the same 16,000-document corpus with a different random 1% (160 documents) held out, from the same seed and the same data order as the base model in retrained/base. Retraining is deterministic within one environment, so the models differ only by the documents removed. That is the expensive part of any leave-k-out… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-N16k-bs256.tabular1K<n<10K0 likes625 downloads28d agoHugging Face10EleutherAI /LDS-retrain-bank-muon-N32k-bs256tabular1K<n<10K0 likes616 downloads27d agoHugging Face11EleutherAI /LDS-retrain-bank-adamw-N16k-bs256-scale0.25 Retrain bank: plan_adam_eps1e17_16k_scale0.25 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on the same 16,000-document corpus with a different random 1% (160 documents) held out, from the same seed and the same data order as the base model in retrained/base. Retraining is deterministic within one environment, so the models differ only by the documents removed. That is the expensive part of any leave-k-out… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-N16k-bs256-scale0.25.tabular1K<n<10K0 likes604 downloads28d agoHugging Face12EleutherAI /PARTIAL_LDS-retrain-bank-muon-N64k-bs2560 likes537 downloads20d agoHugging Face13ghanaopenai /fante-speech-text-multispeaker_lds Fante Speech-Text Multispeaker Dataset (LDS) Sentence-level aligned Fante (fat) speech dataset sourced from the Church of Jesus Christ of Latter-day Saints General Conference translations. Dataset Statistics Split Clips Hours Talks Train 29,992 58.32 405 Eval 2,028 4.09 28 Total 32,020 62.41 433 Features audio: 16 kHz mono FLAC sentence-level clips text: Fante transcript (sentence-aligned) talk_id: Source conference talk… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/fante-speech-text-multispeaker_lds.audioautomatic-speech-recognition10K<n<100K1 likes531 downloads2mo agoHugging Face14EleutherAI /LDS-retrain-bank-muon-N16k-bs256tabular1K<n<10K0 likes527 downloads29d agoHugging Face15EleutherAI /LDS-retrain-bank-adamw-N16k-bs256-ep4tabular1K<n<10K0 likes468 downloads27d agoHugging Face16EleutherAI /LDS-retrain-bank-adamw-N32k-bs256tabular1K<n<10K0 likes445 downloads27d agoHugging Face17EleutherAI /LDS-retrain-bank-adamw-N16k-bs256-gpt2-mediumtabular1K<n<10K0 likes445 downloads27d agoHugging Face18EleutherAI /LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1005 Retrain bank: WikiText-2 / GPT-2, random halves, seed 1005 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents) of the 4,656-document WikiText-2 training set from EleutherAI/bergson-wikitext-2-4656-chunks, following the recipe of Bae et al. 2024, Training Data Attribution via Approximate Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1005.tabular10K<n<100K0 likes418 downloads23d agoHugging Face19EleutherAI /PARTIAL_LDS-retrain-bank-gpt2medium-16k-bs32tabular1K<n<10K0 likes400 downloads20d agoHugging Face20EleutherAI /PARTIAL_LDS-retrain-bank-london16k-bs256-adamwtabular1K<n<10K0 likes394 downloads20d agoHugging Face21EleutherAI /LDS-retrain-bank-muon-N8k-bs256 Retrain bank: plan_muon_eps1e17_8k_bs256 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on the same 8,000-document corpus with a different random 1% (80 documents) held out, from the same seed and the same data order as the base model in retrained/base. Retraining is deterministic within one environment, so the models differ only by the documents removed. That is the expensive part of any leave-k-out… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-muon-N8k-bs256.tabular1K<n<10K0 likes362 downloads28d agoHugging Face22EleutherAI /LDS-retrain-bank-adamw-N8k-bs256 Retrain bank: plan_adam_eps1e17_8k_bs256 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on the same 8,000-document corpus with a different random 1% (80 documents) held out, from the same seed and the same data order as the base model in retrained/base. Retraining is deterministic within one environment, so the models differ only by the documents removed. That is the expensive part of any leave-k-out… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-N8k-bs256.tabular1K<n<10K0 likes312 downloads28d agoHugging Face23EleutherAI /LDS-retrain-bank-muon-N4k-bs256 Retrain bank: plan_muon_eps1e17_4k_bs256 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on the same 4,000-document corpus with a different random 1% (40 documents) held out, from the same seed and the same data order as the base model in retrained/base. Retraining is deterministic within one environment, so the models differ only by the documents removed. That is the expensive part of any leave-k-out… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-muon-N4k-bs256.tabular1K<n<10K0 likes292 downloads28d agoHugging Face24EleutherAI /LDS-retrain-bank-adamw-N16k-bs64tabular1K<n<10K0 likes278 downloads1mo agoHugging Face25EleutherAI /LDS-retrain-bank-adamw-N64k-bs256tabular1K<n<10K0 likes269 downloads24d agoHugging Face26EleutherAI /LDS-retrain-bank-adamw-N16k-bs512tabular1K<n<10K0 likes227 downloads29d agoHugging Face27EleutherAI /LDS-retrain-bank-london16k-bs256-muontabular1K<n<10K0 likes186 downloads20d agoHugging Face28EleutherAI /LDS-retrain-bank-adamw-N16k-bs256-clip1.0tabular1K<n<10K0 likes127 downloads29d agoHugging Face29ghananlpcommunity /fante-speech-text-multispeaker_lds Fante Speech-Text Multispeaker Dataset (LDS) Sentence-level aligned Fante (fat) speech dataset sourced from the Church of Jesus Christ of Latter-day Saints General Conference translations. Dataset Statistics Split Clips Hours Talks Train 29,992 58.32 405 Eval 2,028 4.09 28 Total 32,020 62.41 433 Features audio: 16 kHz mono FLAC sentence-level clips text: Fante transcript (sentence-aligned) talk_id: Source conference talk… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/fante-speech-text-multispeaker_lds.audioautomatic-speech-recognition10K<n<100K0 likes125 downloads2mo agoHugging Face30WeiDai-David /Office-Home-LDS Office-Home-LDS Dataset Paper: “Geometric Knowledge-Guided Localized Global Distribution Alignment for Federated Learning” Github: 2025CVPR_GGEUR The Office-Home-LDS dataset is constructed by introducing label skew on top of the domain skew present in the Office-Home dataset. The goal is to create a more challenging and realistic dataset that simultaneously exhibits both label skew and domain skew. 🔗 Citation The article has been accepted by 2025CVPR, if you use… See the full description on the dataset page: https://huggingface.co/datasets/WeiDai-David/Office-Home-LDS.imageimage-classification2 likes104 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.