CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01EleutherAI /LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1007 Retrain bank: WikiText-2 / GPT-2, random halves, seed 1007 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents) of the 4,656-document WikiText-2 training set from EleutherAI/bergson-wikitext-2-4656-chunks, following the recipe of Bae et al. 2024, Training Data Attribution via Approximate Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1007.tabular10K<n<100K0 likes890 downloads25d agoHugging Face02EleutherAI /LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1004 Retrain bank: WikiText-2 / GPT-2, random halves, seed 1004 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents) of the 4,656-document WikiText-2 training set from EleutherAI/bergson-wikitext-2-4656-chunks, following the recipe of Bae et al. 2024, Training Data Attribution via Approximate Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1004.tabular10K<n<100K0 likes854 downloads25d agoHugging Face03EleutherAI /LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1006 Retrain bank: WikiText-2 / GPT-2, random halves, seed 1006 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents) of the 4,656-document WikiText-2 training set from EleutherAI/bergson-wikitext-2-4656-chunks, following the recipe of Bae et al. 2024, Training Data Attribution via Approximate Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1006.tabular10K<n<100K0 likes846 downloads25d agoHugging Face04EleutherAI /LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1008 Retrain bank: WikiText-2 / GPT-2, random halves, seed 1008 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents) of the 4,656-document WikiText-2 training set from EleutherAI/bergson-wikitext-2-4656-chunks, following the recipe of Bae et al. 2024, Training Data Attribution via Approximate Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1008.tabular10K<n<100K0 likes839 downloads25d agoHugging Face05EleutherAI /LDS-retrain-bank-adamw-N4k-bs256 Retrain bank: plan_adam_eps1e17_4k_bs256 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on the same 4,000-document corpus with a different random 1% (40 documents) held out, from the same seed and the same data order as the base model in retrained/base. Retraining is deterministic within one environment, so the models differ only by the documents removed. That is the expensive part of any leave-k-out… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-N4k-bs256.tabular1K<n<10K0 likes723 downloads1mo agoHugging Face06adamsmike /P23977tabular10M<n<100M0 likes632 downloads1y agoHugging Face07EleutherAI /LDS-retrain-bank-adamw-N16k-bs256 Retrain bank: sm_adamw_eps1e17_16k_bs256 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on the same 16,000-document corpus with a different random 1% (160 documents) held out, from the same seed and the same data order as the base model in retrained/base. Retraining is deterministic within one environment, so the models differ only by the documents removed. That is the expensive part of any leave-k-out… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-N16k-bs256.tabular1K<n<10K0 likes629 downloads1mo agoHugging Face08EleutherAI /LDS-retrain-bank-adamw-N16k-bs256-scale0.25 Retrain bank: plan_adam_eps1e17_16k_scale0.25 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on the same 16,000-document corpus with a different random 1% (160 documents) held out, from the same seed and the same data order as the base model in retrained/base. Retraining is deterministic within one environment, so the models differ only by the documents removed. That is the expensive part of any leave-k-out… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-N16k-bs256-scale0.25.tabular1K<n<10K0 likes608 downloads1mo agoHugging Face09EleutherAI /LDS-retrain-bank-adamw-N16k-bs256-ep4tabular1K<n<10K0 likes473 downloads29d agoHugging Face10EleutherAI /LDS-retrain-bank-adamw-N32k-bs256tabular1K<n<10K0 likes451 downloads29d agoHugging Face11EleutherAI /LDS-retrain-bank-adamw-N16k-bs256-gpt2-mediumtabular1K<n<10K0 likes447 downloads29d agoHugging Face12EleutherAI /LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1005 Retrain bank: WikiText-2 / GPT-2, random halves, seed 1005 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents) of the 4,656-document WikiText-2 training set from EleutherAI/bergson-wikitext-2-4656-chunks, following the recipe of Bae et al. 2024, Training Data Attribution via Approximate Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1005.tabular10K<n<100K0 likes422 downloads25d agoHugging Face13EleutherAI /PARTIAL_LDS-retrain-bank-london16k-bs256-adamwtabular1K<n<10K0 likes403 downloads21d agoHugging Face14EleutherAI /LDS-retrain-bank-adamw-N8k-bs256 Retrain bank: plan_adam_eps1e17_8k_bs256 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on the same 8,000-document corpus with a different random 1% (80 documents) held out, from the same seed and the same data order as the base model in retrained/base. Retraining is deterministic within one environment, so the models differ only by the documents removed. That is the expensive part of any leave-k-out… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-N8k-bs256.tabular1K<n<10K0 likes327 downloads1mo agoHugging Face15EleutherAI /LDS-retrain-bank-adamw-N64k-bs256tabular1K<n<10K0 likes276 downloads26d agoHugging Face16EleutherAI /LDS-retrain-bank-adamw-N16k-bs64tabular1K<n<10K0 likes255 downloads1mo agoHugging Face17EleutherAI /LDS-retrain-bank-adamw-N16k-bs512tabular1K<n<10K0 likes227 downloads1mo agoHugging Face18EleutherAI /LDS-retrain-bank-adamw-N16k-bs256-clip1.0tabular1K<n<10K0 likes127 downloads1mo agoHugging Face19adamm-hf /levelbot-datatabular100K<n<1M2 likes80 downloads1y agoHugging Face20EleutherAI /LDS-retrain-bank-adamw-N16k-bs16tabular1K<n<10K0 likes73 downloads1mo agoHugging Face21EleutherAI /LDS-retrain-bank-adamw-N16k-bs256-wd0.1tabular1K<n<10K0 likes72 downloads1mo agoHugging Face22AdamAtractor /btc-l2-orderbook 📊 Hyperliquid Bitcoin (BTC) Level 2 Orderbook Depth - Free Sample 🚀 GET THE FULL DATASET: > You are currently viewing a 7-day free sample. Stop wasting weeks on data engineering. Get the complete, institutional-grade dataset featuring 12+ months of continuous history across 24 crypto assets (~2.3 Million rows) directly at 👉 ImbalanceLabs.com 🛑 Stop Trading on "Liquidity Illusions" Standard OHLCV candles hide the true market intent, spread, spoofing walls, and… See the full description on the dataset page: https://huggingface.co/datasets/AdamAtractor/btc-l2-orderbook.tabulartime-series-forecasting1K<n<10K1 likes64 downloads6mo agoHugging Face23EleutherAI /LDS-retrain-bank-adamw-N16k-bs256-wd0.0tabular1K<n<10K0 likes53 downloads1mo agoHugging Face24Adam159 /data_jobs 🧠 data_jobs Dataset A dataset of real-world data analytics job postings from 2023, collected and processed by Luke Barousse. Background I've been collecting data on data job postings since 2022. I've been using a bot to scrape the data from Google, which come from a variety of sources. You can find the full dataset at my app datanerd.tech. Serpapi has kindly supported my work by providing me access to their API. Tell them I sent you and get 20% off paid plans.… See the full description on the dataset page: https://huggingface.co/datasets/Adam159/data_jobs.tabular100K<n<1M0 likes42 downloads5mo agoHugging Face25AdamJovine /LISTEN-benchmark LISTEN Multi-Objective Selection Benchmark Datasets accompanying the paper "LISTEN to Your Preferences: An LLM Framework for Multi-Objective Selection" (IJCAI-ECAI 2026). 📄 Paper: https://huggingface.co/papers/2510.25799 💻 Code: https://github.com/AdamJovine/LISTEN Overview LISTEN is a benchmark for evaluating how well an LLM can elicit a user's preferences and select a top option from a large candidate set with many numeric and categorical attributes. The… See the full description on the dataset page: https://huggingface.co/datasets/AdamJovine/LISTEN-benchmark.tabularother1K<n<10K0 likes41 downloads4mo agoHugging Face26jason1966 /adamvakar_irish-rent-prices-2020-2025-rtb-official-data Irish Rent Prices 2020-2025 (RTB Official Data) Average monthly rent across 26 Irish counties - ML ready dataset Dataset Info Source: Kaggle Original Size: 0.92 MB Kaggle Downloads: 687 Files: 3 Files irish_rent_by_county.csv irish_rent_full.csv irish_rent_specific.csv Mirrored from Kaggle tabular10K<n<100K0 likes36 downloads6mo agoHugging Face27AdamLucek /legal-rag-positives-synthetic Synthetic QnA Chunk Pairs from Legal Documents This dataset contains excerpts from legal cases' court opinions that mention artificial intelligence, along with corresponding question-answer pairs derived from the content. The data was sourced from CourtListener's public API and processed to create a structured dataset suitable for question-answering tasks. Specifically including cases: Senetas Corporation, Ltd. v. DeepRadiology Corporation Electronic Privacy Information Center v.… See the full description on the dataset page: https://huggingface.co/datasets/AdamLucek/legal-rag-positives-synthetic.tabular1K<n<10K4 likes35 downloads2y agoHugging Face28EleutherAI /LDS-retrain-bank-adamw-N16k-bs128tabular1K<n<10K0 likes30 downloads1mo agoHugging Face29EleutherAI /LDS-retrain-bank-adamw-N16k-bs32tabular1K<n<10K0 likes27 downloads1mo agoHugging Face30AdamLucek /twittersentiment-llama-3.1-405B-labelsFiltered and processed subset of mteb/tweet_sentiment_extraction First 5000 entries were gathered for the train subset, and then 5001-6000 for test. Blanks were removed from this subset, and further filtered to remove innapropriate content via Llama 3.1 405B's inherent harmful/explicit content flagging during the below label processing. This results in a split ofTrain: 4992Test: 998 Original labels have been kept, and further labels have been generated using Llama 3.1 405B, via the prompt:… See the full description on the dataset page: https://huggingface.co/datasets/AdamLucek/twittersentiment-llama-3.1-405B-labels.tabular1K<n<10K1 likes25 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.