datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1007
Retrain bank: WikiText-2 / GPT-2, random halves, seed 1007
This repository contains 100 fully retrained language models, not just scores.
Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents)
of the 4,656-document WikiText-2 training set from
EleutherAI/bergson-wikitext-2-4656-chunks,
following the recipe of Bae et al. 2024, Training Data Attribution via Approximate
Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1007.LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1004
Retrain bank: WikiText-2 / GPT-2, random halves, seed 1004
This repository contains 100 fully retrained language models, not just scores.
Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents)
of the 4,656-document WikiText-2 training set from
EleutherAI/bergson-wikitext-2-4656-chunks,
following the recipe of Bae et al. 2024, Training Data Attribution via Approximate
Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1004.LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1006
Retrain bank: WikiText-2 / GPT-2, random halves, seed 1006
This repository contains 100 fully retrained language models, not just scores.
Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents)
of the 4,656-document WikiText-2 training set from
EleutherAI/bergson-wikitext-2-4656-chunks,
following the recipe of Bae et al. 2024, Training Data Attribution via Approximate
Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1006.LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1008
Retrain bank: WikiText-2 / GPT-2, random halves, seed 1008
This repository contains 100 fully retrained language models, not just scores.
Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents)
of the 4,656-document WikiText-2 training set from
EleutherAI/bergson-wikitext-2-4656-chunks,
following the recipe of Bae et al. 2024, Training Data Attribution via Approximate
Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1008.LDS-retrain-bank-adamw-N4k-bs256
Retrain bank: plan_adam_eps1e17_4k_bs256
This repository contains 100 fully retrained language models, not just scores.
Each model is GPT-2 (gpt2) fine-tuned on the same 4,000-document corpus with a different random 1% (40 documents) held out, from the same seed and the same data order as the base model in retrained/base. Retraining is deterministic within one environment, so the models differ only by the documents removed.
That is the expensive part of any leave-k-out… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-N4k-bs256.P23977LDS-retrain-bank-adamw-N16k-bs256
Retrain bank: sm_adamw_eps1e17_16k_bs256
This repository contains 100 fully retrained language models, not just scores.
Each model is GPT-2 (gpt2) fine-tuned on the same 16,000-document corpus with a different random 1% (160 documents) held out, from the same seed and the same data order as the base model in retrained/base. Retraining is deterministic within one environment, so the models differ only by the documents removed.
That is the expensive part of any leave-k-out… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-N16k-bs256.LDS-retrain-bank-adamw-N16k-bs256-scale0.25
Retrain bank: plan_adam_eps1e17_16k_scale0.25
This repository contains 100 fully retrained language models, not just scores.
Each model is GPT-2 (gpt2) fine-tuned on the same 16,000-document corpus with a different random 1% (160 documents) held out, from the same seed and the same data order as the base model in retrained/base. Retraining is deterministic within one environment, so the models differ only by the documents removed.
That is the expensive part of any leave-k-out… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-N16k-bs256-scale0.25.LDS-retrain-bank-adamw-N16k-bs256-ep4LDS-retrain-bank-adamw-N32k-bs256LDS-retrain-bank-adamw-N16k-bs256-gpt2-mediumLDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1005
Retrain bank: WikiText-2 / GPT-2, random halves, seed 1005
This repository contains 100 fully retrained language models, not just scores.
Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents)
of the 4,656-document WikiText-2 training set from
EleutherAI/bergson-wikitext-2-4656-chunks,
following the recipe of Bae et al. 2024, Training Data Attribution via Approximate
Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1005.PARTIAL_LDS-retrain-bank-london16k-bs256-adamwLDS-retrain-bank-adamw-N8k-bs256
Retrain bank: plan_adam_eps1e17_8k_bs256
This repository contains 100 fully retrained language models, not just scores.
Each model is GPT-2 (gpt2) fine-tuned on the same 8,000-document corpus with a different random 1% (80 documents) held out, from the same seed and the same data order as the base model in retrained/base. Retraining is deterministic within one environment, so the models differ only by the documents removed.
That is the expensive part of any leave-k-out… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-N8k-bs256.LDS-retrain-bank-adamw-N64k-bs256LDS-retrain-bank-adamw-N16k-bs64LDS-retrain-bank-adamw-N16k-bs512LDS-retrain-bank-adamw-N16k-bs256-clip1.0levelbot-dataLDS-retrain-bank-adamw-N16k-bs16LDS-retrain-bank-adamw-N16k-bs256-wd0.1btc-l2-orderbook
📊 Hyperliquid Bitcoin (BTC) Level 2 Orderbook Depth - Free Sample
🚀 GET THE FULL DATASET: > You are currently viewing a 7-day free sample.
Stop wasting weeks on data engineering. Get the complete, institutional-grade dataset featuring 12+ months of continuous history across 24 crypto assets (~2.3 Million rows) directly at 👉 ImbalanceLabs.com
🛑 Stop Trading on "Liquidity Illusions"
Standard OHLCV candles hide the true market intent, spread, spoofing walls, and… See the full description on the dataset page: https://huggingface.co/datasets/AdamAtractor/btc-l2-orderbook.LDS-retrain-bank-adamw-N16k-bs256-wd0.0data_jobs
🧠 data_jobs Dataset
A dataset of real-world data analytics job postings from 2023, collected and processed by Luke Barousse.
Background
I've been collecting data on data job postings since 2022. I've been using a bot to scrape the data from Google, which come from a variety of sources.
You can find the full dataset at my app datanerd.tech.
Serpapi has kindly supported my work by providing me access to their API. Tell them I sent you and get 20% off paid plans.… See the full description on the dataset page: https://huggingface.co/datasets/Adam159/data_jobs.LISTEN-benchmark
LISTEN Multi-Objective Selection Benchmark
Datasets accompanying the paper "LISTEN to Your Preferences: An LLM Framework for Multi-Objective Selection" (IJCAI-ECAI 2026).
📄 Paper: https://huggingface.co/papers/2510.25799
💻 Code: https://github.com/AdamJovine/LISTEN
Overview
LISTEN is a benchmark for evaluating how well an LLM can elicit a user's preferences and select a top option from a large candidate set with many numeric and categorical attributes. The… See the full description on the dataset page: https://huggingface.co/datasets/AdamJovine/LISTEN-benchmark.adamvakar_irish-rent-prices-2020-2025-rtb-official-data
Irish Rent Prices 2020-2025 (RTB Official Data)
Average monthly rent across 26 Irish counties - ML ready dataset
Dataset Info
Source: Kaggle
Original Size: 0.92 MB
Kaggle Downloads: 687
Files: 3
Files
irish_rent_by_county.csv
irish_rent_full.csv
irish_rent_specific.csv
Mirrored from Kaggle
legal-rag-positives-synthetic
Synthetic QnA Chunk Pairs from Legal Documents
This dataset contains excerpts from legal cases' court opinions that mention artificial intelligence, along with corresponding question-answer pairs derived from the content. The data was sourced from CourtListener's public API and processed to create a structured dataset suitable for question-answering tasks.
Specifically including cases:
Senetas Corporation, Ltd. v. DeepRadiology Corporation
Electronic Privacy Information Center v.… See the full description on the dataset page: https://huggingface.co/datasets/AdamLucek/legal-rag-positives-synthetic.LDS-retrain-bank-adamw-N16k-bs128LDS-retrain-bank-adamw-N16k-bs32twittersentiment-llama-3.1-405B-labelsFiltered and processed subset of mteb/tweet_sentiment_extraction
First 5000 entries were gathered for the train subset, and then 5001-6000 for test.
Blanks were removed from this subset, and further filtered to remove innapropriate content via Llama 3.1 405B's inherent harmful/explicit content flagging during the below label processing.
This results in a split ofTrain: 4992Test: 998
Original labels have been kept, and further labels have been generated using Llama 3.1 405B, via the prompt:… See the full description on the dataset page: https://huggingface.co/datasets/AdamLucek/twittersentiment-llama-3.1-405B-labels.
