CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ceselder /loracle-pretrain-v5-qwen14b-tokenstabular10K<n<100K0 likes1.2k downloads2mo agoHugging Face02kothasuhas /dclm_10B_tokenstabular1M<n<10M0 likes480 downloads1y agoHugging Face03BlockDB /ERC20-Tokens-Ethereum-Cryptocurrency-Data ERC20-Tokens-Ethereum-Cryptocurrency-Data Hive-partitioned Parquet export of BlockDB erc20_tokens (Ethereum). Load from datasets import load_dataset ds = load_dataset("BlockDB/ERC20-Tokens-Ethereum-Cryptocurrency-Data", split="train") Files live under data/year=YYYY/month=MM/part-NNNN.parquet. Range: 2023-08 .. 2026-06 (UTC calendar months). Schema column type block_timestamp timestamp block_number int64 tx_index int32 contract_id… See the full description on the dataset page: https://huggingface.co/datasets/BlockDB/ERC20-Tokens-Ethereum-Cryptocurrency-Data.tabular1M<n<10M0 likes372 downloads1mo agoHugging Face04marin-community /open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n8 Open Thoughts 4 - Math (Qwen3-32B, 32K tokens, n=8) This dataset contains math reasoning problems with 8 independent responses generated by Qwen3-32B. Overview Source: marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens Model: Qwen/Qwen3-32B Temperature: 0.8 Max tokens: 32,768 Columns Column Description instruction_seed The math problem prompt _source Source dataset identifier gpt41_mini_response Reference response from… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n8.tabular10K<n<100K1 likes279 downloads8mo agoHugging Face05treadon /speech-dac-tokens-3cb Speech DAC Tokens (3 Codebooks) Pre-tokenized speech dataset using the Descript Audio Codec (DAC). Each audio clip has been encoded into discrete codebook tokens from DAC's first 3 residual vector quantization codebooks, paired with its text transcription. Dataset Summary Stat Value Total samples 241,451 Total audio ~780 hours Language English Codebooks 3 (of DAC's 9) Codebook size 1,024 entries each DAC model 44kHz Tokens per second ~258 (86 frames… See the full description on the dataset page: https://huggingface.co/datasets/treadon/speech-dac-tokens-3cb.tabulartext-to-speech100K<n<1M0 likes262 downloads6mo agoHugging Face06marin-community /open-thoughts-4-30k-code-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8 Open Thoughts 4 - Code (Qwen3-30B-A3B-Thinking-2507, 32K tokens, n=8) This dataset contains code reasoning problems with 8 independent responses generated by Qwen3-30B-A3B-Thinking-2507. Overview Source: marin-community/open-thoughts-4-30k-code-qwen3-32b-annotated (prompts only) Model: Qwen/Qwen3-30B-A3B-Thinking-2507 Temperature: 0.8 Max tokens: 32,768 Columns Column Description instruction_seed The code problem prompt _source Source dataset… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-code-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8.tabular10K<n<100K0 likes247 downloads7mo agoHugging Face07bethrezen /ru-big-russian-dataset-16k-tokens-limittabular1M<n<10M0 likes242 downloads1y agoHugging Face08wmatejuk /midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512-epr-labs midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512-epr-labs Pre-tokenized MIDI pieces for IsoFLOP scaling-law runs. Each row is one full piece (no time-windowing); training crops sequences from packed token bins. The source column is the original piece metadata as JSON so a row can be traced back to its EPR Labs source dataset. Based on MIDI datasets gathered by EPR Labs. Codec name: dyadic tokenizer vocab size: 512 max_time_step: 1.0 n_velocity_bins: 32… See the full description on the dataset page: https://huggingface.co/datasets/wmatejuk/midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512-epr-labs.tabulartext-generation1M<n<10M0 likes232 downloads20d agoHugging Face09marin-community /open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8 Open Thoughts 4 - Math (Qwen3-30B-A3B-Thinking-2507, 32K tokens, n=8) This dataset contains math reasoning problems with 8 independent responses generated by Qwen3-30B-A3B-Thinking-2507. Overview Source: marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated (base prompts) Model: Qwen/Qwen3-30B-A3B-Thinking-2507 Temperature: 0.8 Max tokens: 32,768 Columns Column Description instruction_seed The math problem prompt _source Source dataset… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8.tabular10K<n<100K0 likes220 downloads7mo agoHugging Face10princeton-nlp /QuRatedPajama-1B_tokens_for_analysis QuRatedPajama Paper: QuRating: Selecting High-Quality Data for Training Language Models This dataset is a 1B token subset derived from princeton-nlp/QuRatedPajama-260B, which is a subset of cerebras/SlimPajama-627B annotated by princeton-nlp/QuRater-1.3B with sequence-level quality ratings across 4 criteria: Educational Value - e.g. the text includes clear explanations, step-by-step reasoning, or questions and answers Facts & Trivia - how much factual and trivia knowledge the text… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/QuRatedPajama-1B_tokens_for_analysis.tabular1M<n<10M6 likes208 downloads2y agoHugging Face11marin-community /open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8-reformatted Dataset Card for Open-Thoughts-4-30K-Math-Qwen3-30B-A3B-Thinking-2507-Annotated-32768-Tokens-N8-Reformatted Overview This dataset is a reformatted version of marin-community/open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8. The original dataset contained 29,963 samples, each with 8 responses generated by the same model with different random seeds (stored in generated_text, generated_text2, ..., generated_text8 columns). This reformatted… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8-reformatted.tabular100K<n<1M0 likes205 downloads6mo agoHugging Face12NoahEJ /fineweb-sample-100BT_over-1024-tokenstabular10M<n<100M0 likes203 downloads1y agoHugging Face13marin-community /open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8 Open Thoughts 4 - Math (Qwen3-4B, 32K tokens, n=8) This dataset contains math reasoning problems with 8 independent responses generated by Qwen3-4B. Overview Source: marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens (n=1 version with 1 response per prompt) Model: Qwen/Qwen3-4B Temperature: 0.8 Max tokens: 32,768 Columns Column Description instruction_seed The math problem prompt _source Source dataset identifier… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8.tabular10K<n<100K0 likes193 downloads7mo agoHugging Face14NoahEJ /fineweb-sample-100BT_over-1024-tokens-subset-xsA subset of FineWeb sample-100BT with sequence length >= 1024 when tokenized with the Llama 2 tokenizer (including special tokens) tabular10M<n<100M0 likes190 downloads1y agoHugging Face15NoahEJ /fineweb-sample-100BT_over-4096-tokenstabular1M<n<10M0 likes185 downloads1y agoHugging Face16NoahEJ /fineweb-sample-100BT_over-8192-tokenstabular100K<n<1M0 likes184 downloads1y agoHugging Face17wmatejuk /midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512-aug-lesstabular100K<n<1M0 likes178 downloads1mo agoHugging Face18Self-GRIT /open-hermes-2.5-sft-mixture-llama3-inference-retrieval-tokenstabular1M<n<10M0 likes166 downloads2y agoHugging Face19BlockDB /ERC1155-Tokens-Ethereum-Cryptocurrency-Data ERC1155-Tokens-Ethereum-Cryptocurrency-Data Hive-partitioned Parquet export of BlockDB erc1155_tokens (Ethereum). Load from datasets import load_dataset ds = load_dataset("BlockDB/ERC1155-Tokens-Ethereum-Cryptocurrency-Data", split="train") Files live under data/year=YYYY/month=MM/part-NNNN.parquet. Range: 2015-08 .. 2026-06 (UTC calendar months). Schema column type block_timestamp timestamp block_number int64 tx_index int32… See the full description on the dataset page: https://huggingface.co/datasets/BlockDB/ERC1155-Tokens-Ethereum-Cryptocurrency-Data.tabular100K<n<1M0 likes162 downloads1mo agoHugging Face20marin-community /open-thoughts-4-6865-math-kimi-k2pt5-annotated-32768-tokens-n8-reformatted open-thoughts-4-6865-math-kimi-k2pt5-annotated-32768-tokens Math reasoning responses generated by Kimi K2.5 (moonshotai/Kimi-K2.5) via a Together AI dedicated instance. Overview Total rows: 54,920 Unique prompts: 6,865 (each with 8 response annotations) Source prompts: marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n8-reformatted Generation model: moonshotai/Kimi-K2.5 Max tokens: 32,768 Temperature: 0.8 Tokenizer used for stats: Qwen/Qwen2.5-3B… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-6865-math-kimi-k2pt5-annotated-32768-tokens-n8-reformatted.tabular10K<n<100K0 likes160 downloads6mo agoHugging Face21marin-community /open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens Dataset Card for Open-Thoughts-4-30K-Math-Qwen3-4B-Annotated-32768-Tokens Overview This dataset is a variant of the OpenThoughts-4 30K math subset with responses generated by Qwen/Qwen3-4B using max output tokens = 32768, allowing for longer and more complete chain-of-thought reasoning. Generation Details Model: Qwen/Qwen3-4B Temperature: 0.8 Max Output Tokens: 32768 Dataset Statistics Number of Samples: 29,963 Split: train Dataset… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens.tabular10K<n<100K0 likes159 downloads7mo agoHugging Face22BlockDB /ERC721-Tokens-Ethereum-Cryptocurrency-Data ERC721-Tokens-Ethereum-Cryptocurrency-Data Hive-partitioned Parquet export of BlockDB erc721_tokens (Ethereum). Load from datasets import load_dataset ds = load_dataset("BlockDB/ERC721-Tokens-Ethereum-Cryptocurrency-Data", split="train") Files live under data/year=YYYY/month=MM/part-NNNN.parquet. Range: 2023-08 .. 2026-06 (UTC calendar months). Schema column type block_timestamp timestamp block_number int64 tx_index int32… See the full description on the dataset page: https://huggingface.co/datasets/BlockDB/ERC721-Tokens-Ethereum-Cryptocurrency-Data.tabular10M<n<100M0 likes153 downloads1mo agoHugging Face23marin-community /open-thoughts-4-30k-code-qwen3-32b-annotated-32768-tokens Dataset Card for Open-Thoughts-4-30K-Code-Qwen3-32B-Annotated-32768-Tokens Overview This dataset is a variant of marin-community/open-thoughts-4-30k-code-qwen3-32b-annotated with an extended maximum sequence length. The responses in the generated_text column were generated with max output tokens = 32768 (instead of 7500 in the original dataset), allowing for longer and more complete chain-of-thought reasoning. Generation Details Model: Qwen/Qwen3-32B… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-code-qwen3-32b-annotated-32768-tokens.tabulartext-generation10K<n<100K0 likes147 downloads8mo agoHugging Face24Bradley /fineweb-over-20k-l3b-tokenstabular1M<n<10M0 likes146 downloads1y agoHugging Face25marin-community /open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n8-reformatted-selfconsistency Dataset Card for Open-Thoughts-4-30K-Math-Qwen3-32B-Annotated-32768-Tokens-N8-Reformatted-SelfConsistency Overview This dataset is a self-consistency filtered version of marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n8-reformatted. For each prompt, 8 responses were generated by Qwen3-32B with different random seeds. A majority vote was taken over the final answers (extracted from \boxed{...}) to determine the most popular answer, and only… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n8-reformatted-selfconsistency.tabular100K<n<1M2 likes124 downloads8mo agoHugging Face26NoahEJ /fineweb-sample-100BT_over-1024-tokens-subsettabular10M<n<100M0 likes121 downloads1y agoHugging Face27marin-community /open-thoughts-4-30k-math-qwen3-235b-a22b-annotated-32768-tokens Dataset Card for Open-Thoughts-4-30K-Math-Qwen3-235B-A22B-Annotated-32768-Tokens Overview This dataset is a variant of marin-community/open-thoughts-4-30k-math-qwen3-235b-a22b-annotated with an extended maximum sequence length. The responses in the qwen235b_generated_text column were regenerated with max output tokens = 32768 (instead of 16000 in the original dataset), allowing for longer and more complete chain-of-thought reasoning. The conversations column has been… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-235b-a22b-annotated-32768-tokens.tabular10K<n<100K0 likes121 downloads7mo agoHugging Face28NoahEJ /fineweb-sample-100BT_over-2048-tokens-subsettabular1M<n<10M0 likes116 downloads1y agoHugging Face29jackyk02 /nemotron-cc-v2.1-hq-dqa-qwen3-tokens Nemotron-CC-v2.1 / High-Quality-DQA — tokenized with the Qwen3-8B tokenizer Question/answer pairs extracted from nvidia/Nemotron-CC-v2.1 (High-Quality-DQA subset) and tokenized with the Qwen/Qwen3-8B tokenizer. In the source data each row is a web document whose tail carries synthetic QA pairs marked Question: / Answer:. Here that document is split into its original prose (context) and the individual QA pairs, each tokenized separately. The Question: / Answer: marker keywords… See the full description on the dataset page: https://huggingface.co/datasets/jackyk02/nemotron-cc-v2.1-hq-dqa-qwen3-tokens.tabular100M<n<1B0 likes115 downloads2mo agoHugging Face30AmanPriyanshu /GPT-OSS-20B-benchmark-rollouts-512-tokens GPT-OSS-20B Benchmark Rollouts (512 tokens) This dataset contains text generation outputs from OpenAI's GPT-OSS-20B model across multiple evaluation benchmarks, with generation limited to 512 tokens. Dataset Description The dataset captures GPT-OSS-20B's text generation behavior when responding to prompts from established AI evaluation benchmarks. Each example includes the original prompt, the model's generated response, and token statistics. Benchmark Coverage… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/GPT-OSS-20B-benchmark-rollouts-512-tokens.tabular10K<n<100K0 likes112 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.