CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01JackHsieh /32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids 32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen3-32B with thinking mode off. Each thought is a few dense sentences of reasoning about the next 8 tokens after a cut, written from the document prefix alone — the generator never sees the continuation. Stored thought_text includes the <thought>/</thought> wrapper. The 4B parity counterpart is… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.tabulartext-generation10M<n<100M0 likes593 downloads24d agoHugging Face02JackHsieh /4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids 4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen3-4B-Instruct-2507. Each thought is a few dense sentences of reasoning about the next 8 tokens after a cut, written from the document prefix alone — the generator never sees the continuation. Stored thought_text includes the <thought>/</thought> wrapper. This is the small-generator parity counterpart of… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.tabulartext-generation10M<n<100M0 likes517 downloads8d agoHugging Face03JackHsieh /4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids 4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen/Qwen3-4B in thinking mode. The prompted task: reason about what comes IMMEDIATELY next — the next k=8 tokens after the cut — and answer with a few dense, declarative sentences (about two to four) inside a <thought>…</thought> block, focused on the exact state at the cut and what the local grammar… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.tabular10M<n<100M0 likes336 downloads17d agoHugging Face04JackHsieh /4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained 4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained A pre-tokenized, tag-wrapped variant of JackHsieh/4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids. The source thoughts are two parts: a <think> block, then a few dense sentences that the generator wrapped in a literal <thought>…</thought> block (plain text, not special tokens). Only the part after </think> becomes the VALUE, with every <thought> / </thought> tag removed… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.tabular10M<n<100M0 likes336 downloads17d agoHugging Face05ZipLime /sec-8k-events SEC Form 8-K Corporate Events Every Form 8-K filed since the modern item taxonomy took effect — and, for each one, the second the SEC accepted it, which is not the date printed on it. 1 761 353 filings · 3 676 835 item-level events · 23 August 2004 to today The pipeline lives in recipe/ at the same revision as the data. See PIPELINE.md for the method. The problem this dataset exists to solve Apple filed its June-quarter results on 30 July 2026. Here is the filing… See the full description on the dataset page: https://huggingface.co/datasets/ZipLime/sec-8k-events.tabulartext-classification1M<n<10M0 likes321 downloads9d agoHugging Face06JackHsieh /4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained 4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained A pre-tokenized, tag-wrapped variant of JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids. The source thoughts are two parts: a <think> block, then a single paragraph of dense reasoning about the immediate continuation. Only the part after </think> — the paragraph — becomes the VALUE. The reasoning inside the think block is dropped. The VALUE is capped at 512 tokens.… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.tabular10M<n<100M0 likes306 downloads17d agoHugging Face07JackHsieh /4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids 4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen/Qwen3-4B in thinking mode. The prompted task: reason about what comes IMMEDIATELY next — the next k=8 tokens after the cut — and answer with a single unconstrained paragraph of dense reasoning, focused on the exact state at the cut and what the local grammar, notation, or argument forces next. Both the… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.tabular10M<n<100M0 likes303 downloads17d agoHugging Face08hazyresearch /OT_8K_seed_all_responsestabular100K<n<1M0 likes285 downloads11mo agoHugging Face09InfiniAILab /gsm_infinite_hard_8ktabular10K<n<100K1 likes267 downloads2y agoHugging Face10JackHsieh /4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained 4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained Tokenized, tag-wrapped form of JackHsieh/4B-reason-only.rule-r-1.0-k-8.L-512.statml-arxiv. Each thought is wrapped as <|note|> This is a hint about a span that appears later in this document. KEY is the text immediately before that span; VALUE is a note about what might come next. KEY: <last 8 prefix tokens> VALUE: <thought> <|/note|> and stored both as text (thought_text) and as Qwen/Qwen3-4B-Base… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.tabulartext-generation10M<n<100M0 likes209 downloads8d agoHugging Face11anirudhb11 /qwen3_30b_instruct_lcbv6_hardest_to_easiest_s_0_e_65_8kx160_t_1_gepatabular10K<n<100K0 likes193 downloads7mo agoHugging Face12JackHsieh /luna-reason-only.k-8.statml-arxiv-llama32.llama32-ids.tags-pausefilled luna-reason-only.k-8.statml-arxiv-llama32.llama32-ids.tags-pausefilled Every non-first chunk of every document carries a thought: the gpt-5.6-luna reasoning thought where one was generated, and a content-free pause thought everywhere else. luna chunk <|reserved_special_token_1|> luna reasoning <|reserved_special_token_2|> filler chunk <|reserved_special_token_1|> 256x <|reserved_special_token_0|> <|reserved_special_token_2|> The filler is 258 tokens. Chunk 0 is excluded… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/luna-reason-only.k-8.statml-arxiv-llama32.llama32-ids.tags-pausefilled.tabulartext-generation1M<n<10M0 likes191 downloads2mo agoHugging Face13JackHsieh /pause.rule-r-1.0-k-8.L-128.statml-arxivtabular1M<n<10M0 likes180 downloads2mo agoHugging Face14JackHsieh /4B-ranked-v7.rule-stride-train4-test32.k-8.L-4096.statml-arxivtabular1M<n<10M0 likes169 downloads24d agoHugging Face15Flinchlab /sec-8k-market-reaction-dataset Historical SEC 8-K Market-Reaction Dataset — by FlinchLab Know what each kind of company news historically did to the stock — before you build on top of it. 29,331 SEC 8-K filings from 660 US companies, read and re-classified finer than the official item codes, with each category's price reaction measured by a market-model event study: direction, magnitude, overnight-vs-session split, pre-filing baselines — and every published finding validated on two independent out-of-sample… See the full description on the dataset page: https://huggingface.co/datasets/Flinchlab/sec-8k-market-reaction-dataset.tabularn<1K0 likes155 downloads1mo agoHugging Face16anirudhb11 /qwen3_4b_instruct_lcbv6_hardest_to_easiest_s_0_e_65_8kx160_t_1_gepatabular10K<n<100K0 likes143 downloads7mo agoHugging Face17playwithmino /alimeeting-eval-8k AliMeeting Eval — 8 kHz CH0 clips Unknown-(N) eval clips from AliMeeting Eval (M2MeT / OpenSLR 119), far-field channel 0, resampled to 8 kHz. Manifest Clips manifests/eval_all.jsonl 1280 (full local eval) manifests/eval_n100.jsonl 100 stratified subset manifests/eval_n200.jsonl 200 stratified subset manifests/eval_*mix.jsonl by (N=1\ldots4) Headset s{k}.wav stems (near, TextGrid-gated) are present for the n200 subset (eval_n200_headset.jsonl). Other clips… See the full description on the dataset page: https://huggingface.co/datasets/playwithmino/alimeeting-eval-8k.audioaudio-to-audio1K<n<10K0 likes126 downloads22d agoHugging Face18YangZhoumill /infini_gsm_8k_noise_closetabular10K<n<100K0 likes120 downloads2y agoHugging Face19JackHsieh /luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by gpt-5.6-luna. Each thought is visible reasoning about the next 8 Qwen3 tokens after a cut, written without ever seeing that continuation. Intended to be spliced into the document before the chunk so a small model (Qwen3-4B-Base) can read the reasoning and predict the chunk. Thoughts are strings, not token ids;… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids.tabulartext-generation100K<n<1M0 likes119 downloads24d agoHugging Face20YangZhoumill /factorrandom_8k_hardtabular10K<n<100K0 likes114 downloads2y agoHugging Face21hkust-nlp /drkernel-coldstart-8k DR.Kernel Cold-Start Dataset Paper | Code This directory documents the format of hkust-nlp/drkernel-coldstart-8k. The cold-start set is used for supervised fine-tuning (SFT) before RL in DR.Kernel. As described in the paper, it is built from 5-turn multi-turn trajectories collected with KernelGYM feedback. Overview Purpose: initialize kernel-generation ability (Triton coding + iterative optimization) before TRLOO/MRS/PR/PRS RL.Data form: one row per full multi-turn… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/drkernel-coldstart-8k.tabulartext-generation1K<n<10K2 likes105 downloads8mo agoHugging Face22MVRL /TaxaBench-8k Paper: TaxaBind: A Unified Embedding Space for Ecological Applications Venue: WACV 2025 Github: https://github.com/mvrl/TaxaBind Dataset Name: TaxaBench-8k Dataset Description: TaxaBench-8k is a multimodal dataset containing six modalities - image, text, satellite image, audio, geographic location, and environmental features for evaluating large ecological models. Usage: Please use the test_df.csv for reading data which… See the full description on the dataset page: https://huggingface.co/datasets/MVRL/TaxaBench-8k.audiozero-shot-classification1K<n<10K1 likes103 downloads2y agoHugging Face23YangZhoumill /factor_medium_8ktabular10K<n<100K0 likes101 downloads2y agoHugging Face24JackHsieh /4B-Instruct-distill-jx739p0v.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained 4B-Instruct-distill-jx739p0v.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained A pre-tokenized, tag-wrapped variant of JackHsieh/4B-Instruct-distill-jx739p0v.stride-1.k-8.statml-arxiv.qwen3-ids. Each thought_text is wrapped as <|note|> This is a hint about a span that appears later in this document. KEY is the text immediately before that span; VALUE is a note about what might come next. KEY: {last 8 prefix tokens} VALUE: {thought_text} <|/note|> and stored both as text… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-distill-jx739p0v.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.tabular10M<n<100M0 likes98 downloads23d agoHugging Face25YangZhoumill /factor_hard_8ktabular10K<n<100K0 likes97 downloads2y agoHugging Face26alea-institute /kl3m-index-edgar-filings-8-ktabular1M<n<10M0 likes97 downloads2y agoHugging Face27InfiniAILab /gsm_infinite_medium_8ktabular10K<n<100K0 likes96 downloads2y agoHugging Face28YangZhoumill /factoranimalgptworld_8k_hardtabular10K<n<100K0 likes94 downloads2y agoHugging Face29YangZhoumill /infini_gsm_8k_noise_generaltabular10K<n<100K0 likes91 downloads2y agoHugging Face30YangZhoumill /factoranimalgptworld_8k_mediumtabular10K<n<100K0 likes87 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.