datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids
32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids
Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen3-32B
with thinking mode off. Each thought is a few dense sentences of reasoning about the next
8 tokens after a cut, written from the document prefix alone — the generator never sees the
continuation. Stored thought_text includes the <thought>/</thought> wrapper.
The 4B parity counterpart is… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids
4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids
Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by
Qwen3-4B-Instruct-2507. Each thought is a few dense sentences of reasoning about the next
8 tokens after a cut, written from the document prefix alone — the generator never sees the
continuation. Stored thought_text includes the <thought>/</thought> wrapper.
This is the small-generator parity counterpart of… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids
4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids
Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by
Qwen/Qwen3-4B in thinking mode.
The prompted task: reason about what comes IMMEDIATELY next — the next k=8 tokens after the
cut — and answer with a few dense, declarative sentences (about two to four) inside a
<thought>…</thought> block, focused on the exact state at the cut and what the local
grammar… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
A pre-tokenized, tag-wrapped variant of JackHsieh/4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.
The source thoughts are two parts: a <think> block, then a few dense sentences that the
generator wrapped in a literal <thought>…</thought> block (plain text, not special tokens).
Only the part after </think> becomes the VALUE, with every <thought> / </thought> tag
removed… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.sec-8k-events
SEC Form 8-K Corporate Events
Every Form 8-K filed since the modern item taxonomy took effect — and, for each
one, the second the SEC accepted it, which is not the date printed on it.
1 761 353 filings · 3 676 835 item-level events · 23 August 2004 to today
The pipeline lives in recipe/ at the same revision as the data.
See PIPELINE.md for the method.
The problem this dataset exists to solve
Apple filed its June-quarter results on 30 July 2026. Here is the filing… See the full description on the dataset page: https://huggingface.co/datasets/ZipLime/sec-8k-events.4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
A pre-tokenized, tag-wrapped variant of JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.
The source thoughts are two parts: a <think> block, then a single paragraph of dense
reasoning about the immediate continuation. Only the part after </think> — the paragraph —
becomes the VALUE. The reasoning inside the think block is dropped.
The VALUE is capped at 512 tokens.… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids
4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids
Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by
Qwen/Qwen3-4B in thinking mode.
The prompted task: reason about what comes IMMEDIATELY next — the next k=8 tokens after the
cut — and answer with a single unconstrained paragraph of dense reasoning, focused on the
exact state at the cut and what the local grammar, notation, or argument forces next. Both the… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.OT_8K_seed_all_responsesgsm_infinite_hard_8k4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
Tokenized, tag-wrapped form of JackHsieh/4B-reason-only.rule-r-1.0-k-8.L-512.statml-arxiv.
Each thought is wrapped as
<|note|>
This is a hint about a span that appears later in this document. KEY is the text immediately before that span; VALUE is a note about what might come next.
KEY: <last 8 prefix tokens>
VALUE: <thought>
<|/note|>
and stored both as text (thought_text) and as
Qwen/Qwen3-4B-Base… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.qwen3_30b_instruct_lcbv6_hardest_to_easiest_s_0_e_65_8kx160_t_1_gepaluna-reason-only.k-8.statml-arxiv-llama32.llama32-ids.tags-pausefilled
luna-reason-only.k-8.statml-arxiv-llama32.llama32-ids.tags-pausefilled
Every non-first chunk of every document carries a thought: the gpt-5.6-luna reasoning thought
where one was generated, and a content-free pause thought everywhere else.
luna chunk <|reserved_special_token_1|> luna reasoning <|reserved_special_token_2|>
filler chunk <|reserved_special_token_1|> 256x <|reserved_special_token_0|> <|reserved_special_token_2|>
The filler is 258 tokens. Chunk 0 is excluded… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/luna-reason-only.k-8.statml-arxiv-llama32.llama32-ids.tags-pausefilled.pause.rule-r-1.0-k-8.L-128.statml-arxiv4B-ranked-v7.rule-stride-train4-test32.k-8.L-4096.statml-arxivsec-8k-market-reaction-dataset
Historical SEC 8-K Market-Reaction Dataset — by FlinchLab
Know what each kind of company news historically did to the stock —
before you build on top of it. 29,331 SEC 8-K filings from 660 US
companies, read and re-classified finer than the official item codes, with
each category's price reaction measured by a market-model event study:
direction, magnitude, overnight-vs-session split, pre-filing baselines —
and every published finding validated on two independent out-of-sample… See the full description on the dataset page: https://huggingface.co/datasets/Flinchlab/sec-8k-market-reaction-dataset.qwen3_4b_instruct_lcbv6_hardest_to_easiest_s_0_e_65_8kx160_t_1_gepaalimeeting-eval-8k
AliMeeting Eval — 8 kHz CH0 clips
Unknown-(N) eval clips from AliMeeting Eval (M2MeT / OpenSLR 119), far-field channel 0, resampled to 8 kHz.
Manifest
Clips
manifests/eval_all.jsonl
1280 (full local eval)
manifests/eval_n100.jsonl
100 stratified subset
manifests/eval_n200.jsonl
200 stratified subset
manifests/eval_*mix.jsonl
by (N=1\ldots4)
Headset s{k}.wav stems (near, TextGrid-gated) are present for the n200 subset (eval_n200_headset.jsonl). Other clips… See the full description on the dataset page: https://huggingface.co/datasets/playwithmino/alimeeting-eval-8k.infini_gsm_8k_noise_closeluna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids
luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids
Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by
gpt-5.6-luna. Each thought is visible reasoning about the next 8 Qwen3 tokens after a cut,
written without ever seeing that continuation. Intended to be spliced into the document
before the chunk so a small model (Qwen3-4B-Base) can read the reasoning and predict the
chunk. Thoughts are strings, not token ids;… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids.factorrandom_8k_harddrkernel-coldstart-8k
DR.Kernel Cold-Start Dataset
Paper | Code
This directory documents the format of hkust-nlp/drkernel-coldstart-8k.
The cold-start set is used for supervised fine-tuning (SFT) before RL in DR.Kernel. As described in the paper, it is built from 5-turn multi-turn trajectories collected with KernelGYM feedback.
Overview
Purpose: initialize kernel-generation ability (Triton coding + iterative optimization) before TRLOO/MRS/PR/PRS RL.Data form: one row per full multi-turn… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/drkernel-coldstart-8k.TaxaBench-8k
Paper: TaxaBind: A Unified Embedding Space for Ecological Applications
Venue: WACV 2025
Github: https://github.com/mvrl/TaxaBind
Dataset Name: TaxaBench-8k
Dataset Description:
TaxaBench-8k is a multimodal dataset containing six modalities - image, text, satellite image, audio, geographic location, and environmental features for evaluating large ecological models.
Usage:
Please use the test_df.csv for reading data which… See the full description on the dataset page: https://huggingface.co/datasets/MVRL/TaxaBench-8k.factor_medium_8k4B-Instruct-distill-jx739p0v.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
4B-Instruct-distill-jx739p0v.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
A pre-tokenized, tag-wrapped variant of JackHsieh/4B-Instruct-distill-jx739p0v.stride-1.k-8.statml-arxiv.qwen3-ids. Each thought_text is wrapped as
<|note|>
This is a hint about a span that appears later in this document. KEY is the text immediately before that span; VALUE is a note about what might come next.
KEY: {last 8 prefix tokens}
VALUE: {thought_text}
<|/note|>
and stored both as text… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-distill-jx739p0v.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.factor_hard_8kkl3m-index-edgar-filings-8-kgsm_infinite_medium_8kfactoranimalgptworld_8k_hardinfini_gsm_8k_noise_generalfactoranimalgptworld_8k_medium
