datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
encoder-decoder-trial-stat
Encoder/decoder trial: encoder-marginal report
Dataset: G-reen/encoder-decoder-trial-stat
Rows analysed: 122,933 (every kept (encoder, decoder, source row) triple; source G-reen/cc-re-2021-filtered shard 0, 2000 rows of at most 4000 words)
Prompt file: prompts/indirect_reference_dataset_train.json (turn 0 encodes the document, turn 1 reconstructs it from the encoding alone)
Encoders: 9 (granite-4.2-30b-nvfp4 [0], Ornith-1.5-35B-A3B-NVFP4 [1], Llama-3.3-70B-Instruct-NVFP4 [2]… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/encoder-decoder-trial-stat.encoder-decoder-trial-rewritten
Encoder/decoder trial: decoded texts
Turn-1 outputs: each config shard_<index> is one decoder's reconstruction of every encoder's encodings from G-reen/encoder-decoder-trial-encodings. encoder_model / decoder_model name the pair; response_0 is the encoding the decoder saw. Rows failing post-processing are in shard_<index>_trashed_data; per-run counters are in runs/.
encoder-decoder-trial-encodings
Encoder/decoder trial: encodings
Turn-0 outputs of the indirect-reference prompts: each config encoder_<index> is one model's encoding of the same 2000 rows (source_row of G-reen/cc-re-2021-filtered shard 0, at most 4000 words). prompt records the template each row was given, identically across encoders. Per-run counters are in runs/.
encoder-decoder-trial-human-editlensfintime-decoder-dataset
Dataset Card for FinTime Dataset
Dataset Summary
FinTime Dataset is a comprehensive, large-scale financial time series dataset designed for training and evaluating decoder-only models on financial forecasting tasks across diverse asset classes and market conditions.
Composed of real-world market data spanning equities, cryptocurrencies, forex, commodities, and indices, the dataset captures the complexity, volatility, and multi-scale dynamics typical of financial markets.… See the full description on the dataset page: https://huggingface.co/datasets/thesven/fintime-decoder-dataset.decoderstack-gsm8k
decoderstack-gsm8k
GSM8K, pre-tokenized for stacks/decoder-rtx/train_gsm8k.py (the RL sanity-check
pipeline of the DecoderStack backward-pass speedrun): ClimbMix 32k ids with the
nanochat chat template already applied. The trainer downloads these files and never
tokenizes.
file
rows
what
prompts.parquet
7473 train + 1319 test
[bos, user_start, *question, user_end, assistant_start], gold answer, is_val (256 seeded test problems = the in-loop validation tracker)… See the full description on the dataset page: https://huggingface.co/datasets/ChrisMcCormick/decoderstack-gsm8k.eval_dit_posttrainv2_union_norm_bs128_seqlora_dit_all_decoder_real_1_on_real_1_seed1000This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "franka",
"total_episodes": 10,
"total_frames": 3498,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 1,"video_files_size_in_mb": 1,
"fps": 15,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/continuallearning/eval_dit_posttrainv2_union_norm_bs128_seqlora_dit_all_decoder_real_1_on_real_1_seed1000.DecoderLoraEncoder-DecoderL7_8_sampled_news_headlines_decoder_classificationDecoder_onlyDecoder-OnlyDecoderNormalTraining
