datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
encoder-decoder-floresp-scorestemp-decoder-train-tokenizedSylReg-Decoderencoder-decoder-trial-stat
Encoder/decoder trial: encoder-marginal report
Dataset: G-reen/encoder-decoder-trial-stat
Rows analysed: 122,933 (every kept (encoder, decoder, source row) triple; source G-reen/cc-re-2021-filtered shard 0, 2000 rows of at most 4000 words)
Prompt file: prompts/indirect_reference_dataset_train.json (turn 0 encodes the document, turn 1 reconstructs it from the encoding alone)
Encoders: 9 (granite-4.2-30b-nvfp4 [0], Ornith-1.5-35B-A3B-NVFP4 [1], Llama-3.3-70B-Instruct-NVFP4 [2]… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/encoder-decoder-trial-stat.encoder-decoder-trial-rewritten
Encoder/decoder trial: decoded texts
Turn-1 outputs: each config shard_<index> is one decoder's reconstruction of every encoder's encodings from G-reen/encoder-decoder-trial-encodings. encoder_model / decoder_model name the pair; response_0 is the encoding the decoder saw. Rows failing post-processing are in shard_<index>_trashed_data; per-run counters are in runs/.
dream-decoder-dataset
Dream Decoder Synthetic Dataset
Size: 1,200 examplesModality: Text (dream_text, interpretation)Fields: id, dream_text, interpretation, symbols, emotions, setting, actions, tags, source
How it was created
Base data generated with templated combinations (symbols, emotions, settings, actions).
~300 dreams were paraphrased with google/flan-t5-base to satisfy the "use a HF model" requirement.
Intended use
For demo/building a dream similarity & recommendation app.… See the full description on the dataset page: https://huggingface.co/datasets/samvlad/dream-decoder-dataset.nav2tex-decoder-latex-pretrainencoder-decoder-trial-encodings
Encoder/decoder trial: encodings
Turn-0 outputs of the indirect-reference prompts: each config encoder_<index> is one model's encoding of the same 2000 rows (source_row of G-reen/cc-re-2021-filtered shard 0, at most 4000 words). prompt records the template each row was given, identically across encoders. Per-run counters are in runs/.
encoder-decoder-trial-human-editlensfintime-decoder-dataset
Dataset Card for FinTime Dataset
Dataset Summary
FinTime Dataset is a comprehensive, large-scale financial time series dataset designed for training and evaluating decoder-only models on financial forecasting tasks across diverse asset classes and market conditions.
Composed of real-world market data spanning equities, cryptocurrencies, forex, commodities, and indices, the dataset captures the complexity, volatility, and multi-scale dynamics typical of financial markets.… See the full description on the dataset page: https://huggingface.co/datasets/thesven/fintime-decoder-dataset.riemannian-generative-decoder
Riemannian generative decoder dataset
This repository contains the data related to the paper "Riemannian generative decoder".
Project Page: https://yhsure.github.io/riemannian-generative-decoder
Code Repository: https://github.com/yhsure/riemannian-generative-decoder
Abstract
Riemannian representation learning typically relies on an encoder to estimate densities on chosen manifolds. This involves optimizing numerically brittle objectives, potentially harming model… See the full description on the dataset page: https://huggingface.co/datasets/yhsure/riemannian-generative-decoder.decoderstack-gsm8k
decoderstack-gsm8k
GSM8K, pre-tokenized for stacks/decoder-rtx/train_gsm8k.py (the RL sanity-check
pipeline of the DecoderStack backward-pass speedrun): ClimbMix 32k ids with the
nanochat chat template already applied. The trainer downloads these files and never
tokenizes.
file
rows
what
prompts.parquet
7473 train + 1319 test
[bos, user_start, *question, user_end, assistant_start], gold answer, is_val (256 seeded test problems = the in-loop validation tracker)… See the full description on the dataset page: https://huggingface.co/datasets/ChrisMcCormick/decoderstack-gsm8k.MixAtis_for_DecoderOnly
Dataset Card for "MixAtis_for_DecoderOnly"
More Information needed
pet-decoder-audio-fixtures
Pet Decoder Audio Test Fixtures (v1.0) 🧪
This repository contains audio test fixtures used for validating the ingestion pipeline of the Pet Decoder AI application.
These are verified, public domain samples used to test our audio visualization and classification algorithms against known baselines (e.g., Low Frequency vs. High Frequency vocalizations).
Dataset Contents
The dataset consists of 5 reference audio files representing distinct spectral patterns:… See the full description on the dataset page: https://huggingface.co/datasets/petdecoder/pet-decoder-audio-fixtures.MixSnips_for_DecoderOnly_90-10_split-HALF
Dataset Card for "MixSnips_for_DecoderOnly_90-10_split-HALF"
More Information needed
mineru-decoder-finetune-dataset
MinerU Decoder Fine-Tune Dataset (real data)
Note: the dataset viewer is disabled — the data is distributed as zip bundles, not a
HuggingFace-loadable table. Download + unpack with the reproduce script (below), don't use load_dataset.
68,798 real image→tagged-text samples for teaching a VLM to emit inline formatting tags
(bold, italic, sup, sub, underline, strike). Sources: 22 arXiv categories (bold/italic/sup/sub) +
Washington & Texas legislative bills (underline/strike).… See the full description on the dataset page: https://huggingface.co/datasets/ahamad-ai/mineru-decoder-finetune-dataset.pantheon-ui-decoder-conversations
Pantheon UI Decoder Conversations
Training dataset for the decoder half of the Pantheon UI round-trip translator. The encoder turns natural language into emoji; the decoder takes emoji back to natural language.
Inspired by Anthropic's Natural Language Autoencoders — emoji as a discrete, human-legible intermediate between two model passes.
How it was built
Each row is derived from shreyask/pantheon-ui-conversations by inverting the encoder pairs:
Encoder pair:… See the full description on the dataset page: https://huggingface.co/datasets/shreyask/pantheon-ui-decoder-conversations.eval_dit_posttrainv2_union_norm_bs128_seqlora_dit_all_decoder_real_1_on_real_1_seed1000This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "franka",
"total_episodes": 10,
"total_frames": 3498,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 1,"video_files_size_in_mb": 1,
"fps": 15,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/continuallearning/eval_dit_posttrainv2_union_norm_bs128_seqlora_dit_all_decoder_real_1_on_real_1_seed1000.MixSnips_for_DecoderOnly
Dataset Card for "MixSnips_for_DecoderOnly"
More Information needed
MixAtis_for_DecoderOnly_90-10_split
Dataset Card for "MixAtis_for_DecoderOnly_90-10_split"
More Information needed
MixAtis_for_DecoderOnly_90-10_split-HALF
Dataset Card for "MixAtis_for_DecoderOnly_90-10_split-HALF"
More Information needed
E31_decoder_gen_fsqprescription-decoder-dataset-cleanDecoderLoraEncoder-Decoderprescription-decoder-datasetdecoder_patch_attributionDoctor_miniMixSnips_for_DecoderOnly_90-10_split
Dataset Card for "MixSnips_for_DecoderOnly_90-10_split"
More Information needed
low-decoder
low-decoding
Author: Surpem
This dataset contains 1200 unique, clean synthetic audio signals representing decoded text commands.
The audio signals represent synthesized Morse Code message blocks.
Dataset Structure
id: A unique UUID string.
audio: The audio wav bytes (16kHz Mono).
text: The decoded string transcription.
Dataset Level: LOW
Low: Slow WPM (~12 WPM), high signal-to-noise ratio (clean), short command strings.
Medium: Fast WPM (~24 WPM), background… See the full description on the dataset page: https://huggingface.co/datasets/Surpem/low-decoder.
