datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
compressed_filesDiscord-Unveiled-Compressed
.hf-sanitized.hf-sanitized-uCWd6SwyNH8FCkRETeRYS .container { --bg-primary: #0d0511; --bg-secondary: #1a0f1f; --bg-tertiary: #2d1b35; --bg-card: #3d2847; --text-primary: #fef7ff; --text-secondary: #f0d9ff; --text-muted: #c084fc; --pink-soft: #fce7f3; --pink-medium: #f9a8d4; --pink-bright: #ec4899; --pink-hot: #e91e63; --pink-neon: #ff1493; --purple-soft: #e879f9; --purple-bright: #c026d3; --purple-deep: #7c3aed; --border-glow: #f472b6; --shadow-pink: rgba(244, 114, 182, 0.4);… See the full description on the dataset page: https://huggingface.co/datasets/SaisExperiments/Discord-Unveiled-Compressed.finance_data_compressed
Dataset Card for Finance Data Compressed
This dataset is created using the methodology introduced in LLMLingua-2 (Pan et al., 2024), and is collected to construct the training data for LLMLingua-2 compressor.
It consists of 5000 instances from AdaptLLM/finance-tasks, with their GPT-3.5-turbo compressed versions.
This dataset consists of pairs of original prompts and their compressed versions (specifically for financial data).
🎯 Usage
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/anshkhandelwal/finance_data_compressed.article-digests-compressedTheRoyalCarpet-Deduped-Compressed
The Royal Carpet: Cleaned
Original Dataset: KaraKaraWitch/the-royal-carpet
Thanks KaraKaraWitch!
Why?
The original dataset is 40GB in size (at least in decimal), which is quite the large set! 40GB would likely overwhelm most systems since I doubt most people have more than 32GB of RAM.
Not to mention too, this dataset contained a redundant row, html, which is just a exact copy of the text row with HTML crap, removing that row brought down most of the dataset to… See the full description on the dataset page: https://huggingface.co/datasets/LyraNovaHeart/TheRoyalCarpet-Deduped-Compressed.nq-item-id-llm-compressed-refinednq-item-id-llm-compressed-hardneg-3shot-v4
Token Statistics
===== Token Statistics =====
Tokenizer: Abner0803/Qwen3-1.7B-icl-3shot-v4_128k-copy_tag
Input file: train_3shot.jsonl
Format: conversations
Text field: text
Message scope: all
Operation filter: None
Eligible examples: 417748
Sample size: 1000
Random seed: 42
Add special tokens: False
Examples: 1000
Skipped: 0
Skipped by operation: 0
Total tokens: 892617
Average tokens/example: 892.62
Min tokens/example: 539
Max tokens/example: 1772
P50 tokens/example: 883
P90… See the full description on the dataset page: https://huggingface.co/datasets/Lala8383/nq-item-id-llm-compressed-hardneg-3shot-v4.mydataset_compressed_gsm8k_llmlingua2_qwen_3Bnq-item-id-llm-compressed-hardneg-10shot-v4
Token Statistics
===== Token Statistics =====
Tokenizer: Abner0803/Qwen3-1.7B-icl-3shot-v4_128k-copy_tag
Input file: train_10shot.jsonl
Format: conversations
Text field: text
Message scope: all
Operation filter: None
Eligible examples: 417748
Sample size: 100
Random seed: 42
Add special tokens: False
Examples: 100
Skipped: 0
Skipped by operation: 0
Total tokens: 273820
Average tokens/example: 2738.20
Min tokens/example: 2110
Max tokens/example: 3917
P50 tokens/example: 2724
P90… See the full description on the dataset page: https://huggingface.co/datasets/Lala8383/nq-item-id-llm-compressed-hardneg-10shot-v4.nq-item-id-llm-compressed-hardneg-20shot-v4
Token Statistics
===== Token Statistics =====
Tokenizer: Abner0803/Qwen3-1.7B-icl-3shot-v4_128k-copy_tag
Input file: train_20shot.jsonl
Format: conversations
Text field: text
Message scope: all
Operation filter: None
Eligible examples: 417748
Sample size: 100
Random seed: 42
Add special tokens: False
Examples: 100
Skipped: 0
Skipped by operation: 0
Total tokens: 537835
Average tokens/example: 5378.35
Min tokens/example: 4165
Max tokens/example: 6844
P50 tokens/example: 5401
P90… See the full description on the dataset page: https://huggingface.co/datasets/Lala8383/nq-item-id-llm-compressed-hardneg-20shot-v4.mydataset_compressed_gsm8k_llmlingua2_qwen_3B_ner_enhancedaria-math-reasoning-compressednq-item-id-compressed-convnq-item-id-llm-compressed_v1nq-item-id-llm-compressed_refine_v2
