datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hq-20smsmarco-item-id-hardneg-filter-20shot-v4nq-item-id-llm-refined-hardneg-20shot-v4nq-item-id-llm-compressed-hardneg-20shot-v4
Token Statistics
===== Token Statistics =====
Tokenizer: Abner0803/Qwen3-1.7B-icl-3shot-v4_128k-copy_tag
Input file: train_20shot.jsonl
Format: conversations
Text field: text
Message scope: all
Operation filter: None
Eligible examples: 417748
Sample size: 100
Random seed: 42
Add special tokens: False
Examples: 100
Skipped: 0
Skipped by operation: 0
Total tokens: 537835
Average tokens/example: 5378.35
Min tokens/example: 4165
Max tokens/example: 6844
P50 tokens/example: 5401
P90… See the full description on the dataset page: https://huggingface.co/datasets/Lala8383/nq-item-id-llm-compressed-hardneg-20shot-v4.
