datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
long_context_hin_22klong_context_hindi
Dataset
This dataset was filtered from AI4BHarat dataset sangraha,which is the largest high-quality, cleaned Indic language pretraining data containing 251B tokens summed up over 22 languages, extracted from curated sources, existing multilingual corpora and large scale translations.
This dataset contains only Hindi as of now
Information
First this dataset is mainly for long context training
The minimum len is 6000 and maximum len is 3754718
Getting started… See the full description on the dataset page: https://huggingface.co/datasets/damerajee/long_context_hindi.LongMagpie_multidoc_longcontext_datasetlongcontext-haldetect
Long-Context Hallucination Detection Benchmark
A synthetic benchmark dataset for evaluating hallucination detection models on long documents (8K-24K tokens). This dataset is specifically designed to test models that can handle contexts beyond the typical 8K token limit.
Dataset Summary
Property
Value
Total samples
3,366
Token range
8,005 - 23,998
Average tokens
17,852
Hallucinated
1,681 (49.9%)
Supported
1,685 (50.1%)
Splits… See the full description on the dataset page: https://huggingface.co/datasets/llm-semantic-router/longcontext-haldetect.long-context-retrieval-training-pool
Long context retrieval training pool
Long prompts with short, checkable answers. Each row is one complete message: a task instruction,
a long body of text that hides what the question is about, and the question itself, together with
every string an answer has to contain for it to be right. The bodies run from four thousand to
thirty-two thousand tokens. Three sources, laid out twice. Train on either layer or on
both.
pool.jsonl
Every source rewritten into one… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/long-context-retrieval-training-pool.longcontext_datasetlongcontext-cmrc2018-zhlong-context-text-summarization-alpaca-formatlong-context-reasoning-v1long-context-qa-df
Dataset Card for "long-context-qa-df"
More Information needed
Long_Context_MT_ALL_EGrankzephyr_longcontext_range80-100_merged_80klongcontext-cmrc2018-zh-qrelslong-contextrankzephyr_longcontext_merged_140krankzephyr_longcontext_range80-100_merged_140kllama-2-7b-LongContext-mixed-32k-30APRIL2024longcontext-summarization-v1llama-2-7b-LongContext-mixed-24k-30APRIL2024rankzephyr_longcontext_merged_80kD-ExpTracker__1022_longcontext__maxlen4096_0epoch_3and4arg__v1train_sft_longcontext_ver2D-ExpTracker__1022_longcontext__maxlen8192_1e_3args__v1llama-2-7b-LongContext-mixed-64k-30APRIL2024long_context_hin_10klong-context-patternizedD-ExpTracker__1022_longcontext__maxlen4096_0epoch_3args__v1qwen36-adapter-longcontext-sfttrain_sft_longcontextD-ExpTracker__1022_longcontext__maxlen8192_0epoch_3args__v1
