CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01damerajee /long_context_hin_22ktext100K<n<1M0 likes399 downloads2y agoHugging Face02damerajee /long_context_hindi Dataset This dataset was filtered from AI4BHarat dataset sangraha,which is the largest high-quality, cleaned Indic language pretraining data containing 251B tokens summed up over 22 languages, extracted from curated sources, existing multilingual corpora and large scale translations. This dataset contains only Hindi as of now Information First this dataset is mainly for long context training The minimum len is 6000 and maximum len is 3754718 Getting started… See the full description on the dataset page: https://huggingface.co/datasets/damerajee/long_context_hindi.texttext-generation100K<n<1M1 likes209 downloads2y agoHugging Face03Seerkfang /LongMagpie_multidoc_longcontext_datasettext100K<n<1M5 likes168 downloads1y agoHugging Face04llm-semantic-router /longcontext-haldetect Long-Context Hallucination Detection Benchmark A synthetic benchmark dataset for evaluating hallucination detection models on long documents (8K-24K tokens). This dataset is specifically designed to test models that can handle contexts beyond the typical 8K token limit. Dataset Summary Property Value Total samples 3,366 Token range 8,005 - 23,998 Average tokens 17,852 Hallucinated 1,681 (49.9%) Supported 1,685 (50.1%) Splits… See the full description on the dataset page: https://huggingface.co/datasets/llm-semantic-router/longcontext-haldetect.texttoken-classificationn<1K0 likes148 downloads9mo agoHugging Face05Emulated-Inc /long-context-retrieval-training-pool Long context retrieval training pool Long prompts with short, checkable answers. Each row is one complete message: a task instruction, a long body of text that hides what the question is about, and the question itself, together with every string an answer has to contain for it to be right. The bodies run from four thousand to thirty-two thousand tokens. Three sources, laid out twice. Train on either layer or on both. pool.jsonl Every source rewritten into one… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/long-context-retrieval-training-pool.texttext-generation10K<n<100K1 likes128 downloads14d agoHugging Face06crellis /longcontext_datasettext1M<n<10M0 likes119 downloads5mo agoHugging Face07jinaai /longcontext-cmrc2018-zhtext1K<n<10K2 likes65 downloads3y agoHugging Face08antash420 /long-context-text-summarization-alpaca-formattext100K<n<1M1 likes42 downloads2y agoHugging Face09AIGym /long-context-reasoning-v1text10K<n<100K0 likes41 downloads1y agoHugging Face10Abzu /long-context-qa-df Dataset Card for "long-context-qa-df" More Information needed textn<1K2 likes39 downloads3y agoHugging Face11RanaGaber /Long_Context_MT_ALL_EGtext10K<n<100K0 likes35 downloads2mo agoHugging Face12jaehyeokdoo2 /rankzephyr_longcontext_range80-100_merged_80ktext10K<n<100K0 likes27 downloads2y agoHugging Face13jinaai /longcontext-cmrc2018-zh-qrelstext1K<n<10K1 likes25 downloads3y agoHugging Face14tilde-research /long-contexttabularn<1K1 likes22 downloads1y agoHugging Face15jaehyeokdoo2 /rankzephyr_longcontext_merged_140ktext100K<n<1M0 likes21 downloads2y agoHugging Face16jaehyeokdoo2 /rankzephyr_longcontext_range80-100_merged_140ktext100K<n<1M0 likes21 downloads2y agoHugging Face17thusinh1969 /llama-2-7b-LongContext-mixed-32k-30APRIL2024text10K<n<100K0 likes20 downloads2y agoHugging Face18AIGym /longcontext-summarization-v1text100K<n<1M0 likes20 downloads1y agoHugging Face19thusinh1969 /llama-2-7b-LongContext-mixed-24k-30APRIL2024text10K<n<100K0 likes19 downloads2y agoHugging Face20jaehyeokdoo2 /rankzephyr_longcontext_merged_80ktext10K<n<100K0 likes18 downloads2y agoHugging Face21TAUR-dev /D-ExpTracker__1022_longcontext__maxlen4096_0epoch_3and4arg__v1tabularn<1K0 likes18 downloads11mo agoHugging Face22nguyentatdat /train_sft_longcontext_ver2text1K<n<10K0 likes14 downloads1y agoHugging Face23TAUR-dev /D-ExpTracker__1022_longcontext__maxlen8192_1e_3args__v1tabularn<1K0 likes13 downloads11mo agoHugging Face24thusinh1969 /llama-2-7b-LongContext-mixed-64k-30APRIL2024text10K<n<100K1 likes12 downloads2y agoHugging Face25damerajee /long_context_hin_10ktextn<1K0 likes12 downloads2y agoHugging Face26dadu /long-context-patternizedtext1K<n<10K0 likes12 downloads2y agoHugging Face27TAUR-dev /D-ExpTracker__1022_longcontext__maxlen4096_0epoch_3args__v1tabularn<1K0 likes12 downloads11mo agoHugging Face28keypa /qwen36-adapter-longcontext-sfttext1K<n<10K0 likes12 downloads5mo agoHugging Face29nguyentatdat /train_sft_longcontexttext1K<n<10K0 likes9 downloads1y agoHugging Face30TAUR-dev /D-ExpTracker__1022_longcontext__maxlen8192_0epoch_3args__v1tabularn<1K0 likes8 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.