CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01caskcsg /NExtLong-128K-dataset NExtLong: Toward Effective Long-Context Training without Long Documents This repository contains the code ,models and datasets for our paper NExtLong: Toward Effective Long-Context Training without Long Documents. [Github] Quick Links Overview NExtLong Models NExtLong Datasets Datasets list How to use NExtLong datasets Bugs or Questions? Overview Large language models (LLMs) with extended context windows have made significant strides yet remain a… See the full description on the dataset page: https://huggingface.co/datasets/caskcsg/NExtLong-128K-dataset.text10K<n<100K1 likes726 downloads1y agoHugging Face02jinofy-corp /jora_corpus1_FR_tokenized_128ktabularn<1K2 likes612 downloads2mo agoHugging Face03donmaclean /LongMIT-128K LongMIT: Essential Factors in Crafting Effective Long Context Multi-Hop Instruction Datasets [ArXiv] Download LongMIT Datasets def download_longmit_datasets(dataset_name: str, save_dir: str): qa_pairs = [] dataset = load_dataset(dataset_name, split='train', cache_dir=HFCACHEDATASETS, trust_remote_code=True) for d in dataset: all_docs = d['all_docs'] if d['type'] in ['inter_doc', 'intra_doc']: if… See the full description on the dataset page: https://huggingface.co/datasets/donmaclean/LongMIT-128K.textquestion-answering10K<n<100K8 likes335 downloads2y agoHugging Face04aldea-ai /ruler_eval_data_128k RULER evaluation data — 128K context only This dataset is a subset of aldea-ai/ruler-eval-data, containing only the 128K context length (131072 tokens). The full multi-length dataset also includes 1M and other lengths. Files are published under 131072/ (numeric token count) for compatibility with benchmark_ruler.py --context_length 131072, even though the source snapshot uses a 128k/ folder name. Layout Same as the upstream RULER on-disk layout, compatible with… See the full description on the dataset page: https://huggingface.co/datasets/aldea-ai/ruler_eval_data_128k.tabular1K<n<10K0 likes120 downloads5mo agoHugging Face05caskcsg /entropylong_128ktext10K<n<100K0 likes118 downloads10mo agoHugging Face06Tongyi-Zhiwen /ruler-128k-subset ruler-128k-subset This is the partial dataset for evaluating QwenLong-CPRS textn<1K0 likes110 downloads1y agoHugging Face07Lala8383 /msmarco-item-id-hardneg-100shot-v4_128ktexttext-generation100K<n<1M0 likes79 downloads5mo agoHugging Face08caskcsg /Litelong_Nextlong_128ktext10K<n<100K1 likes69 downloads1y agoHugging Face09Lala8383 /msmarco-atomic-id-3shot-v4_128k_few_shot msmarco-atomic-id-3shot-v4_128k MSMARCO few-shot evaluation dataset for in-context learning generative retrieval, atomic-id variant. Identical construction to Lala8383/msmarco-item-id-3shot-v4_128k_few_shot, except every document's Identifier (and the answer target) is an arbitrary unique integer (Tay et al. DSI "Atomic Docid") instead of the natural-language document title. The id carries no semantics, so a retriever can only answer by matching the query to a document in… See the full description on the dataset page: https://huggingface.co/datasets/Lala8383/msmarco-atomic-id-3shot-v4_128k_few_shot.text10K<n<100K0 likes61 downloads3mo agoHugging Face10open-llm-leaderboard /EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-COT-detailsgated Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-COT Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-COT The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-COT-details.tabular10K<n<100K0 likes42 downloads2y agoHugging Face11open-llm-leaderboard /EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-0.001-128K-auto-detailsgated Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-0.001-128K-auto Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-0.001-128K-auto The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-0.001-128K-auto-details.tabular10K<n<100K0 likes41 downloads2y agoHugging Face12open-llm-leaderboard /EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-ds-auto-detailsgated Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-ds-auto Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-ds-auto The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-ds-auto-details.tabular10K<n<100K0 likes40 downloads2y agoHugging Face13open-llm-leaderboard /microsoft__Phi-3-mini-128k-instruct-detailsgated Dataset Card for Evaluation run of microsoft/Phi-3-mini-128k-instruct Dataset automatically created during the evaluation run of model microsoft/Phi-3-mini-128k-instruct The dataset is composed of 40 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/microsoft__Phi-3-mini-128k-instruct-details.tabular10K<n<100K0 likes37 downloads2y agoHugging Face14open-llm-leaderboard /NousResearch__Yarn-Llama-2-7b-128k-detailsgated Dataset Card for Evaluation run of NousResearch/Yarn-Llama-2-7b-128k Dataset automatically created during the evaluation run of model NousResearch/Yarn-Llama-2-7b-128k The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NousResearch__Yarn-Llama-2-7b-128k-details.tabular10K<n<100K0 likes36 downloads2y agoHugging Face15open-llm-leaderboard /NousResearch__Yarn-Llama-2-13b-128k-detailsgated Dataset Card for Evaluation run of NousResearch/Yarn-Llama-2-13b-128k Dataset automatically created during the evaluation run of model NousResearch/Yarn-Llama-2-13b-128k The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NousResearch__Yarn-Llama-2-13b-128k-details.tabular10K<n<100K0 likes36 downloads2y agoHugging Face16open-llm-leaderboard /EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-detailsgated Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-details.tabular10K<n<100K1 likes36 downloads2y agoHugging Face17open-llm-leaderboard /EpistemeAI2__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math-detailsgated Dataset Card for Evaluation run of EpistemeAI2/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math-details.tabular10K<n<100K0 likes36 downloads2y agoHugging Face18open-llm-leaderboard /EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-detailsgated Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-details.tabular10K<n<100K0 likes36 downloads2y agoHugging Face19open-llm-leaderboard /EpistemeAI__Polypsyche-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-Empathy-detailsgated Dataset Card for Evaluation run of EpistemeAI/Polypsyche-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-Empathy Dataset automatically created during the evaluation run of model EpistemeAI/Polypsyche-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-Empathy The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Polypsyche-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-Empathy-details.tabular10K<n<100K0 likes36 downloads2y agoHugging Face20open-llm-leaderboard /EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-detailsgated Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 5 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-details.tabular10K<n<100K2 likes35 downloads2y agoHugging Face21open-llm-leaderboard /EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-detailsgated Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-details.tabular10K<n<100K0 likes34 downloads2y agoHugging Face22open-llm-leaderboard /EpistemeAI2__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.005-128K-code-COT-detailsgated Dataset Card for Evaluation run of EpistemeAI2/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.005-128K-code-COT Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.005-128K-code-COT The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.005-128K-code-COT-details.tabular10K<n<100K0 likes34 downloads2y agoHugging Face23open-llm-leaderboard /EpistemeAI__Fireball-Meta-Llama-3.2-8B-Instruct-agent-003-128k-code-DPO-detailsgated Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.2-8B-Instruct-agent-003-128k-code-DPO Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.2-8B-Instruct-agent-003-128k-code-DPO The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.2-8B-Instruct-agent-003-128k-code-DPO-details.tabular10K<n<100K0 likes33 downloads2y agoHugging Face24open-llm-leaderboard /NousResearch__Yarn-Mistral-7b-128k-detailsgated Dataset Card for Evaluation run of NousResearch/Yarn-Mistral-7b-128k Dataset automatically created during the evaluation run of model NousResearch/Yarn-Mistral-7b-128k The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NousResearch__Yarn-Mistral-7b-128k-details.tabular10K<n<100K0 likes32 downloads2y agoHugging Face25teamsleeping /udio-128Ktabular100K<n<1M0 likes31 downloads2mo agoHugging Face26open-llm-leaderboard /microsoft__Phi-3-medium-128k-instruct-detailsgated Dataset Card for Evaluation run of microsoft/Phi-3-medium-128k-instruct Dataset automatically created during the evaluation run of model microsoft/Phi-3-medium-128k-instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/microsoft__Phi-3-medium-128k-instruct-details.tabular10K<n<100K0 likes26 downloads2y agoHugging Face27Renjie-Ranger /open_r1_math_all_sampled_128ktext100K<n<1M0 likes23 downloads11mo agoHugging Face28open-llm-leaderboard /EpistemeAI__Polypsyche-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-Logic-detailsgated Dataset Card for Evaluation run of EpistemeAI/Polypsyche-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-Logic Dataset automatically created during the evaluation run of model EpistemeAI/Polypsyche-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-Logic The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Polypsyche-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-Logic-details.tabular10K<n<100K0 likes19 downloads2y agoHugging Face29Lala8383 /msmarco-item-id-3shot-v4_128k_mixtext100K<n<1M0 likes19 downloads5mo agoHugging Face30open-llm-leaderboard /microsoft__Phi-3-small-128k-instruct-detailsgated Dataset Card for Evaluation run of microsoft/Phi-3-small-128k-instruct Dataset automatically created during the evaluation run of model microsoft/Phi-3-small-128k-instruct The dataset is composed of 35 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/microsoft__Phi-3-small-128k-instruct-details.tabular10K<n<100K0 likes14 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.