datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Argimi-Ardian-Finance-10k-text
The ArGiMI Ardian datasets : Text-only version
The ArGiMi project is committed to open-source principles and data sharing.
Thanks to our generous partners, we are releasing several valuable datasets to the public.
Dataset description
This text-only dataset comprises 34,000 financial annual reports, written in English, meticulously
extracted from their original PDF format to provide a valuable resource for researchers and developers in financial
analysis and natural… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text.Light-Omni-Training
Light-Omni Training Dataset
This repository contains the training data used by Light-Omni, a multimodal
agent framework for reflexive video understanding with long-term memory.
Light-Omni uses memory-augmented multimodal streams to train adapters for
memory construction, response generation, and reaction/action control.
Links
Project page: https://clare-nie.github.io/Light-Omni/
Code: https://github.com/Clare-Nie/Light-Omni
Dataset:… See the full description on the dataset page: https://huggingface.co/datasets/ClareNie/Light-Omni-Training.Argimi-Ardian-Finance-10k-text-image
The ArGiMI Ardian datasets : text and images
The ArGiMi project is committed to open-source principles and data sharing.
Thanks to our generous partners, we are releasing several valuable datasets to the public.
Dataset description
This dataset comprises 34,000 financial annual reports, written in English, meticulously
extracted from their original PDF format to provide a valuable resource for researchers and developers in financial
analysis and natural language… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text-image.xlsum-subset
Dataset Card for "XL-Sum"
Dataset Summary
We present XLSum, a comprehensive and diverse dataset comprising 1.35 million professionally annotated article-summary pairs from BBC, extracted using a set of carefully designed heuristics. The dataset covers 45 languages ranging from low to high-resource, for many of which no public dataset is currently available. XL-Sum is highly abstractive, concise, and of high quality, as indicated by human and intrinsic evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/xlsum-subset.cs415-twitch-chatsefficient_llm
Data V4 for NeurIPS LLM Challenge
Contains 70949 samples collected from Huggingface:
Math: 1273
gsm8k
math_qa
math-eval/TAL-SCQ5K
TAL-SCQ5K-EN
meta-math/MetaMathQA
TIGER-Lab/MathInstruct
Science: 42513
lighteval/mmlu - 'all', "split": 'auxiliary_train'
lighteval/bbq_helm - 'all'
openbookqa - 'main'
ComplexQA: 2940
ARC-Challenge
ARC-Easy
piqa
social_i_qa
Muennighoff/babi
Rowan/hellaswag
ComplexQA1: 2060
medmcqa
winogrande_xl,
winogrande_debiased
boolq
sciq
CNN: 2787… See the full description on the dataset page: https://huggingface.co/datasets/transZ/efficient_llm.
