datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cite-llm-multi-cite-trainllmpicto-commonvoice-v2-multiAutomatic translation of benoitfavre/llmpicto-commonvoice-v2 from French to languages where the Arasaac lexicon is available. Generated with facebook/nllb-200-distilled-600m. Contains about 500k sentences for 35 languages.
NASDAQ-News-Multi-LLM-Scores
NASDAQ News Multi-LLM Scores
127,176 financial news articles scored by 11 state-of-the-art LLMs for sentiment and risk assessment.
This dataset takes the same articles from FNSPID / FinRL_DeepSeek and re-scores them using multiple LLMs with varying reasoning effort levels and summary inputs. It enables direct cross-model comparison of financial sentiment analysis on identical articles.
Motivation
When we began using the FNSPID dataset for RL trading agent research, we… See the full description on the dataset page: https://huggingface.co/datasets/HYL/NASDAQ-News-Multi-LLM-Scores.harmful-prompts-multillm-before-guardrail-evaluationharmful-prompts-multillm-after-guardrail-evaluation-newmulti_llm_dpogn-multi-affective-alpaca
Guaraní Multidimensional Affective Alpaca Dataset
This dataset is a multidimensional affective instruction-tuning dataset for Guaraní (gn), created by preprocessing and unifying three original Guaraní classification datasets into the Alpaca format (instruction/input/output). It is designed for fine-tuning large language models (LLMs) on low-resource indigenous language tasks.
📚 Source Datasets
This dataset combines and transforms the following datasets from Marvin M.… See the full description on the dataset page: https://huggingface.co/datasets/Capibara-LLM/gn-multi-affective-alpaca.multillm-route-instruct2rebel_multi_llm_as_userseed_math_multiple_samples_scale_up_scaredy_cat_multiple_samples_llm_verifier_multi_domainharmful-prompts-multillm-after-guardrail-evaluationcite-llm-multi-cite-evall1-10-multi-label-tagger-datafft-multi-tool-router-datamultillm-route-instructharmful-prompts-multillm-after-guardrail-evaluation-newharmful-prompts-multillm-after-guardrail-evaluation-newvirtuoussy_multi_subject_rlvr_llm_judgewiki_rephrase_multillm_brazil
wiki_rephrase_multillm_brazil
Dataset em portugues derivado de conteudos da Wikipedia em portugues, com um subset
para os textos originais e subsets separados para reescritas sinteticas. A reescrita
foi feita utilizando quatro modelos diferentes de três famílias de llms:
google/gemma-4-E4B-it
google/gemma-4-E2B-it
mistralai/Ministral-3-3B-Instruct-2512
Qwen/Qwen3.5-4B
Repositorio alvo: wiki_rephrase_multillm_brazil.
Visibilidade pretendida no Hugging Face Hub: private.… See the full description on the dataset page: https://huggingface.co/datasets/br-llm-data/wiki_rephrase_multillm_brazil.
