datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
legalbench.br
LegalBench.BR ⚖️🇧🇷
LegalBench.BR is a benchmark for evaluating large language models (LLMs) on tasks grounded in Brazilian Law and written in Brazilian Portuguese.
The dataset was designed to test whether general-purpose and legal-domain LLMs can answer, classify, infer, and recall legal information in the context of the Brazilian legal system. It covers multiple legal areas and combines synthetic legal questions, court-decision excerpts, legal entailment tasks, and closed-book… See the full description on the dataset page: https://huggingface.co/datasets/celsowm/legalbench.br.DomesticNames_AllStates_Text Domestic Names from the Federal Government's repository of official geographic names [CSV dataset]
This Dataset includes 980,065 geographic names as of September 10, 2023.
It is apparent that no currently released LLMs are pretrained on datasets with many of these geographic names (i.e., features), descriptions, and histories.
Example: feature_name: Abercrombie Gulch
GPT-3.5 responds "I'm not aware of a specific location called Abercrombie Gulch in my training data,..." when prompted… See the full description on the dataset page: https://huggingface.co/datasets/cellos/DomesticNames_AllStates_Text.aurynoab_geminiCelebrityThis dataset accompanies the paper:
When Do LLMs Admit Their Mistakes? Understanding the Role of Model Belief in Retraction
This dataset contains the original Celebrity questions with train/test split. Please see the paper for more details.
Code: https://github.com/ayyyq/llm-retraction
Citation
@misc{yang2025llmsadmitmistakesunderstanding,
title={When Do LLMs Admit Their Mistakes? Understanding the Role of Model Belief in Retraction},
author={Yuqing Yang and Robin Jia}… See the full description on the dataset page: https://huggingface.co/datasets/ayyyq/Celebrity.CelebA_RoBERTa_Sp
Corpus Summary
This corpus contains 250000 entries made up of a pair of sentences in Spanish and their respective similarity value in the range 0 to 1. This corpus was used in the training of the
sentence-transformer library to improve the efficiency of the RoBERTa-large-bne base model.
Each of the pairs of sentences are textual descriptions of the faces of the CelebA dataset, which were previously translated into Spanish. The process followed to generate it was:
First, a… See the full description on the dataset page: https://huggingface.co/datasets/oeg/CelebA_RoBERTa_Sp.gemini_orpo_dpo_ptbrsimulado_oabreddit_dataset_190
Bittensor Subnet 13 Reddit Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks.
For more… See the full description on the dataset page: https://huggingface.co/datasets/CelestialWandererOfTheVoid/reddit_dataset_190.CelebA_Sent2Vect_Sp
Corpus Summary
This corpus has 192050 entries made up of descriptive sentences of the faces of the CelebA dataset.
The preprocessing of the corpus has been to translate into Spanish the captions of the CelebA dataset with the algorithm used in Text2FaceGAN.
In particular, all sentences are combined to generate a larger corpus. Additionally, a data preprocessing was applied that consists of eliminating stopwords, separation symbols and complementary elements that are not useful for… See the full description on the dataset page: https://huggingface.co/datasets/oeg/CelebA_Sent2Vect_Sp.demons_megaten_fandom_orpo_dpo_englishminecraft_qa_es
Minecraft Q&A (Spanish)
A Spanish, chat-formatted question/answer dataset about Minecraft. Each example is a short conversation with a single user question and a single assistant answer (plus a system prompt).
Data format
The dataset is provided as JSON Lines (.jsonl): one JSON object per line.
Each record has a single key:
messages: an array of chat messages, each with:
role: one of "system", "user", "assistant"
content: the message text
Typical structure:… See the full description on the dataset page: https://huggingface.co/datasets/CelesteLove/minecraft_qa_es.reddit_dataset_231
Bittensor Subnet 13 Reddit Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks.
For more… See the full description on the dataset page: https://huggingface.co/datasets/CelestialWandererOfTheVoid/reddit_dataset_231.cleand_sequelbox_Celestia3-DeepSeek-R1-0528元データ: https://huggingface.co/datasets/sequelbox/Celestia3-DeepSeek-R1-0528
データ件数: 88,443
平均トークン数: 2143
最大トークン数: 31,680
合計トークン数: 189,577,005
ファイル形式: JSONL
ファイルサイズ: 812.4 MB
x_dataset_231
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/CelestialWandererOfTheVoid/x_dataset_231.auryn_dpo_orpo_englishmeu_site_juridico_perguntas_e_respostas
Dataset Card for "meu_site_juridico_perguntas_e_respostas"
More Information needed
enunciados_pge_rj_orpox_dataset_190
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/CelestialWandererOfTheVoid/x_dataset_190.auryn_dpo_orpo
