datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
holobench
HoloBench (Holistic Reasoning Benchmark)
HoloBench is a benchmark designed to evaluate the ability of long-context language models (LCLMs) to perform holistic reasoning over extended text contexts.
Unlike standard models that retrieve isolated information, HoloBench tests how well LCLMs handle complex reasoning tasks that require aggregating and synthesizing information across multiple documents or large text segments.
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/megagonlabs/holobench.megaunscene
Emergent Extreme-View Geometry in 3D Foundation Models
Yiwen Zhang¹ Joseph Tung² Ruojin Cai³ David Fouhey² Hadar Averbuch-Elor¹
¹Cornell University ²New York University ³Kempner Institute, Harvard University
MegaUnScene Benchmark
Overview
MegaUnScene is a dataset of Internet scenes unseen by existing 3DFMs for benchmarking. There are three test splits split across two evaluation tasks:
Relative Pose Estimation: UnScenePairs and UnScenePairs-t… See the full description on the dataset page: https://huggingface.co/datasets/cornell-vailab/megaunscene.megalith-10m-florence2
Megalith-10M with Florence-2 Caption
日本語はこちら
This reposity is the supplymentary of Megalith-10M.
Megalith-10M is an CC-0 like image dataset. However, the dataset does not contain the image caption.
Therefore, we caption the images by Florence 2.
Usage
from datasets import load_dataset
dataset = load_dataset("aipicasso/megalith-10m-florence2")
How to get images
git lfs install
git clone https://huggingface.co/datasets/drawthingsai/megalith-10m… See the full description on the dataset page: https://huggingface.co/datasets/aipicasso/megalith-10m-florence2.MEGA-cleaned-prompts
Cleaned Prompts Mega Dataset
Created by Aipresso LIMITED, London, UK
⚠️ IMPORTANT: By using this dataset, you agree to our Terms of Use
You must provide attribution when using this data in publications, research, or commercial products.
Dataset Overview
A comprehensive collection of 2.7 million cleaned English prompts, meticulously processed for training advanced language models and AI systems.
📊 Dataset Statistics
Metric
Value
Total Rows
2,689… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/MEGA-cleaned-prompts.MegaTempQAMegaVul_simplefake-news-mega-split-4adamgrey88_megastore-sales-data
MegaStore Sales Data
Retail Sales Dataset: Two-Year Record of Transactions for In-Depth Analysis
Dataset Info
Source: Kaggle
Original Size: 8.56 MB
Kaggle Downloads: 4,397
Files: 1
Files
superstoredata.csv
Mirrored from Kaggle
fake-news-mega-split-1adamgrey88_megastore-sales-data
MegaStore Sales Data
Retail Sales Dataset: Two-Year Record of Transactions for In-Depth Analysis
Dataset Info
Source: Kaggle
Original Size: 8.56 MB
Kaggle Downloads: 4,397
Files: 1
Files
superstoredata.csv
Mirrored from Kaggle
kult4e-lieddg_megadatasetfake-news-mega-split-3fake-news-mega-split-5mega-dataset-5000-training-data
mega-dataset-5000 Training Dataset
Dataset humaniste pour l'affinage de modèles IA avec 5000 exemples authentiques.
Contenu
5000 paires question-réponse authentiques
Générées par Mistral AI avec scores d'alignement humaniste
Focus sur les valeurs humanistes : diversité, science, droits, écologie
Utilisation
Ce dataset contient des données réelles pour l'affinage de modèles selon les principes humanistes.
Données
Les données sont stockées dans… See the full description on the dataset page: https://huggingface.co/datasets/peopleofverso/mega-dataset-5000-training-data.fake-news-mega-split-2MegaDatabaseMegalex
[!NOTE]
Dataset origin: https://openlexicon.fr/datasets-info/FrenchLexiconProject/README-FrenchLexiconProject.html
MEGALEX : méga-étude de la reconnaissance des mots écrits et parlés
Megalex provides visual and auditory lexical decision times and accuracy rates several thousands of words: Visual lexical decision data are available for 28466 French words and the same number of pseudowords, and auditory lexical decision data are available for 17876 French words and the same number of… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/Megalex.Tamazight-Mega-Corpusmegasthenes-eval-dataset
