datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ps2_hf1basilbasil-2basil-instances
maximilian-franz/basil-instances
Per-plant-instance segmented crops derived from ['maximilian-franz/basil', 'maximilian-franz/basil-2'], one row per
(original frame × confirmed plant instance).
Layout
ImageFolder-style dataset: metadata.csv at the repo root, with a file_name column pointing
to each row's masked crop under images/<plant_instance_id>/<NNNN>.png, and a
bbox_file_name column pointing to the same row's plain rectangular crop under… See the full description on the dataset page: https://huggingface.co/datasets/maximilian-franz/basil-instances.v1_300hours_maximum_5spks_precompute_25600framesMetaMedQA
MetaMedQA Dataset
Overview
MetaMedQA is an enhanced medical question-answering benchmark that builds upon the MedQA-USMLE dataset. It introduces uncertainty options and addresses issues with malformed or incorrect questions in the original dataset. Additionally, it incorporates questions from the Glianorex benchmark to assess models' ability to recognize the limits of their knowledge.
Key Features
Extended version of MedQA-USMLE
Incorporates uncertainty… See the full description on the dataset page: https://huggingface.co/datasets/maximegmd/MetaMedQA.PolaRiS-Hub
PolaRiS Hub
PolaRiS Hub is a lightweight collection of reconstructed real-to-sim manipulation environments used for evaluating generalist robot policies in simulation.
It is designed to be used with the PolaRiS evaluation framework.
Instructions on how to use can be found at the Github PolaRiS Codebase
What’s in the Hub
Each environment in PolaRiS Hub includes:
A reconstructed scene (mesh + Gaussian splats)
Canonical initial conditions
A language instruction describing… See the full description on the dataset page: https://huggingface.co/datasets/MaximumY/PolaRiS-Hub.basil-instances-archive-2
maximilian-franz/basil-instances
Per-plant-instance segmented crops derived from ['maximilian-franz/basil', 'maximilian-franz/basil-2'], one row per
(original frame × confirmed plant instance).
Layout
ImageFolder-style dataset: metadata.csv at the repo root, with a file_name column pointing
to each row's image under images/<plant_instance_id>/<NNNN>.png. Load with:
from datasets import load_dataset
ds = load_dataset("maximilian-franz/basil-instances")… See the full description on the dataset page: https://huggingface.co/datasets/maximilian-franz/basil-instances-archive-2.timeseries-1m-QQQ-5ysick_nl
Dataset Summary
An automatically translated, manually corrected translation of the SICK dataset of Marelli et al. 2014, intended to boost research in Dutch NLP.
Languages
The dataset is in Dutch.
Dataset Structure
Data Fields
pair_ID: sentence pair ID
sentence_A: sentence A
sentence_B: sentence B
label: textual entailment gold label: entailment (0), neutral (1) or contradiction (2)
relatedness_score: semantic relatedness gold score (on a 1-5… See the full description on the dataset page: https://huggingface.co/datasets/maximedb/sick_nl.basil-instances-archive-3
maximilian-franz/basil-instances
Per-plant-instance segmented crops derived from ['maximilian-franz/basil', 'maximilian-franz/basil-2'], one row per
(original frame × confirmed plant instance).
Layout
ImageFolder-style dataset: metadata.csv at the repo root, with a file_name column pointing
to each row's masked crop under images/<plant_instance_id>/<NNNN>.png, and a
bbox_file_name column pointing to the same row's plain rectangular crop under… See the full description on the dataset page: https://huggingface.co/datasets/maximilian-franz/basil-instances-archive-3.timeseries-QQQ-1d-25yral2basil-segmentation-plantcv
maximilian-franz/basil-segmentation-plantcv
Per-instance segmented basil crops produced by the plantcv backend. This is a Hugging Face ImageFolder dataset: file_name points to the black-background masked crop used for downstream image analysis and bbox_file_name points to the corresponding unmasked rectangular crop. Empty masks are omitted. Bounding boxes use native source-frame coordinates.
Load it with:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/maximilian-franz/basil-segmentation-plantcv.mfaq_lightMQA is a multilingual corpus of questions and answers parsed from the Common Crawl. Questions are divided between Frequently Asked Questions (FAQ) pages and Community Question Answering (CQA) pages.follie
Dataset Card for FOLLIE dataset
Dataset Details
Dataset Description
FOLLIE (First-Order Logic for Language Inference and Entailment) is the first dataset for French with natural language sentences and their corresponding first-order logic (FOL) formulas.
The sentences in the dataset were drawn from all the existing French Natural Language Inference (NLI) datasets, namely DACCORD, FraCaS-FR, GQNLI-FR, RTE3-FR (dev and test), SICK-FR, and XNLI… See the full description on the dataset page: https://huggingface.co/datasets/maximoss/follie.so100_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 54,
"total_frames": 16724,
"total_tasks": 1,
"total_videos": 162,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:54"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/maximilienroberti/so100_test.basil-segmentation-hybrid
maximilian-franz/basil-segmentation-hybrid
Per-instance segmented basil crops produced by the hybrid backend. This is a Hugging Face ImageFolder dataset: file_name points to the black-background masked crop used for downstream image analysis and bbox_file_name points to the corresponding unmasked rectangular crop. Empty masks are omitted. Bounding boxes use native source-frame coordinates.
Load it with:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/maximilian-franz/basil-segmentation-hybrid.so100_lego_red_boxThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 150,
"total_frames": 40711,
"total_tasks": 1,
"total_videos": 450,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:150"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/maximilienroberti/so100_lego_red_box.timeseries-1m-QQQ-10yomx_multicubesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "omx_follower",
"total_episodes": 176,
"total_frames": 137392,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:176"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/maximellerbach/omx_multicubes.task1148_maximum_ascii_value
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1148_maximum_ascii_value
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1148_maximum_ascii_value.lingnli-multi
Dataset Card for Dataset Name
Dataset Summary
This repository contains a collection of machine translations of LingNLI dataset
into 9 different languages (Bulgarian, Finnish, French, Greek, Italian, Korean, Lithuanian, Portuguese, Spanish). The goal is to predict textual entailment (does sentence A
imply/contradict/neither sentence B), which is a classification task (given two sentences,
predict one of three labels). It is here formatted in the same manner as the… See the full description on the dataset page: https://huggingface.co/datasets/maximoss/lingnli-multi.English-Valid-Words
English Valid Words
This repository contains CSV files with valid English words along with their frequency, stem, and stem valid probability.
Dataset Github link: https://github.com/Maximax67/English-Valid-Words
Files included
valid_words_sorted_alphabetically.csv:
N: Counter for each word entry.
Word: The English word itself.
Frequency count: The number of occurrences of the word in the 1-grams dataset.
Stem: The stem of the word.
Stem valid probability: Probability… See the full description on the dataset page: https://huggingface.co/datasets/Maximax67/English-Valid-Words.basil-segmentation-sam3
maximilian-franz/basil-segmentation-sam3
Per-instance segmented basil crops produced by the sam3 backend. This is a Hugging Face ImageFolder dataset: file_name points to the black-background masked crop used for downstream image analysis and bbox_file_name points to the corresponding unmasked rectangular crop. Empty masks are omitted. Bounding boxes use native source-frame coordinates.
Load it with:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/maximilian-franz/basil-segmentation-sam3.multilingual_librispeech_frgrok-demon-dataset-ESmaxima-1-power-spectrum
MAXIMA-1 power spectrum and calibrated map
This dataset contains the served LAMBDA products for the ten-bin MAXIMA-1
angular power spectrum, the calibrated sky map, and their numeric ancillary
products. Cl.txt has
no header or field labels; Column 1, Column 2, and Column 3 are generated
Parquet column names for its three source positions. The calibrated-map FITS
fields are separate configurations named exactly PIXEL RA [DEGREE],
PIXEL DEC [DEGREE], MAP [UK], and NOISE [UK^2].… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/maxima-1-power-spectrum.MAS
MAS: A Millennium of Arabic Manuscripts in Three Styles
MAS (Medieval Arabic Script) is a line-level OCR benchmark of 11,841
expert-annotated lines from nine authentic Arabic manuscript works. It covers
Naskh, Taliq, and Nastaliq. Copies date from the 12th to the 20th century;
compositions from the 10th to the early 20th.
This dataset accompanies the ICDAR 2026 paper
A Millennium of Arabic Manuscripts in Three Styles: A Line-Level OCR Benchmark
for Naskh, Taliq, and Nastaliq.… See the full description on the dataset page: https://huggingface.co/datasets/maximazzik/MAS.glianorex
Multiple Choice Questions and Large Languages Models: A Case Study with Fictional Medical Data
This multiple choice question dataset on a fictional organ, the Glianorex, is used to assess the capabilities of models to answer questions on knowledge they have never encountered.
We only provide a test dataset as training models on this dataset would defeat the purpose of isolating linguistic capabilities from knowledge.
Motivation
We designed this dataset to evaluate the… See the full description on the dataset page: https://huggingface.co/datasets/maximegmd/glianorex.
