datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nectar
Dataset Card for Nectar
Developed by: Banghua Zhu * , Evan Frick * , Tianhao Wu * , Hanlin Zhu and Jiantao Jiao.
License: Apache-2.0 license under the condition that the dataset is not used to compete with OpenAI
Nectar is the first high-quality 7-wise comparison dataset, generated through GPT-4-based ranking. Nectar contains diverse chat prompts, high-quality and diverse responses, and accurate ranking labels. Nectar's prompts are an amalgamation of diverse sources, including… See the full description on the dataset page: https://huggingface.co/datasets/berkeley-nest/Nectar.nesteo-prototype
NestEO: Modular and Hierarchical EO Dataset Framework
NestEO is a hierarchical, resolution-aligned, UTM-based nested grid dataset framework supporting general-purpose, multi-scale multimodal Earth Observation workflows. Built from diverse EO sources and enriched with metadata for landcover, climate zones, and population, it enables scalable, representative and progressive sampling for AI4EO.
Grid Levels: 120000m, 12000m, 2400m, 1200m, 600m, 300m, 150mGrid Metadata: ESA WorldCover… See the full description on the dataset page: https://huggingface.co/datasets/nesteo-datasets/nesteo-prototype.nestful
NESTFUL: Nested Function-Calling Dataset
NESTFUL is a benchmark to evaluate LLMs on nested sequences of API calls, i.e., sequences where the output of one API call is passed as input to
a subsequent call.
The NESTFUL dataset includes over 1800 nested sequences from two main areas: mathematical reasoning and coding tools. The mathematical reasoning portion is generated from
the MathQA dataset, while the coding portion is generated from the
StarCoder2-Instruct dataset.
All… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/nestful.nestle1904-quotation-refs
NuBerea Nestle 1904 Quotation References
Quotation reference annotations mapping Old Testament quotations cited in the New Testament to their original source locations, extracted from the Nestle 1904 Greek New Testament critical apparatus. Each entry records where a New Testament passage cites the Old Testament, together with the apparatus's canonical source-location reference — supporting work in intertextuality, reception history, and the New Testament's use of the Old… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/nestle1904-quotation-refs.berkeley-nest-Nectar-DPOSource berkeley-nest/Nectar
nestedclinbr
NestedClinBr Corpus
NestedClinBr is a new corpus containing nested and discontinuous entities in Brazilian Portuguese clinical narratives.
The main goal of NestedClinBr is to provide a human-annotated corpus that can be used for learning and evaluating different machine learning models to extract valuable medical information in the Portuguese language, in special nested and discontinuous entities, an important but less explored task.
In the context of clinical NLP, the recognition… See the full description on the dataset page: https://huggingface.co/datasets/pucpr-br/nestedclinbr.beir-minus-nanobeir-queries-random-nested-subsetsNestedCiphersnesting-tasks-2d
Nesting Tasks Dataset for 2D Nesting Efficiency Estimation
This is the official Hugging Face Hub version of the 2D Nesting Tasks Dataset originally published on Zenodo (DOI: 10.5281/zenodo.7030786), in 2022, during the my PhD research.
👥 Authors & Affiliations
Corentin Lallier (University of Bordeaux / @ Lectra) — 🎓 Google Scholar | 💼 LinkedIn | 💻 GitHub | 🆔 ORCID
Laurent Vézard (Data Science Manager @ Lectra) — 🎓 Google Scholar | 💼 LinkedIn
Bruno Pinaud… See the full description on the dataset page: https://huggingface.co/datasets/clallier/nesting-tasks-2d.PrimeVulentero_genes_longdyck3_fully_nestedCNEC2_0_nestedlabel_names = [
'O',
'B-P', 'I-P', 'B-T', 'I-T', 'B-A', 'I-A', 'B-C', 'I-C',
'B-ah', 'I-ah', 'B-at', 'I-at', 'B-az', 'I-az',
'B-g_', 'I-g_', 'B-gc', 'I-gc', 'B-gh', 'I-gh',
'B-gl', 'I-gl', 'B-gq', 'I-gq', 'B-gr', 'I-gr',
'B-gs', 'I-gs', 'B-gt', 'I-gt', 'B-gu', 'I-gu',
'B-i_', 'I-i_', 'B-ia', 'I-ia', 'B-ic', 'I-ic',
'B-if', 'I-if', 'B-io', 'I-io', 'B-me', 'I-me'… See the full description on the dataset page: https://huggingface.co/datasets/stulcrad/CNEC2_0_nested.CNEC1_1_Supertypes_nestedberkeley-nest__Starling-LM-7B-alpha-details
Dataset Card for Evaluation run of berkeley-nest/Starling-LM-7B-alpha
Dataset automatically created during the evaluation run of model berkeley-nest/Starling-LM-7B-alpha
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/berkeley-nest__Starling-LM-7B-alpha-details.nesting-optimizer-benchmark
Nesting Optimizer Benchmark
Benchmark results for NestForge 2D nesting strategies across MDF, glass, and steel use cases.
Reference Dataset
Built on synthetic instances aligned with clallier/nesting-tasks-2d geometry distributions.
Strategies Evaluated
Family
Methods
Heuristic
Bottom Left, First Fit, Guillotine, Skyline
Optimization
CP-SAT, MIP
Metaheuristic
GA, Simulated Annealing, Differential Evolution
Neural… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/nesting-optimizer-benchmark.beir-minus-nanobeir-random-nested-subsetsNEST_250323_250411
NEST_250323_250411
NEST (Novel Emerging Segmentation Task) is a benchmark dataset for segmenting (i) novel entities that MLLMs fail to recognize due to their absence from training data, and (ii) emerging entities that exist within the model’s knowledge but demand up-to-date external information for accurate recognition, introduced in the CVPR 2026 Findings paper ROSE: Retrieval-Oriented Segmentation Enhancement.
Dataset Description
NEST targets two categories of… See the full description on the dataset page: https://huggingface.co/datasets/FudanCVL/NEST_250323_250411.RESCAST-100k-NEST
NEST
Processed real residential time-series subset used with RESCAST-100k.
Download
hf download Jainam03/RESCAST-100k-NEST --repo-type dataset --local-dir NEST
License
This processed release is distributed under CC BY 4.0. The original dataset/article is distributed under CC BY 4.0.
Files
timeseries_data/*.parquet
nest_processed_metadata.csv
house_features_nest.parquet
nest_processing_notes.txt
Time-Series Schema
Schema observed from… See the full description on the dataset page: https://huggingface.co/datasets/Jainam03/RESCAST-100k-NEST.nestednestar-ppobeir-minus-nanobeir-random-nested-subsets-embeddingsnested_smallberkeley-nest-Nectarmistral_chat_nesting_datasetdata_prep_2021_12_26___t1_7.csvbertMasked_nestingDatasetsft-ready-berkeley-nest-NectarMASTER_METAPHOR_LIST_NESTED
http://araw.mede.uic.edu/~alansz/metaphor/METAPHORLIST.pdf
Master Metaphor List processed by section using Gemini 2.5 Flash Preview 04-17 based on the prompt below without any post-processing for accuracy.
1. ROLE AND GOAL
You are a specialized AI assistant for parsing academic documents. Your primary objective is to meticulously analyze sections from Lakoff's "Metaphor Master List" and extract the conceptual metaphor data into a structured JSON format. Your processing must be… See the full description on the dataset page: https://huggingface.co/datasets/adelevett/MASTER_METAPHOR_LIST_NESTED.nest-commits
