datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HaluEvalFineWeb-Edu-10B-PMI-Filteredtrue-falsePMIndiaSum
Dataset Card for "PMIndiaSum"
Dataset Description
Summary
PMIndiaSum is a new multilingual and massively parallel headline summarization corpus focused on languages in India. Our corpus covers four language families, 14 languages, and the largest to date, 196 language pairs. It provides a testing ground for all cross-lingual pairs.
Supported tasks
Monolingual, multilingual and cross-lingual summarization for languages in India.
Languages… See the full description on the dataset page: https://huggingface.co/datasets/PMIndiaData/PMIndiaSum.NQ-Swapnist-gdt-pmi-vlm-benchmark
NIST GD&T/PMI VLM Benchmark
This benchmark measures exact-match transcription of geometric dimensioning and tolerancing (GD&T) and product and manufacturing information (PMI) from rendered NIST Fully-Toleranced Test Case drawing pages. A row supplies the page image and target element_id; the expected output is one engineering-significant specification string.
The reported evaluation uses open transcription: image plus element_id.
page_answer_choices is included for anyone who… See the full description on the dataset page: https://huggingface.co/datasets/CLARKBENHAM/nist-gdt-pmi-vlm-benchmark.averitecinverse-scalingVQAv2hl-feverFineWeb-Edu-10B-PMI-Filtered-DeletePMIndia-SpeechDataset-2004-2024
🇮🇳 PMIndia-SpeechDataset-2004-2024
A curated dataset of official speeches delivered by the Prime Ministers of India from 2004 to 2024, covering the terms of:
Dr. Manmohan Singh (2004–2014)
Shri Narendra Modi (2014–2024)
These speeches reflect key moments in India’s domestic and international policies, national celebrations (like Independence Day), public addresses, and government program launches.
📂 Dataset Overview
Languages: English, Hindi
Time Range: 2004 to… See the full description on the dataset page: https://huggingface.co/datasets/anirudhsankar/PMIndia-SpeechDataset-2004-2024.africa-synth-economic-indicators-manufacturing-pmi-data-all
African Manufacturing Purchasing Managers' Index (PMI) Data | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-economic-indicators-manufacturing-pmi-data-all.7B_iter2_pmistral_N1_random_pair_new_v2tokenizer7B_iter1_pmistral_N1_sft_v01pminervini_NQ_Swap_org_answer_question_openai_google_gemma_2_9b
pminervini_NQ_Swap
Dataset Description
This dataset contains evaluation results for pminervini_NQ_Swap with label column org_answer, with various model performance metrics and samples.
Dataset Summary
The dataset contains original samples from the evaluation process, along with metadata like model names, input columns, and scores. This helps with understanding model performance across different tasks and datasets.
Features
id: Unique identifier for… See the full description on the dataset page: https://huggingface.co/datasets/gallifantjack/pminervini_NQ_Swap_org_answer_question_openai_google_gemma_2_9b.7b_pmistral_iter4_1_new_v2tokenizershroom7B_iter2_pmistral_N1_random_pair7B_iter2_mask_kto_pmistral_N1_random_pair_new_v2tokenizer7b_pmistral_iter3_1_new_v2tokenizerpminervini_NQ_Swap_sub_answer_question_openai_google_gemma_2_9b
pminervini_NQ_Swap
Dataset Description
This dataset contains evaluation results for pminervini_NQ_Swap with label column sub_answer, with various model performance metrics and samples.
Dataset Summary
The dataset contains original samples from the evaluation process, along with metadata like model names, input columns, and scores. This helps with understanding model performance across different tasks and datasets.
Features
id: Unique identifier for… See the full description on the dataset page: https://huggingface.co/datasets/gallifantjack/pminervini_NQ_Swap_sub_answer_question_openai_google_gemma_2_9b.7B_iter3_pmistral_N1_random_pair_new_v2tokenizerpminervini_NQ_Swap_sub_answer_question
pminervini_NQ_Swap
Dataset Description
This dataset contains evaluation results for pminervini_NQ_Swap with label column sub_answer, with various model performance metrics and samples.
Dataset Summary
The dataset contains original samples from the evaluation process, along with metadata like model names, input columns, and scores. This helps with understanding model performance across different tasks and datasets.
Features
id: Unique identifier for… See the full description on the dataset page: https://huggingface.co/datasets/gallifantjack/pminervini_NQ_Swap_sub_answer_question.7B_iter1_pmistral_N1_random_pairpminervini_NQ_Swap_org_answer_question
pminervini_NQ_Swap
Dataset Description
This dataset contains evaluation results for pminervini_NQ_Swap with label column org_answer, with various model performance metrics and samples.
Dataset Summary
The dataset contains original samples from the evaluation process, along with metadata like model names, input columns, and scores. This helps with understanding model performance across different tasks and datasets.
Features
id: Unique identifier for… See the full description on the dataset page: https://huggingface.co/datasets/gallifantjack/pminervini_NQ_Swap_org_answer_question.PMID_CITED_forKGDataset collected from PGB: A PubMed Graph Benchmark for Heterogeneous Network Representation Learning
Description :
inbound_citation: List List of PMID that cites the paper
outbound_citation: List References of the paper
PMID : Pubmed ID
7B_iter1_pmistral_N1_random_pair_new7B_iter4_pmistral_N1_random_pair_new_v2tokenizer7B_iter1_pmistral_N1_random_pair_new_v01
