datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GreekMMLU
GreekMMLU
GreekMMLU is a native-sourced benchmark for evaluating massive multitask language understanding in Greek, built from authentic Greek exam-style multiple-choice questions (MCQ) rather than machine-translated English benchmarks.
21,805 questions across 45 subjects
4 high-level groups: STEM, Humanities, Social Sciences, Other
Difficulty/education levels spanning Primary → Secondary → University → Professional (+ an N/A bucket)
Public vs. private split for… See the full description on the dataset page: https://huggingface.co/datasets/dascim/GreekMMLU.Institutional-Holdings-Dashboard
📊 Institution Holdings Dashboard
SEC EDGAR 13F filings — cleaned, structured, and ready to use.
42 top hedge funds · 10+ years of history · Weekly auto-updates · Zero auth required
📌 Overview
This dataset contains cleaned, structured institutional holdings data parsed directly from SEC EDGAR 13F-HR filings. It powers a public intelligence platform tracking what the world's top hedge funds are buying and selling — quarter by quarter.Everything in this… See the full description on the dataset page: https://huggingface.co/datasets/Kasher13/Institutional-Holdings-Dashboard.quants
QuAnTS: Question Answering on Time Series
QuAnTS is a challenging dataset designed to bridge the gap in question-answering research on time series data.
The dataset features a wide variety of questions and answers concerning human movements, presented as tracked skeleton trajectories.
QuAnTS also includes human reference performance to benchmark the practical usability of models trained on this dataset.
At present, there is no official leaderboard for this dataset.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/dasyd/quants.FStarDataset-V2-Conversation
F* Proof Completion Dataset (Chat Format)
This dataset is a preprocessed version of microsoft/FStarDataSet-V2. It has been reformatted into a chat-style JSONL structure for supervised fine-tuning of language models on F* function synthesis and proof completion.
Dataset Structure
The dataset consists of three splits:
fstar_train.jsonl
fstar_validation.jsonl
fstar_test.jsonl
Each line in these files is a JSON object with the following schema (where the keys correspond to… See the full description on the dataset page: https://huggingface.co/datasets/dassarthak18/FStarDataset-V2-Conversation.DAS-Mediacal-Red-Teaming-Data
DAS Medical Red-Teaming Test Suites
Accompanies the paper Beyond Benchmarks: Dynamic, Automatic and Systematic Red-Teaming Agents for Trustworthy Medical LLMs.
The data samples presented in this repo are used as the initial data seeds and can be mutated further upon requests. It is designed to stress-test Large Language Models (LLMs) in safety-critical medical domains, auditing along four critical axes: Robustness, Privacy, Bias/Fairness, and Hallucination.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/JZPeterPan/DAS-Mediacal-Red-Teaming-Data.das-dpo-data-searchr1-7b
DAS dpo data-searchr1-7b
This dataset provides DPO preference data for post-training SearchR1 7B search agents with DAS.
The file is provided in LLaMA-Factory compatible DPO format with prompt, chosen, rejected, and optional system fields.
dasd-lol-draft-reasoningspore-protocols
Security Protocols Open Repository (SPORE) Dataset
This dataset contains security protocol specifications formatted for training large language models to understand and reason about cryptographic protocols.
Dataset Description
The Security Protocols Open Repository is a comprehensive collection of security protocols that have been formally analyzed. Each protocol specification includes:
Principal declarations (participants in the protocol)
Cryptographic primitives (keys… See the full description on the dataset page: https://huggingface.co/datasets/dassarthak18/spore-protocols.my-distiset-d78f3f37-das
Dataset Card for my-distiset-d78f3f37-das
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/zastixx/my-distiset-d78f3f37-das/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/zastixx/my-distiset-d78f3f37-das.dge🧠 Awesome ChatGPT Prompts [CSV dataset]
This is a Dataset Repository of Awesome ChatGPT Prompts
View All Prompts on GitHub
License
CC-0
