datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
optimot-linguistic-data
Optimot Linguistic Data
This dataset contains 4,011 entries extracted from the public Optimot linguistic consultation service of the Departament de Política Lingüística, Generalitat de Catalunya.
Each record addresses a Catalan language question or linguistic topic and includes an explanation, source metadata, and a direct source URL when available.
Data
The dataset is provided as JSON Lines:
optimot.jsonl
Each row contains:
Fitxa: Optimot card identifier.… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/optimot-linguistic-data.health-optimization-bench-sample
Health Optimization Bench (Sample)
A 30-task public sample of Health Optimization Bench,
a rubric-graded benchmark measuring how well frontier language models handle current clinical
evidence in preventive and optimization medicine. Three tasks from each of the benchmark's ten
micro benches.
The full benchmark is 977 authored tasks with 346 released across ten micro benches. On the
current leaderboard no model scores above 71 of 100 and the field spans 66 points. Rankings:… See the full description on the dataset page: https://huggingface.co/datasets/Arcophos/health-optimization-bench-sample.MedQA-USMLE-4-options-hfsynthetic-code-optimization-1synthetic-code-optimization-1 is a synthetic dataset with a total of ~1136 Question and Answer pairs.
This dataset was generated using the following models:
ChatGPT:
Whatever is hosted on their website
Claude:
Fable 5
Deepseek:
Deepseek "Instant"
Deepseek "Expert"
Gemini:
3.1 Flash Lite
3.5 Flash
3.1 Pro
Grok:
Fast
Mistral:
Thinking enabled
Qwen 3.7 Plus:
Thinking enabled
GLM 5.2:
Thinking "high"
Perplexity:
Whatever is on their website
This dataset follows the following… See the full description on the dataset page: https://huggingface.co/datasets/takenusername32/synthetic-code-optimization-1.qa-dataset-20250411qa-dataset-20250430adaption-marketing-optimized-neural-titans
Adaption Marketing Optimized Dataset - Neural Titans
Competition: Adaption AutoScientist Challenge ($50,000 Prize Pool)Track: MarketingTeam: Neural Titans (HackIndia)
Dataset Details
Metric
Value
Rows
5,000
Size
22.5 MB
Format
JSONL (instruction-tuning)
Pipeline Configuration
Recipes Applied
Deduplication - Removes duplicate and near-duplicate entries
Prompt Rephrasing - Diversifies prompt formulations for robust… See the full description on the dataset page: https://huggingface.co/datasets/rishini/adaption-marketing-optimized-neural-titans.
