datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
oercommons-v1-optimized
OERCommons v1 Optimized
Authors: Junjie Wang and Yuhan SunHosted by: PIN TeamDataset: pin-team/oercommons-v1-optimized
OERCommons v1 Optimized is a provenance-preserving multimodal pretraining corpus built on the OERCommons subset of The Common Pile v0.1, which serves as its upstream data and licensing baseline. We extend it with full-page recovery, canonical Markdown, ordered image/PDF/link metadata, conservative corrections, and integrity evidence.
At a glance… See the full description on the dataset page: https://huggingface.co/datasets/pin-team/oercommons-v1-optimized.spro-optimized-prompts-fullOptimized_Video_Facial_Landmarks
Dataset Card for 478-Point Normalized 3D Facial Landmark Dataset
Dataset Description
This dataset provides pre-extracted, normalized 3D facial landmark features derived from the Video Emotion dataset. It is optimized for efficient training of emotion recognition and facial analysis models, bypassing the need to process large raw video files.
License: The extracted feature data in this Parquet file is licensed under Apache 2.0. Note that the original source video files may… See the full description on the dataset page: https://huggingface.co/datasets/PSewmuthu/Optimized_Video_Facial_Landmarks.Optimized_Reasoning
Optimized_Reasoning
SUPPORT ME ON PATREON
https://www.patreon.com/c/Rombodawg
Optimized_Reasoning was created because even modern LLM's are not very good at handling reasoning very well, and if they are, they still waste tons of tokens in the process. With this dataset I hope to accomplish 2 things:
Reduce token usage
Increase model strength in reasoning
So how does this dataset accomplish that? By Adding a "system_prompt" like reasoning tag to the beggining of every data… See the full description on the dataset page: https://huggingface.co/datasets/Rombo-Org/Optimized_Reasoning.optimized-sd-configmcp-server-bench-gradio-optimized
🔬 Gradio vs FastMCP Benchmark Report
Generated: 2026-03-02T13:04:10.857215
Total scenarios: 48
Executive Summary
echo: Fastmcp wins (96.6 vs 176.5 RPS, 1.83x difference)
Gradio best config: concurrency_limit=nan
fibonacci: Fastmcp wins (43.5 vs 57.1 RPS, 1.31x difference)
Gradio best config: concurrency_limit=nan
async_sleep: Gradio wins (93.1 vs 80.2 RPS, 1.16x difference)
Gradio best config: concurrency_limit=nan
payload_echo: Fastmcp wins (82.5 vs 164.8 RPS… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/mcp-server-bench-gradio-optimized.optimized_adult_census
Documentation of the Dataset - Adult Census Income Dataset Optimized
1. General Description of the Dataset
This dataset, called Adult Census Income Dataset Optimized, is an optimized version of the Adult Census Income Dataset. The latter comes from the UCI Machine Learning Repository and is commonly used in classification tasks to predict whether a person earns more or less than $50,000 per year based on various demographic characteristics.
We optimized the dataset by… See the full description on the dataset page: https://huggingface.co/datasets/Databoost/optimized_adult_census.mcp-server-bench-gradio-optimized-full-bench
🔬 Gradio vs FastMCP Benchmark Report
Generated: 2026-03-02T21:04:48.460108
Total scenarios: 337
Executive Summary
echo: Fastmcp wins (100.9 vs 189.9 RPS, 1.88x difference)
Gradio best config: concurrency_limit=1.0
fibonacci: Fastmcp wins (45.5 vs 55.0 RPS, 1.21x difference)
Gradio best config: concurrency_limit=nan
json_transform: Fastmcp wins (94.5 vs 165.4 RPS, 1.75x difference)
Gradio best config: concurrency_limit=5.0
async_sleep: Fastmcp wins (94.0 vs 104.3… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/mcp-server-bench-gradio-optimized-full-bench.medical-o1-reasoning-SFT-jsonl-optimizedllama-2-optimized-product-titles-esci-4-7-temp
Dataset Card for "llama-2-optimized-product-titles-esci-4-7-temp"
More Information needed
Banking-77-OptimizedArabic-Optimized-Reasoning-Dataset
Arabic Optimized Reasoning Dataset
Dataset Name: Arabic Optimized ReasoningLicense: Apache-2.0Formats: CSVSize: 1600 rowsBase Dataset: cognitivecomputations/dolphin-r1Libraries Used: Datasets, Dask, Croissant
Overview
The Arabic Optimized Reasoning Dataset helps AI models get better at reasoning in Arabic. While AI models are good at many tasks, they often struggle with reasoning in languages other than English. This dataset helps fix this problem by:
Using fewer tokens… See the full description on the dataset page: https://huggingface.co/datasets/Jr23xd23/Arabic-Optimized-Reasoning-Dataset.hiring-analyses-optimized_parameters-enllama_2_optimized_product_titles-esci-test-sft
Dataset Card for "llama_2_optimized_product_titles-esci-test-sft"
More Information needed
llama-2-optimized-product-titles-esci-test-sft-temp
Dataset Card for "llama-2-optimized-product-titles-esci-test-sft-temp"
More Information needed
lean-expert-optimized-2000
lean-expert-optimized-2000
Dataset Description
Optimized 2000-example dataset for training Lean trading algorithm optimization agents with 94%+ success rate target.
Dataset Statistics
Total Examples: 2,000
Training Examples: 1800
Validation Examples: 200
Target Success Rate: 94%+
Expected Performance: 96% (94-98% range)
Category Distribution
JSON Parsing: 1,333 examples (CRITICAL - 0% → 95% impact)
Optimization Workflows: 182 examples (HIGH… See the full description on the dataset page: https://huggingface.co/datasets/Kronu/lean-expert-optimized-2000.Optimized_Reasoning_with_textOptimized-MoveToRedBalloptimized_prompts_llama_3llama_2-optimized-titles-esci-sft-test
Dataset Card for "llama_2-optimized-titles-esci-sft-test"
More Information needed
exams-ocr-optimized-testlichao_tree_optimized_dp_seed_merged
Polygon Dynamics 2D
物理推理数据集
项目信息
项目类型: dataset
上传时间: Windows系统
数据集: 包含7种场景类型(A-G)和5个难度等级的物理推理任务
场景类型
A: 基础碰撞检测
B: 重力影响
C: 摩擦力
D: 弹性碰撞
E: 复合物理
F: 高级动力学
G: 极端复杂场景
难度等级
0: 基础 - 简单场景
1: 简单 - 增加复杂度
2: 中等 - 多物体交互
3: 困难 - 复杂物理规则
4: 极端 - 极限测试场景
使用方法
from datasets import load_dataset
# 加载数据集
dataset = load_dataset("competitioncode/lichao_tree_optimized_dp_seed_merged")
环境配置
方法1: 使用.env文件(推荐)
在项目根目录创建.env文件:… See the full description on the dataset page: https://huggingface.co/datasets/competitioncode/lichao_tree_optimized_dp_seed_merged.hiring-analyses-optimized_parameters-ukexebench_loop_optimized_llvm_ir_1sthalfaegis-2-optimized-safety-evalllm-rag-optimized-schema-templates
Schema.org JSON-LD Templates Optimized for LLM RAG Retrieval (2026)
Curated dataset of Schema.org JSON-LD templates designed, tested, and optimized for Retrieval-Augmented Generation (RAG) systems, SearchGPT, Gemini, and Claude search parsers.
Published by Pixel Office EU.
Purpose
Standard Schema.org markup is often too nested or dense for token-efficient LLM context window ingestion. These templates prioritize high-salience fields that crawlers prioritize when… See the full description on the dataset page: https://huggingface.co/datasets/pixeloffice/llm-rag-optimized-schema-templates.orpo-optimized-math-qa
Mathematics Q&A Dataset for ORPO Fine-Tuning
This dataset contains questions and answers focused on various topics within the field of mathematics. The data is optimized for ORPO (Ordered Response Pair Optimization) fine-tuning and includes columns for tags, question bodies, and structured answer formats.
Dataset Overview
The Mathematics Q&A Dataset consists of questions and their corresponding answers from various mathematical topics. This dataset is valuable for… See the full description on the dataset page: https://huggingface.co/datasets/blesspearl/orpo-optimized-math-qa.adult_optimized_v1_top9bhl-240p-60v-optimized
bhl-240p-60v-optimized
OCR-model comparison published with ocrscout. Each row pairs one source page with one model's normalized output.
Summary
Pages: 224
Models: 3
Rows: 718
Pages with errors: 0
Mean disagreement: 0.163
Median disagreement: 0.092
Source images embedded: yes
Generated: 2026-06-27 10:27 UTC by ocrscout v0.1.0
Per-model metrics
Model
Format
Pages OK
Pages errored
Total tokens
Mean tokens/pg
Mean s/pg
Mean prepare s
Mean chars… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/bhl-240p-60v-optimized.llama_2_optimized_product_titles-esci-test
Dataset Card for "llama_2_optimized_product_titles-esci-test"
More Information needed
