datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tokenization-multiplicity-data
Dataset: Tokenization Multiplicity Leads to Arbitrary Price Variation in LLM-as-a-service
This dataset contains the official experiment inference traces for the paper Tokenization Multiplicity Leads to Arbitrary Price Variation in LLM-as-a-service by Ivi Chatzi, Nina Corvelo Benz, Stratis Tsirtsis and Manuel Gomez-Rodriguez.
📂 Dataset Structure
The dataset is organized into folders as follows:
.\{model}\{task}\{lang}\{seed}_{10*temperature}.jsonl
where {model}… See the full description on the dataset page: https://huggingface.co/datasets/Human-Centric-Machine-Learning/tokenization-multiplicity-data.axis-ego-centric-data-industrial-samplesData-Centric-Visual-AI-Challenge-Train-Set
Dataset Card for Data-Centric-Visual-AI-Train-Set
This is a FiftyOne dataset with 30,000 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
import fiftyone.utils.huggingface as fouh
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = fouh.load_from_hub("Voxel51/Data-Centric-Visual-AI-Challenge-Train-Set")
# Launch the App
session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/Data-Centric-Visual-AI-Challenge-Train-Set.strategic-ttc-data
Dataset: Strategic Test-Time Compute (TTC)
This dataset contains the official experiment inference traces for the paper "Test-Time Compute Games" (arXiv:2601.21839).
It includes full model generations, token counts, and correctness verifications for various Large Language Models (LLMs) across three major reasoning benchmarks: GSM8K, AIME, and GPQA.
This data allows researchers to analyze the relationship between test-time compute and model performance without needing to re-run… See the full description on the dataset page: https://huggingface.co/datasets/Human-Centric-Machine-Learning/strategic-ttc-data.ego-centric-sample-dataset
Ego-Centric Sample Dataset
This dataset contains synchronized stereo egocentric videos, hand/head pose data, LeRobot-formatted data tables, and metadata for office manipulation tasks provided by Digital Divide Data.
📑 Dataset Summary
Task Domain: Egocentric Robotics & Vision-based Manipulation
Video Streams: Synchronized Stereo Views (ego_left.mp4, ego_right.mp4)
Tracking & Pose Data: High-frequency pose and hand tracking (.h5)
Tabular Metadata: Standardized… See the full description on the dataset page: https://huggingface.co/datasets/Digital-Divide-Data/ego-centric-sample-dataset.data-centric-ml-sft
Data Centric Machine Learning Domain SFT dataset
The Data Centric Machine Learning Domain SFT dataset is an example of how to use distilabel to create a domain-specific fine-tuning dataset easily.
In particular using the Domain Specific Dataset Project Space.
The dataset focuses on the domain of data-centric machine learning and consists of chat conversations between a user and an AI assistant.
Its purpose is to demonstrate the process of creating domain-specific… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/data-centric-ml-sft.Data-Centric-AI-Demo-Dataset
📖 UNDER CONSTRUCTION!!!
🧩 EleMo-SFT-Playground: Data-Centric AI Demo Dataset
Hinweis: Dies ist ein didaktischer Sandbox-Datensatz (Demo). Er enthält eine stark limitierte, handverlesene Auswahl an Trainingsdaten, um den Prozess des Supervised Fine-Tunings (SFT) für das Elementarpädagogische Modell (EleMo) transparent und nachvollziehbar zu machen.
📖 Dataset Summary
Dieser Datensatz dient als Lernwerkzeug, um zu demonstrieren, wie… See the full description on the dataset page: https://huggingface.co/datasets/Earlychildhoodeducation/Data-Centric-AI-Demo-Dataset.
