datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
repro-accurate-evaluation-of-quickest-changepoint-detectors-via-non-parametric-survival-analysis
Accurate Evaluation of Quickest Changepoint Detectors via Non-parametric Survival Analysis
This is a reproduction logbook for ICML 2026.
OpenReview ID: LhGxRnGmGJ
Paper Abstract
This logbook reproduces KM-ARL and KM-ADD estimators for changepoint detection.
See logbook.json for full claim verification details.
banking77-oumi-quickstartQuickCoder-Dataset
QuickCoder-Dataset
This dataset repository stores upload-ready JSONL training checkpoints for code
completion and fill-in-the-middle training. Checkpoints are appended in
approximately 20 GiB units so they can also be copied to Google Drive and loaded
from Colab/H100 training jobs. New checkpoints use one JSONL file per 20 GiB
checkpoint. The long-term target is 400 GiB total mirrored to Hugging Face and
Google Drive.
Current Upload Status
Only validation-passing… See the full description on the dataset page: https://huggingface.co/datasets/aisamdasu/QuickCoder-Dataset.QuickTrainrank-with-reasonblockrank-msmarco-train-10p
BlockRank MS MARCO Training Data (10% Sample)
Dataset Description
A 10% sample of MS MARCO passage ranking data formatted for training in-context ranking LLMs. This dataset is used in the training of the BlockRank project: Scalable In-context Ranking with Generative Models.
Format: JSONL (in-context ranking format)
Size: 50k training examples (10% sample)
Documents per query: 30-50 candidates (mix of positives and hard negatives)
Source
Original: MS… See the full description on the dataset page: https://huggingface.co/datasets/quicktensor/blockrank-msmarco-train-10p.diffusion_data_constraint_quickstartcustomer-support-synth-quickdemo
expanded_E-commerce_and_SaaS_customer_support_(orders,_shipping,_billing,_subscriptions,_app_troubleshooting)
Synthetic customer-support dataset generated with synth-dataset-kit.
What Is Included
train.jsonl: training split in JSONL chat format
eval_summary.json: minimal evaluation and generation summary
expanded_E-commerce_and_SaaS_customer_support_(orders,_shipping,_billing,_subscriptions,_app_troubleshooting)_quality_report.html: visual quality report… See the full description on the dataset page: https://huggingface.co/datasets/Enkiboy/customer-support-synth-quickdemo.Capybara-Quicksilver-1KCurated from the Capybara dataset, proportionally sampling from each source dataset (2000 data points). Then rewritten by gpt4o, ranked by gpt4o, and the top 1k results are taken.
hermes3-quick-probes-multilingual
Hermes3 Quick Probes (Multilingual, Reasoning ON/OFF)
Міні-датасет (20 прикладів) для швидкої перевірки Hermes-3 у режимах reasoning ON/OFF (UA/ES/EN/ID).
Ціль — легкі sanity-checks: де потрібне міркування, а де достатньо стислої відповіді.
Формат
Файл: data.jsonl, по 1 JSON-об’єкту на рядок з полями:
id (string) — унікальний ідентифікатор
lang (uk|es|en|id)
reasoning ("on"|"off")
prompt (string)
expect (dict, опційно: keywords/max_sentences/answer)
Як… See the full description on the dataset page: https://huggingface.co/datasets/segs/hermes3-quick-probes-multilingual.Titan-L_QuickTrain-v4489,438 Instructions
quickfill-gemmaquicksviewer2_stage1
Quicksviewer2 Stage1 Training Data
This dataset contains training data for Quicksviewer2 Stage1 multimodal pretraining.
Dataset Structure
quicksviewer2_stage1/
├── metadata/
│ ├── llava_recap_558k.jsonl # 558K image-text pairs metadata
│ ├── obelics_50k_train.jsonl # 50K multimodal documents metadata
│ └── activitynet_caption.jsonl # ActivityNet video captions metadata
├── images/
│ ├── llava_recap_558k.tar.gz # 13GB compressed… See the full description on the dataset page: https://huggingface.co/datasets/EdgePro001/quicksviewer2_stage1.conversation-quick-mode
