datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
keepedit-release-data
KeepEdit Release Data
本仓库存放 KeepEdit 发布版实验数据与中间结果。为避免在 Hugging Face 上平铺十几万个小文件,发布内容采用归档分卷形式;下载后运行解包脚本即可恢复项目根目录下的 data/。
主要内容:
archives/data_processed.tar.000.part
archives/data_candidates.tar.000.part
archives/data_teachers.tar.000.part
...
archives/MANIFEST.sha256
scripts/unpack_release_data_archives.sh 会自动识别分卷,按文件名前缀顺序拼接后解包。
解包后得到:
data/processed/
magicbrush_train/train.jsonl
magicbrush_dev/dev.jsonl
images/
masks/
data/candidates/… See the full description on the dataset page: https://huggingface.co/datasets/Yitaallen/keepedit-release-data.oanda-trading-data
OANDA Trading Data - 10 Year Backfill
Historical forex (FX) candle data collected from OANDA v3 API for machine learning model training.
Dataset Summary
Period: 10 years of historical data
Instruments: EUR_USD, GBP_USD, USD_JPY, AUD_USD, USD_CHF
Granularities: H1 (1-hour candles) - optimized for model training
Total Records: ~310,000 rows (62k rows × 5 instruments)
Format: CSV with OHLCV columns
Data Format
Each CSV file contains:
instrument:… See the full description on the dataset page: https://huggingface.co/datasets/keeprich/oanda-trading-data.R1_keep_source_3_messages_only_final_decontaminatedx402-endpoint-readinessscvd.store x402 endpoint readiness corpus
scvd.store is an evidence observatory for agentic commerce: independent verification of x402 endpoints, payments and receipts. Before an agent pays an x402 endpoint, we check that it can be paid. After it pays, we check the signed receipt. Over time we watch endpoints and publish a dated, signed corpus. Sellers use it to prove a door works; buyers use it before spending. Every artifact is signed, expires, and names what we did not see. Not escrow, not… See the full description on the dataset page: https://huggingface.co/datasets/keeper-scvd/x402-endpoint-readiness.OpenR1-Math-220k-pruned-keep-0.9-end-start-0.5-correctnessOpenR1-Math-220k-pruned-keep-0.5-end-start-0.5keep-it-simple-multimodal
keep-it-simple-multimodal
A mini, standalone multimodal dataset: image+caption, audio+caption, video+caption, lidar, IMU, and
optimal-control state/action pairs. Companion to keep-it-simple
(text), built to feed KairosPretrainingDataset in kairos.
Structure
One generic schema for every row — no per-modality columns, no fixed shape/dtype assumptions:
Column
Type
Description
modality
string
image_caption | audio_caption | video_caption | lidar | imu |… See the full description on the dataset page: https://huggingface.co/datasets/ffurfaro/keep-it-simple-multimodal.vpt_data_8.0_keep_noopThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 14,
"total_frames": 74356,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:14"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/BarryFutureman/vpt_data_8.0_keep_noop.keep-it-simple
keep-it-simple
Objective: An ultra-minimalist dataset for pre-training tiny language models. The logic relies on bidirectional symmetry (A is B and B is A]) to foster deep semantic understanding. By training the model to predict the "prompt" from the "text" and vice versa, we maximize the utility of every pair.
Data Sources
Simple English Wikipedia: Simplified encyclopedic articles.
Vikidia (FR): Educational content for younger audiences.
OPUS Books (en-fr):… See the full description on the dataset page: https://huggingface.co/datasets/ffurfaro/keep-it-simple.OpenR1-Math-220k-pruned-keep-0.75-end-start-1.0OpenR1-Math-220k-pruned-keep-0.75-end-start-0.5OpenR1-Math-220k-pruned-keep-0.5-end-start-0.5-acc-increamentalOpenR1-Math-220k-pruned-keep-0.2-end-start-0.5-accPMB
PMB: a zero-shot benchmark for music understanding in Persian music
13,544 clips (~20 s, 32 kHz mono MP3) of Persian music with labels for
zero-shot evaluation of audio-language models: genre (7 classes),
musical key (24 classes; also tonic-only and mode-only granularities),
emotion (valence 0–100 + 3-class bins; arousal 3-class), tempo
(4 ordered classes), plus reference captions, Spotify popularity, and
artist/song metadata.
genre
clips
pop
5,449
persian rock
3… See the full description on the dataset page: https://huggingface.co/datasets/keepsolid001/PMB.OpenR1-Math-220k-pruned-keep-0.5-end-start-0.5-add-aimeOpenR1-Math-220k-pruned-keep-0.50-1.00OpenR1-Math-220k-pruned-keep-0.1-end-start-1.0OpenR1-Math-220k-pruned-keep-0.75-end-start-0.5-correctnessOpenR1-Math-220k-pruned-keep-0.4-end-start-1.0OpenR1-Math-220k-pruned-keep-0.4-end-start-0.5-correctnessGeneralThought-195K-pruned-keep-0.5-end-start-0.5-cutoffGeneralThought-195K-pruned-keep-0.5-end-start-0.0-cutoffreviewed-qa-keep-discard-pairsOpenR1-Math-220k-pruned-keep-0.5-end-start-0.5-acc-onlyd4rl_adroit_pen_dense_drop_keepThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "adroit",
"total_episodes": 1000,
"total_frames": 57961,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 100,
"splits": {
"train": "0:1000"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/yongjincho/d4rl_adroit_pen_dense_drop_keep.OpenR1-Math-220k-pruned-keep-0.1-end-start-0.0reviewed-qa-keep-discard-tpriftis-pairsd4rl_adroit_pen_dense_drop_keep_clipThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "adroit",
"total_episodes": 1000,
"total_frames": 57961,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 100,
"splits": {
"train": "0:1000"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/yongjincho/d4rl_adroit_pen_dense_drop_keep_clip.keeptrack-run-transcriptsOpenR1-Math-220k-pruned-keep-0.75-end-start-0.0
