datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vlite7-mini-120m-dataset
ISAI - 이사이
I’m an independent developer building and maintaining AI projects on my own.
Everything from model development to server costs, datasets, and feature updates is managed personally.
Any support you can provide greatly helps keep this project running and allows for continuous improvements.
If you find this project helpful, please consider supporting my work. Thank you.
혼자서 AI 프로젝트를 개발하고 운영하고 있습니다.
모델 개발부터 데이터셋 준비, 서버 비용 감당, 기능 업데이트까지 모두 직접 진행하고 있습니다.
보내주시는 따뜻한 후원은 안정적인… See the full description on the dataset page: https://huggingface.co/datasets/aixk/vlite7-mini-120m-dataset.igsm-med-120Mproblems
Overview
This repository contains datasets to partially reproduce the paper Physics of Language Models: Part 2.1.These are artefacts of our independent reproduction effort.
Models trained on this data are available on Hugginf Face here.
Content
Main training dataset: ./igsm_train_120M
~120 million iGSM problems
difficulties (operations needed to solve each problem): 1-15
variable dependency probe training dataset: ./probes/dep/vprobe_dep_train_20k
dep probe… See the full description on the dataset page: https://huggingface.co/datasets/SimulatedScience/igsm-med-120Mproblems.binance_btcusdt_120M_dollar_klines
BTCUSDT 120M dollar spot klines
This dataset is exported daily from origo.binance_spot_dollar_klines using 120M-dollar dollar bars derived from the 1M Origo dollar-kline foundation.
Latest snapshot:
file: btcusdt_120M_dollar_kline_20200101_to_20260923.parquet
start date: 2020-01-01
rows: 47329
end date: 2026-09-23
columns: start_datetime, end_datetime, dollar_bar_id, open, high, low, close, mean, std, volume, maker_ratio, no_of_trades, open_liquidity, high_liquidity… See the full description on the dataset page: https://huggingface.co/datasets/vaquum/binance_btcusdt_120M_dollar_klines.objaverse_120M_avg_embeddingsobjaverse_120m_parquetsCapsFusion-120Mobjaverse_120M_aes_predsMS-MARCO-0G-120M
MS-MARCO-0G-120M
A long-context retrieval and question-answering benchmark with 1,143,371 documents and a repaired 192-question default evaluation (128 + 64), including 12 questions about 0G. The original 10,000-question construction pool (including 50 0G questions) is also retained. It extends the ms_100M bank from MSA-RAG-BENCHMARKS with 20 million additional text tokens and 655 new questions grounded in the added documents.
120M is the benchmark's nominal size. The actual… See the full description on the dataset page: https://huggingface.co/datasets/0G-AI/MS-MARCO-0G-120M.binance_btcusdt_spot_agg_120M_klines
BTCUSDT 120M dollar spot aggregate klines
This dataset is exported daily from origo.binance_spot_aggtrades_dollar_current using 120M-dollar dollar bars derived from the 1M Origo dollar-kline foundation.
Latest snapshot:
file: btcusdt_spot_agg_120M_kline_20200101_to_20260923.parquet
start date: 2020-01-01
rows: 47329
end date: 2026-09-23
columns: start_datetime, end_datetime, dollar_bar_id, open, high, low, close, mean, std, volume, maker_ratio, no_of_trades, open_liquidity… See the full description on the dataset page: https://huggingface.co/datasets/vaquum/binance_btcusdt_spot_agg_120M_klines.binance_btcusdt_perp_120M_klines
BTCUSDT 120M dollar perp klines
This dataset is exported daily from origo.binance_perp_trades_dollar_current using 120M-dollar dollar bars derived from the 1M Origo dollar-kline foundation.
Latest snapshot:
file: btcusdt_perp_120M_kline_20200101_to_20260923.parquet
start date: 2020-01-01
rows: 268972
end date: 2026-09-23
columns: start_datetime, end_datetime, dollar_bar_id, open, high, low, close, mean, std, volume, maker_ratio, no_of_trades, open_liquidity, high_liquidity… See the full description on the dataset page: https://huggingface.co/datasets/vaquum/binance_btcusdt_perp_120M_klines.qorva-120m-shardssilverwing-120m-corpus-v1AMPLIFY_120M_embeddings_tempisttsai4-120m-datasetisttsai2.5-120m-datasetisttsai2.1-120m-dataset
ISAI - 이사이
I’m an independent developer building and maintaining AI projects on my own.
Everything from model development to server costs, datasets, and feature updates is managed personally.
Any support you can provide greatly helps keep this project running and allows for continuous improvements.
If you find this project helpful, please consider supporting my work. Thank you.
혼자서 AI 프로젝트를 개발하고 운영하고 있습니다.
모델 개발부터 데이터셋 준비, 서버 비용 감당, 기능 업데이트까지 모두 직접 진행하고 있습니다.
보내주시는 따뜻한 후원은 안정적인… See the full description on the dataset page: https://huggingface.co/datasets/aixk/isttsai2.1-120m-dataset.isttsai7-120m-datasetisttsai8-120m-dataset
