datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SIFT1B-DiskANN
SIFT-1B Dataset & Disk Index
The SIFT-1B (BigANN) dataset and pre-built disk-based ANN index.
Built February 2026 on Intel Xeon 8462Y+ (Sapphire Rapids) with 800GB RAM.
Build Parameters
Parameter
Value
Dataset
SIFT-1B (1,000,000,000 vectors, 128-dim, uint8)
Graph R
128 (max degree)
Build L
200 (search list size during construction)
PQ chunks
32 (4 dimensions per sub-quantizer)
Build time
~2 days
Files
Raw Data… See the full description on the dataset page: https://huggingface.co/datasets/Nanvivi/SIFT1B-DiskANN.tcl2-disk-archive
tcl2-disk-archive
Redundant data archived from the tcl2 Vast box before local deletion.
Status: placeholder (2026-09-15). Content is being added in verified batches.
Layout (planned)
MANIFEST.jsonl - one JSON line per archived item: local_path, repo, path_in_repo, bytes, sha256, n_files, encrypted, verified_remote, verified_download, deleted_utc.
Tar archives of PNG trees (per-file md5 lists kept in the manifest side files).
*.tar.enc - third-party-derived data… See the full description on the dataset page: https://huggingface.co/datasets/MingzhenL/tcl2-disk-archive.so101_cube_diskThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 12,
"features": {
"observation.images.front": {
"dtype": "video",
"shape": [
480,
640,
3
],
"names": [
"height",
"width",
"channel"
],
"info": {
"video.height":… See the full description on the dataset page: https://huggingface.co/datasets/alexis779/so101_cube_disk.wenhuibao_disk
文汇报光盘1938-1999
光盘收录了1938年-1999年所有文章共计1231692篇。13张光盘中包括扫描的图像和文本数据, setup.iso为安装程序(内含文本数据库),1-12文件夹为原来1-12号光碟。html.7z为爬虫爬取的文章html页面的压缩包。
光盘使用说明
安装程序需要在windows98 简体中文版中打开
进入系统后,插入setup.iso,安装时根据提示插入1-5号光盘;其他光盘在读取插图时使用。
(建议使用VirtualBox虚拟机)
启动爬虫
安装ie6
在安装目录的html/cgi-bin/oneart.htm的""后插入
<SCRIPT language=JavaScript>
var a =function() {
if (document.readyState !="complete") return;
var x = new ActiveXObject("Microsoft.XMLHTTP");
var content =… See the full description on the dataset page: https://huggingface.co/datasets/banned-historical-archives/wenhuibao_disk.v86-disk-imagesdisk_ridgesDiskANN_indexmetaworld_bin_transfer_disk_obstacle_randomized_noise005_3200eps
MetaWorld Bin Transfer Disk Obstacle Randomized Noise005 3200eps
This dataset contains scripted MetaWorld demonstrations for a two-bin colored-block transfer task with a visible green disk between the red and blue bins. The disk is an imaginary obstacle: in avoid-disk tasks the held block is not allowed above the disk, while in no-avoid tasks the same disk remains visible but the prompt explicitly says it does not need to be avoided.
High-Level Facts
HF repo:… See the full description on the dataset page: https://huggingface.co/datasets/ccwatson/metaworld_bin_transfer_disk_obstacle_randomized_noise005_3200eps.zea-rotating-disk
Verasonics ultrasound data (zea format)
This dataset contains raw ultrasound data acquired with Verasonics systems,
converted to the zea HDF5 format.
diskSIFT100M-DiskANN
SIFT-100M Dataset & Disk Index
The SIFT-100M dataset (first 100M vectors of BigANN SIFT-1B) and pre-built disk-based ANN index.
Built February 2026 on Intel Xeon 8462Y+ (Sapphire Rapids).
Build Parameters
Parameter
Value
Dataset
SIFT-100M (100,000,000 vectors, 128-dim, uint8)
Graph R
96 (max degree)
Build L
128 (search list size during construction)
PQ chunks
32 (4 dimensions per sub-quantizer)
Files
Raw Data
File
Size… See the full description on the dataset page: https://huggingface.co/datasets/Nanvivi/SIFT100M-DiskANN.2026-09-09_cheap_shakeit_bench_pzt_diskThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/jogarulfop/2026-09-09_cheap_shakeit_bench_pzt_disk.ru_sentances_pos
Информация о POS-тегах
not - Неизменяемая частица при отрицании
Abbr - Аббревиатуры
Adj - Прилагательное
Adv - Наречие
Adv/action_desс - Наречие образа действия
Adv/action_time - Наречие времени действия
Adv/measure - Наречие меры
Adv/place - Наречие места
Adv/emph - Усилительные наречия
AdvT - Наречия определенного и неопределенного временного периода
AdvT1 - Наречия определенной и неопределенной частоты
Aux - Вспомогательный и связочный глаголы
Bracket - Скобки
Colon - Двоеточие… See the full description on the dataset page: https://huggingface.co/datasets/disk0dancer/ru_sentances_pos.Disk2Planet_noiseDisk2Planet_vanillaso101_cube_disk_scriptedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 12,
"features": {
"observation.images.front": {
"dtype": "video",
"shape": [
480,
640,
3
],
"names": [
"height",
"width",
"channel"
],
"info": {
"video.height":… See the full description on the dataset page: https://huggingface.co/datasets/alexis779/so101_cube_disk_scripted.rollout_2026-09-10_cheap_shakeit_bench_pzt_disk_20260910_160409This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/jogarulfop/rollout_2026-09-10_cheap_shakeit_bench_pzt_disk_20260910_160409.diskos_qa
Dataset Card for DISKOS-QA
Dataset Summary
DISKOS-QA is an open benchmark for question answering in subsurface and petroleum-domain workflows. It was developed in the FORCE ecosystem and is built from public DISKOS-related oil and gas documents. The broader project uses a Neo4j knowledge graph, topic-based retrieval, Azure OpenAI models, and DeepEval-based filtering to generate and score high-quality question-answer pairs.
The public benchmark is distributed as a tabular… See the full description on the dataset page: https://huggingface.co/datasets/porestar/diskos_qa.cohere1536_diskann2026-09-21_cheap_shakeit_bench_3sensors_v2_pzt_disk_nfft_512This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/jogarulfop/2026-09-21_cheap_shakeit_bench_3sensors_v2_pzt_disk_nfft_512.ShareGPT52K
Dataset Card for ShareGPT52K90K
Dataset Summary
This dataset is a collection of approximately 52,00090,000 conversations scraped via the ShareGPT API before it was shut down.
These conversations include both user prompts and responses from OpenAI's ChatGPT.
This repository now contains the new 90K conversations version. The previous 52K may
be found in the old/ directory.
Supported Tasks and Leaderboards
text-generation
Languages… See the full description on the dataset page: https://huggingface.co/datasets/diskrizz/ShareGPT52K.train_clean_diskThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 101,
"total_frames": 76985,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:101"},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Bekhzod/train_clean_disk.so101-put_disk_dualThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 0,
"total_frames": 0,
"total_tasks": 0,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jliu6718/so101-put_disk_dual.DiskStrukt2023-VL
Dataset Card for Dataset Name
Dataset Summary
Embeddings made using transcripts of the Diskrete Strukturen SOSE2023 lectures usable for OpenAis chatGPT and probably other stuff.
Dataset Structure
The Dataset is stored in the form of a Chroma DB
disks45_nocr_trec-robust-2004_fold1
Dataset Card for disks45/nocr/trec-robust-2004/fold1
The disks45/nocr/trec-robust-2004/fold1 dataset, provided by the ir-datasets package.
For more information about the dataset, see the documentation.
Data
This dataset provides:
queries (i.e., topics); count=50
qrels: (relevance assessments); count=62,789
For docs, use irds/disks45_nocr
Usage
from datasets import load_dataset
queries = load_dataset('irds/disks45_nocr_trec-robust-2004_fold1', 'queries')… See the full description on the dataset page: https://huggingface.co/datasets/irds/disks45_nocr_trec-robust-2004_fold1.Grab_the_black_cube_onto_the_red_diskdisks45_nocr_trec-robust-2004_fold5
Dataset Card for disks45/nocr/trec-robust-2004/fold5
The disks45/nocr/trec-robust-2004/fold5 dataset, provided by the ir-datasets package.
For more information about the dataset, see the documentation.
Data
This dataset provides:
queries (i.e., topics); count=50
qrels: (relevance assessments); count=63,841
For docs, use irds/disks45_nocr
Usage
from datasets import load_dataset
queries = load_dataset('irds/disks45_nocr_trec-robust-2004_fold5', 'queries')… See the full description on the dataset page: https://huggingface.co/datasets/irds/disks45_nocr_trec-robust-2004_fold5.disk_monitordisk monitor data
diskos_conocophillips_50this is a dataset
malayalam_disk_dataset_17_01_25
