datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tcl2-disk-archive
tcl2-disk-archive
Redundant data archived from the tcl2 Vast box before local deletion.
Status: placeholder (2026-09-15). Content is being added in verified batches.
Layout (planned)
MANIFEST.jsonl - one JSON line per archived item: local_path, repo, path_in_repo, bytes, sha256, n_files, encrypted, verified_remote, verified_download, deleted_utc.
Tar archives of PNG trees (per-file md5 lists kept in the manifest side files).
*.tar.enc - third-party-derived data… See the full description on the dataset page: https://huggingface.co/datasets/MingzhenL/tcl2-disk-archive.diskos_qa
Dataset Card for DISKOS-QA
Dataset Summary
DISKOS-QA is an open benchmark for question answering in subsurface and petroleum-domain workflows. It was developed in the FORCE ecosystem and is built from public DISKOS-related oil and gas documents. The broader project uses a Neo4j knowledge graph, topic-based retrieval, Azure OpenAI models, and DeepEval-based filtering to generate and score high-quality question-answer pairs.
The public benchmark is distributed as a tabular… See the full description on the dataset page: https://huggingface.co/datasets/porestar/diskos_qa.ShareGPT52K
Dataset Card for ShareGPT52K90K
Dataset Summary
This dataset is a collection of approximately 52,00090,000 conversations scraped via the ShareGPT API before it was shut down.
These conversations include both user prompts and responses from OpenAI's ChatGPT.
This repository now contains the new 90K conversations version. The previous 52K may
be found in the old/ directory.
Supported Tasks and Leaderboards
text-generation
Languages… See the full description on the dataset page: https://huggingface.co/datasets/diskrizz/ShareGPT52K.disk_monitordisk monitor data
diskos_conocophillips_50this is a dataset
malayalam_disk_dataset_17_01_25diskusi_pilihan_perpajakandisky-v1SosatHuiVse
