datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MJ_dataset
Midjourney Dataset
This is a backup of https://huggingface.co/datasets/vivym/midjourney-messages
MegaScale-Tsuboyama2023
MegaScale Tsuboyama 2023 Stability
Advances in DNA sequencing and machine learning are providing insights into protein sequences and structures on an enormous scale. However, the energetics driving folding are invisible in these structures and remain largely unknown. The hidden thermodynamics of folding can drive disease, shape protein evolution and guide protein engineering, and new approaches are needed to reveal these thermodynamics for every sequence and structure. Here we… See the full description on the dataset page: https://huggingface.co/datasets/LiteFold/MegaScale-Tsuboyama2023.Cinematic-DiT-Video-Dataset
Cinematic DiT Video Dataset
This dataset is publicly downloadable under a restricted research license. It
is intended for non-commercial research on AI-generated video detection, media
forensics, authenticity analysis, and content-safety evaluation.
Dataset Summary
This dataset contains 9,000 synthetic text-to-video samples generated from
3,000 Chinese cinematic prompts. Each source prompt has one video in each of
three generation profiles. The prompt pipeline… See the full description on the dataset page: https://huggingface.co/datasets/Tsu7am1/Cinematic-DiT-Video-Dataset.handover-pean-fullbody-palmup-or-room-lerobot-v3-200epThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "nextage_genesis_g1_inspire_pean_handover",
"total_episodes": 200,
"total_frames": 14720,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:200"
},
"data_path":… See the full description on the dataset page: https://huggingface.co/datasets/Michi-Tsubaki/handover-pean-fullbody-palmup-or-room-lerobot-v3-200ep.lekiwi_pick_and_place_block_dayThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
9
],
"names": [
"arm_shoulder_pan.pos",
"arm_shoulder_lift.pos",
"arm_elbow_flex.pos",
"arm_wrist_flex.pos",
"arm_wrist_roll.pos"… See the full description on the dataset page: https://huggingface.co/datasets/tsuyotobi26/lekiwi_pick_and_place_block_day.ahmeduzaki_global-earthquake-tsunami-risk-assessment-dataset
Global Earthquake-Tsunami Risk Assessment Dataset
Seismic Features & Tsunami Classification Dataset for Risk Assessment
Dataset Info
Source: Kaggle
Original Size: 0.02 MB
Kaggle Downloads: 22,533
Files: 1
Files
earthquake_data_tsunami.csv
Mirrored from Kaggle
lekiwi_pick_and_place_pika_earThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
9
],
"names": [
"arm_shoulder_pan.pos",
"arm_shoulder_lift.pos",
"arm_elbow_flex.pos",
"arm_wrist_flex.pos",
"arm_wrist_roll.pos"… See the full description on the dataset page: https://huggingface.co/datasets/tsuyotobi26/lekiwi_pick_and_place_pika_ear.Tsunami-th__Tsunami-0.5-7B-InstructTsunami-th__Tsunami-1.0-7B-InstructTsu_Data
Oral Health & Dental Disease Dataset
Abstract
This dataset provides 30,000 simulated oral health records (10,000 per scenario) from sub-Saharan Africa. Each record contains 40+ variables including dental caries, DMFT score, periodontal disease, noma, oral cancer, treatment access, barriers, and outcomes. Three settings: dental clinic (23% care-seeking), district hospital (16%), and rural health centre (8%).
1. Introduction
Africa bears the largest global… See the full description on the dataset page: https://huggingface.co/datasets/Ontlametse/Tsu_Data.sarm-handkercheif-b-tsuchiya1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so_follower",
"total_episodes": 2,
"total_frames": 3724,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/PID0930/sarm-handkercheif-b-tsuchiya1.Tsunami-th__Tsunami-1.0-14B-Instructseismic-tsunami-event-linkage
Seismic-Tsunami Event Linkage
The Seismic-Tsunami Event Linkage is a comprehensive, machine learning-ready dataset. It contains seismic characteristics and tsunami potential indicators for 782 significant earthquakes recorded globally from 2001 to 2022. This dataset has been specifically processed and structured for applications in tsunami risk prediction, earthquake analysis, and seismic hazard assessment. It is an enhanced version derived from the foundational "Earthquake Dataset"… See the full description on the dataset page: https://huggingface.co/datasets/mnemoraorg/seismic-tsunami-event-linkage.fps_shootingThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 100,
"total_frames": 763,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/tsutsui22/fps_shooting.lekiwi_pick_and_place_pika2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
9
],
"names": [
"arm_shoulder_pan.pos",
"arm_shoulder_lift.pos",
"arm_elbow_flex.pos",
"arm_wrist_flex.pos",
"arm_wrist_roll.pos"… See the full description on the dataset page: https://huggingface.co/datasets/tsuyotobi26/lekiwi_pick_and_place_pika2.lekiwi_pick_and_place_blockThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
9
],
"names": [
"arm_shoulder_pan.pos",
"arm_shoulder_lift.pos",
"arm_elbow_flex.pos",
"arm_wrist_flex.pos",
"arm_wrist_roll.pos"… See the full description on the dataset page: https://huggingface.co/datasets/tsuyotobi26/lekiwi_pick_and_place_block.lekiwi_pick_and_place_block_dark2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
9
],
"names": [
"arm_shoulder_pan.pos",
"arm_shoulder_lift.pos",
"arm_elbow_flex.pos",
"arm_wrist_flex.pos",
"arm_wrist_roll.pos"… See the full description on the dataset page: https://huggingface.co/datasets/tsuyotobi26/lekiwi_pick_and_place_block_dark2.TW-GSAT-Chinese
台灣本土語言模型語料庫:台灣學科能力測驗-中文考科
響應台灣 AI 在地化、AI 十大建設的議題,繁體中文訓練資料最為重要該資料集為 Apache 2.0 開源許可,可用於 商業、研究、私人使用為台灣 AI 在地化盡一份力
使用須知
考試試題之使用符合著作權法
根據中華民國政府的著作權法 - 第9條
下列各款不得為著作權之標的︰ 一、憲法、法律、命令或公文。 二、中央或地方機關就前款著作作成之翻譯物或編輯物。 三、標語及通用之符號、名詞、公式、數表、表格、簿冊或時曆。 四、單純為傳達事實之新聞報導所作成之語文著作。 五、依法令舉行之各類考試試題及其備用試題。
依法舉辦的考試試題是不具備著作權的 適用該條款的考試,包含:學測、會考、學校段考試題,但是不包含補習班、出版商自製的試題 複雜情況:若學校段考考題使用了出版商的題目,那該題目仍然受到著作權的保護,為了規避法律風險,最佳實踐方案是只收集大考考試試題 該資料有調整題目敘述,即重製題目,讓資料更適合 NLP 之任務… See the full description on the dataset page: https://huggingface.co/datasets/TsukiOwO/TW-GSAT-Chinese.tsuyotobi26_lekiwi_pick_and_place_pika_reverseThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
9
],
"names": [
"arm_shoulder_pan.pos",
"arm_shoulder_lift.pos",
"arm_elbow_flex.pos",
"arm_wrist_flex.pos",
"arm_wrist_roll.pos"… See the full description on the dataset page: https://huggingface.co/datasets/tsuyotobi26/tsuyotobi26_lekiwi_pick_and_place_pika_reverse.lekiwi_pick_and_place_pikaThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
9
],
"names": [
"arm_shoulder_pan.pos",
"arm_shoulder_lift.pos",
"arm_elbow_flex.pos",
"arm_wrist_flex.pos",
"arm_wrist_roll.pos"… See the full description on the dataset page: https://huggingface.co/datasets/tsuyotobi26/lekiwi_pick_and_place_pika.lekiwi_pick_and_place_block_darkThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
9
],
"names": [
"arm_shoulder_pan.pos",
"arm_shoulder_lift.pos",
"arm_elbow_flex.pos",
"arm_wrist_flex.pos",
"arm_wrist_roll.pos"… See the full description on the dataset page: https://huggingface.co/datasets/tsuyotobi26/lekiwi_pick_and_place_block_dark.lekiwi_pick_and_place_pika_dayThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
9
],
"names": [
"arm_shoulder_pan.pos",
"arm_shoulder_lift.pos",
"arm_elbow_flex.pos",
"arm_wrist_flex.pos",
"arm_wrist_roll.pos"… See the full description on the dataset page: https://huggingface.co/datasets/tsuyotobi26/lekiwi_pick_and_place_pika_day.lekiwi_pick_and_place_pika_and_block_dayThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
9
],
"names": [
"arm_shoulder_pan.pos",
"arm_shoulder_lift.pos",
"arm_elbow_flex.pos",
"arm_wrist_flex.pos",
"arm_wrist_roll.pos"… See the full description on the dataset page: https://huggingface.co/datasets/tsuyotobi26/lekiwi_pick_and_place_pika_and_block_day.record-testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 30,
"total_frames": 13421,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/tsubamenagata/record-test.lekiwi_pick_and_place_pika0This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
9
],
"names": [
"arm_shoulder_pan.pos",
"arm_shoulder_lift.pos",
"arm_elbow_flex.pos",
"arm_wrist_flex.pos",
"arm_wrist_roll.pos"… See the full description on the dataset page: https://huggingface.co/datasets/tsuyotobi26/lekiwi_pick_and_place_pika0.tsubameoffice2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 50,
"total_frames": 22378,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/tsubamenagata/tsubameoffice2.Tsunami-th__Tsunami-0.5-7B-Instruct-details
Dataset Card for Evaluation run of Tsunami-th/Tsunami-0.5-7B-Instruct
Dataset automatically created during the evaluation run of model Tsunami-th/Tsunami-0.5-7B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Tsunami-th__Tsunami-0.5-7B-Instruct-details.Tsunami-th__Tsunami-0.5x-7B-Instruct-details
Dataset Card for Evaluation run of Tsunami-th/Tsunami-0.5x-7B-Instruct
Dataset automatically created during the evaluation run of model Tsunami-th/Tsunami-0.5x-7B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Tsunami-th__Tsunami-0.5x-7B-Instruct-details.tsubameoffice1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 30,
"total_frames": 13442,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/tsubamenagata/tsubameoffice1.asia-owid-number-of-tsunamis
Number Of Tsunamis | Asia (Our World in Data)
🌏 417 observations · 20 Asia countries · 478–2024 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 417 observations of Number Of Tsunamis data across 20 Asia countries, spanning 478–2024.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Number Of Tsunamis
Geographic coverage
20 Asia countries · top rows shown below… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-number-of-tsunamis.
