datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
refinedweb-embedded_prototypeContaining 768-dimensional embedding vectors derived from the "content" column of the RefinedWeb dataset.
The embeddings were generated using the E5-Base-4k model with a context length of 1024, employing scaled dot-product attention.
Original dataset:
This dataset bases as derivative work on RefinedWeb, an English web dataset created by the TII (Technology Innovation Institute).
Attribution is given to the TII as original authors of the RefinedWeb dataset as per Section 4 of RefinedWeb's Open… See the full description on the dataset page: https://huggingface.co/datasets/Marcus2112/refinedweb-embedded_prototype.nesteo-prototype
NestEO: Modular and Hierarchical EO Dataset Framework
NestEO is a hierarchical, resolution-aligned, UTM-based nested grid dataset framework supporting general-purpose, multi-scale multimodal Earth Observation workflows. Built from diverse EO sources and enriched with metadata for landcover, climate zones, and population, it enables scalable, representative and progressive sampling for AI4EO.
Grid Levels: 120000m, 12000m, 2400m, 1200m, 600m, 300m, 150mGrid Metadata: ESA WorldCover… See the full description on the dataset page: https://huggingface.co/datasets/nesteo-datasets/nesteo-prototype.smart_market_prototype_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "omx_follower",
"total_episodes": 300,
"total_frames": 378133,
"total_tasks": 30,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:300"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kdy93/smart_market_prototype_2.Cantonese-Prototype-Roleplay-Dataset
Prototype-Dataset-Cantonese
角色扮演訓練集(粵語版)
簡介
本數據集是基於 Moemu/Muice-Dataset 進行製作的粵語 (Cantonese) 衍生版本。旨在模擬二次元角色的語言風格,包含日常生活、情感話題、自我認知增強等約 3,500 條對話。
數據處理說明
語言轉換:將原有的簡體中文對話轉換為繁體粵語口語。
格式保持:嚴格遵循原版的多輪對話 JSONL 格式。
限制聲明
日常導向:本訓練集主要圍繞日常話題展開,對專業性問題(如代碼、高級推理)做了簡化處理,模型可能會產生事實性錯誤。
性格特徵:為了模仿動漫角色的說話風格,訓練集可能含有傲嬌、偏見或不禮貌的回答。如需構建高度安全性的模型,請謹慎使用或進行篩選。
倫理風險:使用本數據集訓練模型所引起的任何法律或倫理風險,由訓練者自行承擔。
許可與鳴謝
原作者: Moemu
粵語版本製作: AhYin
許可證:… See the full description on the dataset page: https://huggingface.co/datasets/AhYin/Cantonese-Prototype-Roleplay-Dataset.Material_prototypesmart_market_prototype_4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "omx_follower",
"total_episodes": 20,
"total_frames": 21971,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kdy93/smart_market_prototype_4.umi_prototype_prototype_0717_50epThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
16
],
"names": [
"left_ee.x",
"left_ee.y",
"left_ee.z",
"left_ee.qx",
"left_ee.qy",
"left_ee.qz"… See the full description on the dataset page: https://huggingface.co/datasets/rcvsungmincho/umi_prototype_prototype_0717_50ep.smart_market_prototype_3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "omx_follower",
"total_episodes": 60,
"total_frames": 116807,
"total_tasks": 6,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:60"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kdy93/smart_market_prototype_3.umi_prototype_prototype_0716_test_3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
16
],
"names": [
"left_ee.x",
"left_ee.y",
"left_ee.z",
"left_ee.qx",
"left_ee.qy",
"left_ee.qz"… See the full description on the dataset page: https://huggingface.co/datasets/rcvsungmincho/umi_prototype_prototype_0716_test_3.umi_prototype_prototype_0716_test_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
16
],
"names": [
"left_ee.x",
"left_ee.y",
"left_ee.z",
"left_ee.qx",
"left_ee.qy",
"left_ee.qz"… See the full description on the dataset page: https://huggingface.co/datasets/rcvsungmincho/umi_prototype_prototype_0716_test_v1.umi_prototype_prototype_0717_50ep_v3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
16
],
"names": [
"left_ee.x",
"left_ee.y",
"left_ee.z",
"left_ee.qx",
"left_ee.qy",
"left_ee.qz"… See the full description on the dataset page: https://huggingface.co/datasets/rcvsungmincho/umi_prototype_prototype_0717_50ep_v3.prototype-jenga-pickup-longThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 1,
"total_frames": 612,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/apinguen/prototype-jenga-pickup-long.Filtered_Reasoning_PrototypeThe most repetitive adjectives and other nonsense was fixed.
update:
format
pattern = r"[\d+]"
Token length distribution (for a similar dataset with 0.5% additional vulume of profane questions being answered with spiritual guidance):
Token Range | Line Count
----------+------
450-500 | 3
500-550 | 27
550-600 | 219
600-650 | 939
650-700 | 3242
700-750 | 7107
750-800 | 11357
800-850 | 13732
850-900 | 13403
900-950 | 11152
950-1000 | 8063
1000-1050 |… See the full description on the dataset page: https://huggingface.co/datasets/EsotericsEnjoyer/Filtered_Reasoning_Prototype.cell1_20260512_electronics-packing_esp-cable-2servoblue-bb-jumpersff-mm-prototypeboarThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "starpilot_yam_gripper",
"total_episodes": 5,
"total_frames": 15396,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/starpilot-ai/cell1_20260512_electronics-packing_esp-cable-2servoblue-bb-jumpersff-mm-prototypeboar.smart_market_prototype_1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "omx_follower",
"total_episodes": 173,
"total_frames": 231238,
"total_tasks": 6,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:173"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kdy93/smart_market_prototype_1.multiapi_prototype_VT_dec11Food-Prototype
Dataset Card for "Food-Prototype"
More Information needed
cogbench-learning-tasksFood-Prototype-Bruce
Dataset Card for "Food-Prototype-Bruce"
More Information needed
prototypeqwen_prototype_10_3ppIntent
