datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nostrhost-agentlibrig2p-nostress-space-cmuGrapheme-to-Phoneme training, validation and test setsrecord-dual-nostream-testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 1,
"total_frames": 796,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/choiwoong/record-dual-nostream-test.record-wrist-nostream-testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 1,
"total_frames": 1479,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/choiwoong/record-wrist-nostream-test.librig2p-nostress-spaceGrapheme-to-Phoneme training, validation and test setscve-kev-snapshot-90d-2025-10-29
CVE-KEV Snapshot (one-time, offline bundle)
This bundle lets you rank likely exploited CVEs and cite official sources without any APIs or accounts, fully offline. It’s a one-time snapshot of the last 90 days of NVD, aligned with CISA KEV for immediate focus on likely exploited CVEs. Query-ready Parquet tables and an optional small RAG pack let you rank by severity, pivot by CWE, and fetch references for briefings. Every row includes provenance; validation metrics and an integrity… See the full description on the dataset page: https://huggingface.co/datasets/NostromoHub/cve-kev-snapshot-90d-2025-10-29.librig2p-nostressGrapheme-to-Phoneme training, validation and test setsNos_Transcrispeech-GL
Corpus description
Manually transcribed and speech-to-text aligned Galician ASR corpus containing 50 hours of multi-domain speech. The corpus contains different types of audios: conferences, debates, speeches, and interviews. The corpus is divided into three partitions: train (80%), dev(10%) and test (10%). Each partition contains audio fragments in WAV format, aligned with their respective transcriptions. In the "/data" folder you can find the audio partitions, while the "/transcript"… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/Nos_Transcrispeech-GL.cve-kev-snapshot-90d-2025-09-04
CVE-KEV Snapshot (one-time, offline bundle)
This bundle lets you rank likely exploited CVEs and cite official sources without any APIs or accounts, fully offline. It’s a one-time snapshot of the last 90 days of NVD, aligned with CISA KEV for immediate focus on likely exploited CVEs. Query-ready Parquet tables and an optional small RAG pack let you rank by severity, pivot by CWE, and fetch references for briefings. Every row includes provenance; validation metrics and an integrity… See the full description on the dataset page: https://huggingface.co/datasets/NostromoHub/cve-kev-snapshot-90d-2025-09-04.1C_code_1000ot3-fineweb-40k-qwen3-nostrip-81921c_code_nanonostradamus-propheties
Dataset Card for "nostradamus-propheties"
Dataset Description
Dataset Summary
The Nostradamus propheties dataset is a set of structured files containing the "Propheties" by Nostradamus, translated in modern English.
The original text consists of 10 "Centuries", every century containing 100 numbered quatrains.
In the dataset, every century is a separate file named century**.json. For instance, all the quatrains of Century I are in the file century01.json.
The… See the full description on the dataset page: https://huggingface.co/datasets/wpicard/nostradamus-propheties.cloudflare-imgbednostr_datasets
