datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
carbon-cpu-enriched-sequences
carbon-cpu-enriched-sequences
A CPU-enriched subset of the carbon pretraining corpus (eukaryote_generator), combining original source fields with normalized sequences
and row-level features for quality analysis, GPU enrichment and embedding generation.
Information of Features
Feature
Type
Description
record_id
string
NCBI Identifier linking the row back to the source genomic record. It provides the primary record-level identity.
begin_of_sequence… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-cpu-enriched-sequences.carbon-cpu-enriched-sequences-sampledMixSub-LLaMA-3.2-Text-Only-Overlap-CPU-Scoreharmony-nemotron-cpu-artifacts
Harmony CPU artifacts: Nemotron datasets (normalized + candidate pools)
This dataset repo is an artifact store produced on an EPYC CPU box. It contains:
normalized/ — CPU-normalized Parquet shards with a text-first Harmony format (text) plus meta_* and quality_* fields.
pools/ — candidate pool Parquet shards (subsets) for later GPU scoring (Modal NLL/PPL). No GPU scoring has been run yet.
reports/ — summary tables of counts per dataset/split/pool.
Directory layout… See the full description on the dataset page: https://huggingface.co/datasets/radna0/harmony-nemotron-cpu-artifacts.final_test_stream_encoding_mac_midhres_cpuThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 4,
"total_frames": 3596,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:4"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/imstevenpmwork/final_test_stream_encoding_mac_midhres_cpu.final_test_stream_encoding_linux_highres_cpuThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 4,
"total_frames": 3529,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:4"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/imstevenpmwork/final_test_stream_encoding_linux_highres_cpu.final_test_stream_encoding_mac_highres_cpuThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 4,
"total_frames": 3524,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:4"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/imstevenpmwork/final_test_stream_encoding_mac_highres_cpu.metatree_cpu_act
Dataset Card for "metatree_cpu_act"
More Information needed
metatree_cpu_small
Dataset Card for "metatree_cpu_small"
More Information needed
final_test_stream_encoding_linux_midres_cpuThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 4,
"total_frames": 3567,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:4"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/imstevenpmwork/final_test_stream_encoding_linux_midres_cpu.skewer_luncheon_jan_28_experiment_cpuThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so107_follower",
"total_episodes": 1,
"total_frames": 820,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/thewisp/skewer_luncheon_jan_28_experiment_cpu.MixSub-LLaMA-3.2-Entities-Overlap-CPU-Scoremodel-cards-entities-1k-cpu
davanstrien/model-cards-entities-1k-cpu
Bootstrap NER dataset produced by urchade/gliner_multi-v2.1 over librarian-bots/model_cards_with_metadata.
Generated using uv-scripts/gliner/extract-entities.py.
Provenance
Source dataset
librarian-bots/model_cards_with_metadata (split train)
Text column
card
Bootstrap model
urchade/gliner_multi-v2.1
Entity types
Person, Organization, Dataset, Model, Framework
Confidence threshold
0.6
Samples processed… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/model-cards-entities-1k-cpu.
