datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cncf-raw-data-for-llm-training
CNCF Raw Data for LLM Training
Description
This dataset, named cncf-raw-data-for-llm-training, consists of markdown (MD) and PDF content extracted from various project repositories within the CNCF (Cloud Native Computing Foundation) landscape. The data was collected by fetching MD and PDF files from different CNCF project repositories and converting them into JSON format. This dataset is intended as raw data for training large language models (LLMs).
The dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/Kubermatic/cncf-raw-data-for-llm-training.cnc-iot-trading-public
🏭 JHC Datasets & JARVIS AI Model Suite
📋 Description
A high-quality suite of datasets and AI models fine-tuned for CNC programming, IoT/Industrial automation, and Automated PCB DRC Layout Generation.
🎯 Dataset Breakdown & Accessibility
Dataset File
Scope
Status / License
Examples
dataset_cnc_iot_public.jsonl
CNC Multi-dialect, IoT, Systemd, MQTT, Docker
Public / Dual-License
146
dataset_pcb_drc_sample.jsonl
PCB Component Layout… See the full description on the dataset page: https://huggingface.co/datasets/BreyAIrev/cnc-iot-trading-public.cnc-gcode-expert
🛠️ AInewgen CNC G-code Expert
[English below]
Dataset d'entraînement instruction → G-code expert pour le pilotage de machines CNC : fraisage, tournage, perçage, filetage, compensations d'outil et sécurité machine. Couvre les dialectes Fanuc, GRBL, Marlin, LinuxCNC, Siemens et Heidenhain.
Chaque exemple contient une instruction en français (cas réaliste d'atelier) et une réponse experte : G-code complet commenté, paramètres de coupe justifiés, et vérifications de sécurité avant… See the full description on the dataset page: https://huggingface.co/datasets/BreyAIrev/cnc-gcode-expert.cnc-machiningCNC_oral_ortofon
Introduction
This is a sample from the ORAL2013 and ORTOFON datasets, maintained by the Czech National Corpus project. The datasets were created from shared .vert file format using the convert_ORTOFON_ORAL13.py script.
The versions of the datasets used here were downloaded from the LINDAT Clarin repository:
ORTOFON v1
ORAL2013
About Original Datasets
ORAL2013
The ORAL2013 corpus is spoken corpus available within the framework of the Czech National Corpus… See the full description on the dataset page: https://huggingface.co/datasets/CZLC/CNC_oral_ortofon.CNC_fictree
Introduction
This is a sample from the FicTree dataset, maintained by the Czech National Corpus project. The dataset was created from shared .vert file format using the convert_FICTREE.py script.
About Original Dataset
(Taken from project Wiki). The FicTree treebank is a syntactically annotated corpus of Czech fiction. It consists of 135,000 words (166,000 tokens). The lemmatization, morphological, and syntactic annotation were performed manually.
Composition… See the full description on the dataset page: https://huggingface.co/datasets/CZLC/CNC_fictree.CNC_KSK
Introduction
This is a sample from Corpus of Private Correspondence (KSK-dopisy) dataset, maintained by Czech National Corpus project.
The dataset was created from shared .vert file format using convert_ksk.py script.
About the Dataset
(Taken from project Wiki, translated).
Private Correspondence Corpus (KSK-Letters) allows insight into the language and style of contemporary epistolary texts of a private nature. This corpus captures what might be the final stage in… See the full description on the dataset page: https://huggingface.co/datasets/CZLC/CNC_KSK.fiber-laser-cnc-logs
PopChain Industrial Proofs — On-Chain Verified Machine Data
Real industrial machine data with cryptographic provenance on PopChain L1.
Demo Dataset (100 rows)
This is a demo preview. Full datasets available for purchase at bomwt.io/market.
Format
File
Rows
JSONL
data/train.jsonl
100
CSV
data/train.csv
100
Verification
Every record is:
SHA-256 hashed on PopChain L1
WOTS+ signed (post-quantum cryptography)
Scored by Proof-of-Process… See the full description on the dataset page: https://huggingface.co/datasets/PopChainNetwork/fiber-laser-cnc-logs.CNC_skript12
Introduction
This is the SKRIPT2012 dataset, maintained by the Czech National Corpus project. This dataset corresponds to the version available in the LINDAT repository, where it is named AKCES-1. The dataset was created from public .rtf and .doc file formats using the convert_AKCES.py script.
About Original Dataset
(Taken from project Wiki).
The Corpus SKRIPT2012 is a learner corpus aimed at representing the written language of Czech pupils and students at elementary… See the full description on the dataset page: https://huggingface.co/datasets/CZLC/CNC_skript12.CNC_Dialekt
Introduction
This is a sample from the Dialekt corpus, maintained by the Czech National Corpus project.
The dataset was created from shared .vert file format using convert_dialekt.py script.
About Original Dataset
(Taken from project Wiki).
The DIALEKT corpus presents traditional regional dialects captured across the entire Czech Republic. The dialect material was obtained by transcribing sound recordings from all dialectal regions of the Czech Republic. Additionally… See the full description on the dataset page: https://huggingface.co/datasets/CZLC/CNC_Dialekt.CNC_KHavlicek_HistNews
Introduction
This is a sample from KH-NOVINY dataset, maintained by Czech National Corpus project.
The dataset was created from shared .vert file format using convert_kh-noviny.py script.
About Original Dataset
(Taken from project Wiki). The Karel Havlíček’s Journalism Corpus (KH-Noviny) contains all journalistic text written by Karel Havlíček (1821—1856) and published in his periodicals Pražské noviny (Prague Newspaper, 1846—1848), including its supplement Česká… See the full description on the dataset page: https://huggingface.co/datasets/CZLC/CNC_KHavlicek_HistNews.CNC_Capek
Introduction
This is a sample from Capek dataset, maintained by Czech National Corpus project.
The dataset was created from shared .vert file format using convert_capek.py script.
About Original Dataset
(Taken from project Wiki). The corpus contains all texts that have been undoubtedly written by Karel Čapek himself (i.e. with no co-authors and without possible influence of a partner or translation original).
Citation
If you use this resource, please cite the… See the full description on the dataset page: https://huggingface.co/datasets/CZLC/CNC_Capek.CN_civil_code_training_dataCNC_PrezPrejavy
Introduction
This is a sample from Speeches dataset, maintained by Czech National Corpus project.
The dataset was created from shared .vert file format using convert_speeches.py script.
About Original Dataset
(Taken from project Wiki). Corpus of official presidential speeches has been created in cooperation between the CNC and University of Oslo. It covers Czech presidential speeches given between 1918 and 2015 on the occassion of anniversaries and public holidays… See the full description on the dataset page: https://huggingface.co/datasets/CZLC/CNC_PrezPrejavy.CNCF-DatasetCNC_machine_manualCNC_machine_manualCNCFStuffcnctoolscnctools2
