datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
open_video_datarechtspraak-opendata
Rechtspraak OpenData Archive
This dataset repository stores archival snapshots used by the Maastricht Rechtspraak data pipelines. The repository is organized as source-oriented data, not as a normalized analysis dataset.
Repository Layout
raw_data/ contains raw Rechtspraak OpenData source snapshots as downloaded from the upstream OpenData feeds.
exports/ contains compressed export snapshots produced by the rs-migration pipeline.
Current Contents… See the full description on the dataset page: https://huggingface.co/datasets/davidwickerhf/rechtspraak-opendata.open-lm-instruction-dataOpen-Qwen2VL-Data-InterleavedIfGPT-OPEN-Dataset
IfGPT Dataset
Objectives of the project IfGPT
The IfGPT Dataset is developed within the project IfGPT: Infrastructure for Fine-tuning Pre-trained Large Language Models which aims to establish a freely accessible infrastructure for the selection and pre-processing of large datasets for Bulgarian as well as tailored data for particular industries and fine-tuning suitable freely available large language models for specific purposes.
IfGPT Dataset
IfGPT… See the full description on the dataset page: https://huggingface.co/datasets/DCL-IBL/IfGPT-OPEN-Dataset.Open-Claw_Dataset
