datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
so100_brickThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 1,
"total_frames": 297,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mldev19/so100_brick.Uzbek_news_datasetThis is an Uzbek News Dataset with 512,750 articles (120 million words and in the Latin script) scraped from the web in 2023.
I combined and uploaded the dataset in this HF repo so that the community can fine-tune LLMs based on the Uzbek language.
@proceedings{kuriyozov_elmurod_2023_7677431,
title = {{Text classification dataset and analysis for Uzbek
language}},
year = 2023,
publisher = {Zenodo},
month = feb,
doi =… See the full description on the dataset page: https://huggingface.co/datasets/MLDataScientist/Uzbek_news_dataset.so100_test_4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 2,
"total_frames": 596,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mldev19/so100_test_4.so100_test_brick_1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 1,
"total_frames": 446,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mldev19/so100_test_brick_1.so100_test_brick_5This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 1,
"total_frames": 446,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mldev19/so100_test_brick_5.so100_test_brick_4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 1,
"total_frames": 446,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mldev19/so100_test_brick_4.SlimOrca-Dedup-English-UzbekThis is an Uzbek translated version of https://huggingface.co/datasets/Open-Orca/SlimOrca-Dedup.
It is a single parquet file.
Check here for cleaned Uzbek only slim Orca dataset: https://huggingface.co/datasets/MLDataScientist/SlimOrca-Dedup-Uzbek-cleaned
oasst2_uzbek
Open Assistant Conversations Dataset Release 2 (OASST2) in Uzbek language
This dataset is an Uzbek translated version of OASST2 dataset.
Llama3 chat template + thread formatted dataset based on this translation is also available for model fine-tuning here.
The Uzbek translation was completed in 45 hours using a single T4 GPU and nllb-200-3.3B model.
Based on nllb metrics, you might want to only filter out records that were not originally in English or Russian since… See the full description on the dataset page: https://huggingface.co/datasets/MLDataScientist/oasst2_uzbek.mlds7_dataml_data_test_detection_bank_transaction_frauds_unbalanced
ML Data Test Detection Bank Transaction Frauds Unbalanced
The project provides a quick and accessible dataset designed for learning and experimenting with machine learning algorithms, specifically in the context of detecting fraudulent bank transactions. It is intended for practicing and applying concepts such as Random Forest, Support Vector Machines (SVM), and Synthetic Minority Over-sampling Technique (SMOTE) to address unbalanced classification problems.
Note: This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/roberto-armas/ml_data_test_detection_bank_transaction_frauds_unbalanced.roberta-base-bne-mldoc-4catml-design-doc-reviewer-data
ml-system-design/ml-design-doc-reviewer-data (v1.0.0)
Evaluation artifacts for the ML Design Doc Reviewer project.
Layout
Path
Description
manifest/sample_manifest.csv
Stratified 100-case sample manifest
manifest/error_topology.csv
Controlled error taxonomy for flawed docs
raw/
Raw markdown exports, metadata sidecars, OCR image blocks
raw/images/
Downloaded article images
normalized/
Canonical 14-section ML design documents
flawed/
Normalized… See the full description on the dataset page: https://huggingface.co/datasets/ml-system-design/ml-design-doc-reviewer-data.mlds7_restaurantsml-data-130kML_dataset
