CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mldev19 /so100_brickThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 1, "total_frames": 297, "total_tasks": 1, "total_videos": 2, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mldev19/so100_brick.tabularrobotics10K<n<100K0 likes122 downloads1y agoHugging Face02MLDataScientist /Uzbek_news_datasetThis is an Uzbek News Dataset with 512,750 articles (120 million words and in the Latin script) scraped from the web in 2023. I combined and uploaded the dataset in this HF repo so that the community can fine-tune LLMs based on the Uzbek language. @proceedings{kuriyozov_elmurod_2023_7677431, title = {{Text classification dataset and analysis for Uzbek language}}, year = 2023, publisher = {Zenodo}, month = feb, doi =… See the full description on the dataset page: https://huggingface.co/datasets/MLDataScientist/Uzbek_news_dataset.tabulartext-generation100K<n<1M4 likes108 downloads2y agoHugging Face03mldev19 /so100_test_4This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 2, "total_frames": 596, "total_tasks": 1, "total_videos": 4, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:2" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mldev19/so100_test_4.tabularroboticsn<1K0 likes88 downloads1y agoHugging Face04mldev19 /so100_test_brick_1This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 1, "total_frames": 446, "total_tasks": 1, "total_videos": 2, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mldev19/so100_test_brick_1.tabularroboticsn<1K0 likes77 downloads1y agoHugging Face05mldev19 /so100_test_brick_5This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 1, "total_frames": 446, "total_tasks": 1, "total_videos": 2, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mldev19/so100_test_brick_5.tabularrobotics1K<n<10K0 likes32 downloads1y agoHugging Face06mldev19 /so100_test_brick_4This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 1, "total_frames": 446, "total_tasks": 1, "total_videos": 2, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mldev19/so100_test_brick_4.tabularroboticsn<1K0 likes29 downloads1y agoHugging Face07MLDataScientist /SlimOrca-Dedup-English-UzbekThis is an Uzbek translated version of https://huggingface.co/datasets/Open-Orca/SlimOrca-Dedup. It is a single parquet file. Check here for cleaned Uzbek only slim Orca dataset: https://huggingface.co/datasets/MLDataScientist/SlimOrca-Dedup-Uzbek-cleaned tabulartext-classification1M<n<10M2 likes25 downloads2y agoHugging Face08MLDataScientist /oasst2_uzbek Open Assistant Conversations Dataset Release 2 (OASST2) in Uzbek language This dataset is an Uzbek translated version of OASST2 dataset. Llama3 chat template + thread formatted dataset based on this translation is also available for model fine-tuning here. The Uzbek translation was completed in 45 hours using a single T4 GPU and nllb-200-3.3B model. Based on nllb metrics, you might want to only filter out records that were not originally in English or Russian since… See the full description on the dataset page: https://huggingface.co/datasets/MLDataScientist/oasst2_uzbek.tabularquestion-answering100K<n<1M2 likes18 downloads2y agoHugging Face09juselara1 /mlds7_datatabular10K<n<100K0 likes13 downloads3y agoHugging Face10roberto-armas /ml_data_test_detection_bank_transaction_frauds_unbalanced ML Data Test Detection Bank Transaction Frauds Unbalanced The project provides a quick and accessible dataset designed for learning and experimenting with machine learning algorithms, specifically in the context of detecting fraudulent bank transactions. It is intended for practicing and applying concepts such as Random Forest, Support Vector Machines (SVM), and Synthetic Minority Over-sampling Technique (SMOTE) to address unbalanced classification problems. Note: This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/roberto-armas/ml_data_test_detection_bank_transaction_frauds_unbalanced.tabular1K<n<10K1 likes13 downloads1y agoHugging Face11aatituanav /roberta-base-bne-mldoc-4cattabularn<1K0 likes8 downloads1y agoHugging Face12ml-system-design /ml-design-doc-reviewer-datagated ml-system-design/ml-design-doc-reviewer-data (v1.0.0) Evaluation artifacts for the ML Design Doc Reviewer project. Layout Path Description manifest/sample_manifest.csv Stratified 100-case sample manifest manifest/error_topology.csv Controlled error taxonomy for flawed docs raw/ Raw markdown exports, metadata sidecars, OCR image blocks raw/images/ Downloaded article images normalized/ Canonical 14-section ML design documents flawed/ Normalized… See the full description on the dataset page: https://huggingface.co/datasets/ml-system-design/ml-design-doc-reviewer-data.imagetext-generationn<1K0 likes5 downloads3mo agoHugging Face13juselara1 /mlds7_restaurantstabular10K<n<100K1 likes3 downloads3y agoHugging Face14randomshit11 /ml-data-130kimage100K<n<1M0 likes2 downloads2y agoHugging Face15SilviuMih21 /ML_datasettabular1K<n<10K0 likes2 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.