CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01astro-legacy-archive /msam-released-products MSAM released flight products This dataset contains the complete numeric contents of the three MSAM1 flight archives served by LAMBDA: the June 1992, June 1994, and June 1995 observation tables, sampled beam maps, and released 1994/1995 covariance matrices. The 38 configuration identifiers preserve the flight directory and source filename stem. How to use python -m venv .venv && .venv/bin/pip install datasets huggingface_hub pyarrow from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/msam-released-products.tabular10K<n<100K0 likes701 downloads18d agoHugging Face02APProjects /us-layoffs-by-metro-area-msa-warn-act US layoffs by metro area: 54,170 WARN notices mapped to 765 metro and micro areas Rebuilt 2026-09-22. 765 of the 935 US core-based statistical areas carry at least one layoff notice on record — 361 metropolitan and 404 micropolitan. Nobody hires, sells or reports by county. A recruiter covers Austin; an account team books the Phoenix metro; a reporter writes Bay Area layoffs. State agencies publish neither — they publish the site of a layoff as free text in 48 different… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-layoffs-by-metro-area-msa-warn-act.tabulartabular-regression10K<n<100K0 likes607 downloads4h agoHugging Face03m-sakka /agripotentialMore information and competition link: https://github.com/MohammadElSakka/agripotential https://www.codabench.org/competitions/12055/ https://zenodo.org/records/15551829 imageimage-segmentation1K<n<10K2 likes244 downloads2mo agoHugging Face04songlab /gpn-msa-sapiens-dataset Training windows for GPN-MSA-Sapiens For more information check out our paper and repository. Path in Snakemake: results/dataset/multiz100way/89/128/64/True/defined.phastCons.percentile-75_0.05_0.001 tabular1M<n<10M0 likes242 downloads2y agoHugging Face05HeshamHaroon /arabic-msa-25k-saudi-male-tashkeel Arabic MSA 25K — Saudi Male (Tashkeel) 25,000 fully-diacritized Arabic MSA text + audio pairs, rendered with a single Saudi male neural voice at 48 kHz / 16-bit PCM, across 10 thematic categories. Dataset Summary arabic-msa-25k-saudi-male-tashkeel is a 25,000-clip Modern Standard Arabic (MSA) speech corpus with matching diacritized text (full tashkeel / ḥarakāt). Every clip is synthesized by the single voice ar-SA-HamedNeural (Azure Neural TTS, Saudi Arabic male) at 48… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/arabic-msa-25k-saudi-male-tashkeel.tabulartext-to-speech10K<n<100K9 likes185 downloads5mo agoHugging Face06clarayyu22 /gpn-msa-microglia-fulltabular1M<n<10M0 likes159 downloads2y agoHugging Face07tunis-ai /tunisian-msa-parallel-corpus Dataset Description This is an ambitious project to create a high-quality, reproducible parallel corpus for Modern Standard Arabic (MSA) and Tunisian Arabic (aeb) through a sophisticated synthetic data generation pipeline. The dataset is being developed by the Tunisia.AI community to address the scarcity of high-quality dialectal data for training and evaluating language models. The primary goal is to provide a rich, well-documented resource for the research and development of:… See the full description on the dataset page: https://huggingface.co/datasets/tunis-ai/tunisian-msa-parallel-corpus.tabulartranslation1K<n<10K0 likes102 downloads1y agoHugging Face08msakota /edisum_dataset Dataset Card for Edisum Dataset Description For more details: Github repository: https://github.com/epfl-dlab/edisum Paper: https://arxiv.org/pdf/2404.03428.pdf Languages Edisum only contains Wikipedia data collected from English Wikipedia. Consequently, synthetic data is also only generated in English. Dataset Structure The Edisum meta-dataset actually comprises 5 datasets: wikiepdia_processed_data (Filtered existing Wikipedia data)… See the full description on the dataset page: https://huggingface.co/datasets/msakota/edisum_dataset.tabular100K<n<1M0 likes96 downloads2y agoHugging Face09sokrypton /afdb-msa-index AlphaFold DB minimizer index A static sequence-search index over 239,602,633 AlphaFold DB v6 entries (99.4% of the database), designed to be queried directly from a browser with HTTP Range requests. No server, no search engine, no MMseqs2. It exists to answer one question quickly: which AFDB entry is ≥90% identical to my sequence? — so that entry's precomputed MSA can be borrowed and re-indexed onto the query instead of computing a new alignment. Used by AFDB MSA.… See the full description on the dataset page: https://huggingface.co/datasets/sokrypton/afdb-msa-index.tabularn<1K0 likes89 downloads2mo agoHugging Face10msavchen-nasa /clr_motion_planning_hwtabular10K<n<100K0 likes52 downloads2mo agoHugging Face11phamluan /vn_msa_sar_v3 VN-MSASar v3 — Multilingual Aspect-Based Sentiment + Sarcasm (hotel & restaurant reviews) Chia split ở mức review. 100% review thật — toàn bộ dữ liệu sinh bằng LLM của bản đầu (nguồn augmented, 7.659 dòng rải ở cả ba split) đã bị loại bỏ. Không tương thích ngược với phamluan/vn_sar_msa. 71% câu trong test nằm trong train của bản cũ. Mọi checkpoint huấn luyện trên bản cũ đều đã nhìn thấy test này — phải huấn luyện lại từ đầu. Splits Split Dòng Review Mỉa… See the full description on the dataset page: https://huggingface.co/datasets/phamluan/vn_msa_sar_v3.tabulartext-classification10K<n<100K0 likes39 downloads1mo agoHugging Face12tunis-ai /tunisian-msa-parallel-corpus-evaluated Dataset Description This dataset is a synthetic parallel corpus of Tunisian Arabic (aeb) and Modern Standard Arabic (arb). It was created with a rigorous multi-stage pipeline to maximize quality and reproducibility, addressing the scarcity of high-quality resources for Tunisian Arabic NLP. The primary goals are to support: Machine translation between Tunisian Arabic and MSA. Research in dialectal-aware text generation and evaluation. Cross-dialect representation learning in… See the full description on the dataset page: https://huggingface.co/datasets/tunis-ai/tunisian-msa-parallel-corpus-evaluated.tabulartranslation1K<n<10K2 likes32 downloads1y agoHugging Face13MSakai /record-bi-11This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "bi_so100_follower", "total_episodes": 2, "total_frames": 569, "total_tasks": 1, "total_videos": 6, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:2" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/MSakai/record-bi-11.tabularroboticsn<1K0 likes24 downloads1y agoHugging Face14MSakai /eval_2-fold-pants-20This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "bi_so100_follower", "total_episodes": 1, "total_frames": 1758, "total_tasks": 1, "total_videos": 1, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/MSakai/eval_2-fold-pants-20.tabularrobotics1K<n<10K0 likes23 downloads1y agoHugging Face15phamluan /vn_msa_sar Dataset Card: Cleaned Multilingual Aspect-Based Sentiment & Sarcasm Reviews Tóm tắt Dataset (Dataset Summary) Dataset này chứa 83.009 đánh giá (reviews) đa ngôn ngữ đã qua quá trình làm sạch (cleaned). Dữ liệu được thu thập từ các nền tảng như Google Maps, Booking, Foody và dữ liệu tăng cường (augmented), tập trung vào các địa điểm kinh doanh dịch vụ (Khách sạn và Nhà hàng). Mỗi đánh giá được gán nhãn chi tiết phục vụ cho các bài toán NLP như: Phân tích cảm xúc… See the full description on the dataset page: https://huggingface.co/datasets/phamluan/vn_msa_sar.tabulartext-classification10K<n<100K0 likes22 downloads5mo agoHugging Face16MSakai /record-bi-100This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "bi_so100_follower", "total_episodes": 20, "total_frames": 35446, "total_tasks": 1, "total_videos": 20, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:20" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/MSakai/record-bi-100.tabularrobotics10K<n<100K0 likes21 downloads1y agoHugging Face17msaramhassan /analogy_concept_questions Analogy Concept Questions This dataset was generated using YourBench (v0.6.0), an open-source framework for generating domain-specific benchmarks from document collections. Pipeline Steps ingestion: Read raw source documents, convert them to normalized markdown and save for downstream steps summarization: Perform hierarchical summarization: chunk-level LLM summaries followed by combine-stage reduction chunking: Split texts into token-based single-hop and multi-hop… See the full description on the dataset page: https://huggingface.co/datasets/msaramhassan/analogy_concept_questions.tabularn<1K0 likes20 downloads11mo agoHugging Face18MSakai /record-bi-7This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "bi_so100_follower", "total_episodes": 20, "total_frames": 35867, "total_tasks": 1, "total_videos": 20, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:20" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/MSakai/record-bi-7.tabularrobotics10K<n<100K0 likes19 downloads1y agoHugging Face19MSaalaamaa /marine_vla_dataset_v2 Marine VLA Dataset — v2 (Fixed Expert) This dataset contains 12,070 frames collected in the Mangalia harbor Gazebo simulation, used to fine-tune a vision-language-action (VLA) model for marine navigation. Built with opencode + DeepSeek V4 Flash — every component of this project (simulation setup, expert controller, data pipeline, ML training, inference nodes) was developed using opencode CLI with DeepSeek V4 Flash as the underlying model. v2 Changes (Fixed Expert)… See the full description on the dataset page: https://huggingface.co/datasets/MSaalaamaa/marine_vla_dataset_v2.tabular10K<n<100K0 likes19 downloads4mo agoHugging Face20MSakai /record-bi-3This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "bi_so100_follower", "total_episodes": 2, "total_frames": 393, "total_tasks": 1, "total_videos": 2, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:2" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/MSakai/record-bi-3.tabularroboticsn<1K0 likes18 downloads1y agoHugging Face21polestarllp /Siamese_Finetune_MSA_SOW_ContractsA dataset prepared for siamese finetuning, to distingush between texts from Legal Contracts text (Majorly SOW, MSA others Legal Algreements, Offer Letters etc) and text scraped from books, news articles, reviews etc license: apache-2.0 tabular10K<n<100K1 likes17 downloads3y agoHugging Face22SadatHossain01 /avian-msa-iqtreetabular10K<n<100K0 likes15 downloads1y agoHugging Face23msaavedra1234 /parise-tts-6h-taggedtabular1K<n<10K0 likes14 downloads2y agoHugging Face24msaavedra1234 /jenny-tts-tags-6htabular1K<n<10K0 likes13 downloads2y agoHugging Face25hbenayed /tunisian-msa-parallel-corpus-evaluated Dataset Description This dataset is a synthetic parallel corpus of Tunisian Arabic (aeb) and Modern Standard Arabic (arb). It was created with a rigorous multi-stage pipeline to maximize quality and reproducibility, addressing the scarcity of high-quality resources for Tunisian Arabic NLP. The primary goals are to support: Machine translation between Tunisian Arabic and MSA. Research in dialectal-aware text generation and evaluation. Cross-dialect representation learning in Arabic… See the full description on the dataset page: https://huggingface.co/datasets/hbenayed/tunisian-msa-parallel-corpus-evaluated.tabulartranslation1K<n<10K0 likes13 downloads6mo agoHugging Face26msaavedra1234 /jenny-tts-6h-taggedtabular1K<n<10K0 likes11 downloads2y agoHugging Face27MSakai /fold-pants-20tabular10K<n<100K0 likes11 downloads1y agoHugging Face28phamluan /vi_msa_sar VN-SarMSA-vi Vietnamese aspect-level sentiment (7-point, −3…+3) and sarcasm (binary) annotations for hotel/restaurant reviews. 17,236 (sentence, aspect) pairs; splits 13,758 / 1,739 / 1,739 (train/validation/test) Splits are leakage-free by construction: reviews sharing any normalized sentence are grouped before splitting, so no sentence or review crosses splits source column marks synthetic augmentation ('augmented'); evaluate sentiment on source != 'augmented' for a… See the full description on the dataset page: https://huggingface.co/datasets/phamluan/vi_msa_sar.tabulartext-classification10K<n<100K0 likes10 downloads3mo agoHugging Face29clarayyu22 /gpn-msa-with-microgliatabular1K<n<10K0 likes9 downloads2y agoHugging Face30tktung /MSAtabular1K<n<10K0 likes8 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.