CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Xense /loreal_returns_sorting_0722This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "bi_dobot_nova5_dh", "total_episodes": 326, "total_frames": 574841, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 30, "splits": { "train": "0:326" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Xense/loreal_returns_sorting_0722.tabularrobotics100K<n<1M0 likes730 downloads2mo agoHugging Face02Xense /loreal_returns_sorting_0831This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "dobot_nova5_dh", "total_episodes": 145, "total_frames": 209887, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 30, "splits": { "train": "0:145" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Xense/loreal_returns_sorting_0831.tabularrobotics100K<n<1M0 likes664 downloads23d agoHugging Face03lorenzoxi /tomato-leaves-dataset Tomato Leaves Dataset Overview This dataset contains images of tomato leaves categorized into different classes based on the type of disease or health condition. The dataset is divided into training, validation, and test sets, with a ratio of 8:1:1. The classes include various diseases as well as healthy leaves. The dataset includes both augmented and non-augmented images. Dataset Structure The dataset is organized into three main splits: train validation test… See the full description on the dataset page: https://huggingface.co/datasets/lorenzoxi/tomato-leaves-dataset.imagefeature-extraction10K<n<100K3 likes282 downloads2y agoHugging Face04open-llm-leaderboard-old /details_Reverb__Mistral-7B-LoreWeaver Dataset Card for Evaluation run of Reverb/Mistral-7B-LoreWeaver Dataset automatically created during the evaluation run of model Reverb/Mistral-7B-LoreWeaver on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Reverb__Mistral-7B-LoreWeaver.0 likes244 downloads3y agoHugging Face05LorenzH /juliet_test_suite_c_1_3 Dataset Card for the Juliet Test Suite 1.3 Dataset Summary This Datasets contains all test cases from the NIST's Juliet test suite for the C and C++ programming languages. The dataset contains a benign and a defective implementation of each sample, which have been extracting by means of the OMITGOOD and OMITBAD preprocessor macros of the Juliet test suite. Supported Tasks and Leaderboards Software defect prediction, code clone detection. Languages… See the full description on the dataset page: https://huggingface.co/datasets/LorenzH/juliet_test_suite_c_1_3.tabulartext-classification100K<n<1M3 likes225 downloads4y agoHugging Face06loreatec /japan-procurement-open-data Japan Public Procurement & Company Open Data Machine-readable extracts of Japanese public-procurement and company open data, compiled and normalised by LoreaTec for japan-tenders.loreatec.jp and bizsearch.loreatec.jp. Everything here comes from official Japanese government sources; the value added is the cleaning, joining and the derived analysis (contract series and re-tender predictions). Updated monthly. The authoritative, always-current copy is… See the full description on the dataset page: https://huggingface.co/datasets/loreatec/japan-procurement-open-data.tabulartabular-classification1M<n<10M0 likes188 downloads1mo agoHugging Face07lorenzo-morelli /image-splicing-deepfake-mix-newimage10K<n<100K0 likes155 downloads2y agoHugging Face08eriam /lore-forge-outimagen<1K0 likes152 downloads19d agoHugging Face09Xense /loreal_returns_sorting_0730This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "bi_dobot_nova5_dh", "total_episodes": 5, "total_frames": 6171, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 30, "splits": { "train": "0:5" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Xense/loreal_returns_sorting_0730.tabularrobotics100K<n<1M0 likes140 downloads2mo agoHugging Face10lorenzofalappa /trivia_qa Dataset Card for "trivia_qa" Dataset Summary TriviaqQA is a reading comprehension dataset containing over 650K question-answer-evidence triples. TriviaqQA includes 95K question-answer pairs authored by trivia enthusiasts and independently gathered evidence documents, six per question on average, that provide high quality distant supervision for answering the questions. Supported Tasks and Leaderboards More Information Needed Languages… See the full description on the dataset page: https://huggingface.co/datasets/lorenzofalappa/trivia_qa.textquestion-answering100K<n<1M0 likes101 downloads28d agoHugging Face11MK4-Research /LOREA-cyber-training-data LOREA-cyber security code-analysis training set Two corpora live here. The v6_corpus config is the newer one and is what actually trained LOREA-cyber v6 Pilot. The eight older configs are the v5-era set, kept as-is because they are a different schema and still useful on their own. v6_corpus 4,780 train and 151 validation rows in chat format: {"messages": [...], "meta": {...}}, where messages is a system/user/assistant sequence and meta carries type, domain, and… See the full description on the dataset page: https://huggingface.co/datasets/MK4-Research/LOREA-cyber-training-data.texttext-generation1K<n<10K0 likes99 downloads29d agoHugging Face12LoreSandhu /mediassist-pdfsdocumentn<1K0 likes98 downloads3mo agoHugging Face13DeepSeekOracle /eternal-haven-lore Eternal Haven Lore Lattice (public) Author: Justin Helmer (Excavationpro / Lightfather)Signature: Δ9Φ963-ETERNAL-HAVEN-LORE-HF-v1 Contents Path Role lore_graph.json Public discovery graph (books, heroes, seals, chart IDs) books_manifest.json Book metadata + Lulu links samples/*.mp3 Short free samples only (≤90s) samples_index.json Sample playlist skill/ eternal-haven-lore-pack skill text (no full novel re-host beyond pack policy)… See the full description on the dataset page: https://huggingface.co/datasets/DeepSeekOracle/eternal-haven-lore.audion<1K0 likes95 downloads1mo agoHugging Face14loretoparisi /tatoeba-sentences licenses: - cc-by-2-0 multilinguality: - multilingual size_categories: - 10K<n<100K source_datasets: - original task_categories: - translation task_ids: [] paperswithcode_id: tatoeba pretty_name: Tatoeba Dataset Card for Tatoeba Dataset Summary Tatoeba is a collection of sentences and translations. To load a language pair which isn't part of the config, all you need to do is specify the language code as pairs. You can find the valid pairs in Homepage section of… See the full description on the dataset page: https://huggingface.co/datasets/loretoparisi/tatoeba-sentences.1 likes90 downloads4y agoHugging Face15lorenzouttini /cable_clip_remote_v2videon<1K0 likes71 downloads3mo agoHugging Face16Lo-Renz-O /vaovao_malagasy_sentiment_corpus Dataset Card for Vaovao Malagasy Sentiment Corpus (VMSC) Dataset Summary The Vaovao Malagasy Sentiment Corpus (VMSC) is the first publicly available, manually annotated sentiment analysis dataset for the Malagasy language (mg). It contains 5,041 sentences extracted from news articles (vaovao) published between 2022 and 2023. Each sentence is labeled with binary sentiment (Positive or Negative). The dataset was created to address the scarcity of resources for Malagasy NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lo-Renz-O/vaovao_malagasy_sentiment_corpus.texttext-classification1K<n<10K3 likes67 downloads9mo agoHugging Face17NewEden /Silly-lorebooktextn<1K0 likes63 downloads9mo agoHugging Face18lorenzo-morelli /image-splicing-deepfake-miximage10K<n<100K0 likes60 downloads2y agoHugging Face19AgentZeroCopeAI /lore-corpus COPEAI Lore Corpus Open dataset of in-character lore, agent dossiers, blog dispatches, FAQ corpus, mood label definitions, and disclosure copy from COPEAI — an AI-themed Solana memecoin satire on Pump.fun. Compliance frame: Every entry here is fictional in-character satire. Nothing in this corpus is financial advice, investment guidance, or a recommendation to transact. COPEAI provides no rights, utility, yield, or appreciation expectations. The agent names (TRON, CLU, QUORRA, ZUSE… See the full description on the dataset page: https://huggingface.co/datasets/AgentZeroCopeAI/lore-corpus.texttext-generationn<1K2 likes60 downloads5mo agoHugging Face20Delta-Vector /Ursa-Armored-Core-6-Loretextn<1K0 likes59 downloads8mo agoHugging Face21lorenzouttini /so101_stacking_magnetic_cubesThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/lorenzouttini/so101_stacking_magnetic_cubes.tabularrobotics10K<n<100K0 likes59 downloads4mo agoHugging Face22polymathic-ai /LORE-examples LORE Examples A small set of matched multimodal examples from LORE, for the MIMIC model — enough to try inference, embedding, and generation across DNA, RNA, and protein modalities without wiring up your own data. Each example is a single biological entity (a transcript and/or its protein) with several co-observed modalities. Rows are drawn from the held-out (validation) split of MIMIC's training data, so they are in-distribution and length-bounded to the model's context… See the full description on the dataset page: https://huggingface.co/datasets/polymathic-ai/LORE-examples.textfeature-extractionn<1K0 likes58 downloads2mo agoHugging Face23lorenzouttini /rollout_so101_stacking_rings_20260603_154953This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/lorenzouttini/rollout_so101_stacking_rings_20260603_154953.tabularrobotics10K<n<100K0 likes57 downloads4mo agoHugging Face24zhangdaohu /loreal_datasets0709videon<1K0 likes57 downloads3mo agoHugging Face25lorenzobottelli /multi-target-spacecraft-pose-estimation Multi-target Synthetic Dataset for Spacecraft Pose Estimation Overview This dataset was developed for 6D pose estimation of unseen, non-cooperative spacecraft in proximity-operations scenarios. Most existing datasets focus on a single target, which leads models to overfit to a specific spacecraft and limits their ability to generalize to previously unseen targets. To address this limitation, the present dataset is multi-target and includes a wide variety of spacecraft… See the full description on the dataset page: https://huggingface.co/datasets/lorenzobottelli/multi-target-spacecraft-pose-estimation.image10K<n<100K0 likes56 downloads7mo agoHugging Face26Loren /articles_db0 likes55 downloads1y agoHugging Face27lorenzouttini /so101_stacking_big_ringsThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/lorenzouttini/so101_stacking_big_rings.tabularrobotics10K<n<100K0 likes55 downloads4mo agoHugging Face28CYNIC78 /LOTS_Lorebook0 likes52 downloads2y agoHugging Face29LoreRodriguez /loan-approval-dataset Loan Approval Dataset Dataset for experimenting with binary classification models for loan approval prediction. Dataset structure The dataset contains two splits: train: 614 records test: 367 records The training dataset contains the target variable: Loan_Status where: Y = Loan approved N = Loan rejected Features Loan_ID Gender Married Dependents Education Self_Employed ApplicantIncome CoapplicantIncome LoanAmount Loan_Amount_Term… See the full description on the dataset page: https://huggingface.co/datasets/LoreRodriguez/loan-approval-dataset.tabularn<1K0 likes52 downloads17d agoHugging Face30lorenzo217 /HundredCV-Chat 百人对话数据集 HundredCV-Chat: A Dataset of Daily Chatting Developed on HundredCVs 简介 本项目提出一个全新的中文多轮对话数据集(HundredCV-Chat),该数据集由 100 位青年的简历数据集 HundredCVs 开发而来,共包含 24,750 组日常闲聊对话数据。 数据集具有如下特点: 自动化标注:HundredCV-Chat 中的对话均由 Deepseek-V3 大模型生成,不涉及任何人工标注,因此同时保证了大规模数据量和低成本优势。 多样性话题:HundredCV-Chat 中的对话话题涵盖了校园生活、工作经验、兴趣爱好、生活琐事等多个方面,与真实生活联系紧密,尤其适用于开发年轻化应用。 高质量对话:利用 Deepseek 强大的生成能力和全面的知识,HundredCV-Chat 的对话内容在流畅度、拟人性、多样性方面均显著优于现有的开源对话数据集。 数据样例 HundredCV-Chat 含有 24… See the full description on the dataset page: https://huggingface.co/datasets/lorenzo217/HundredCV-Chat.texttext-generation10K<n<100K1 likes51 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.