CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01UCSC-VLAA /Recap-DataComp-1B Dataset Card for Recap-DataComp-1B Recap-DataComp-1B is a large-scale image-text dataset that has been recaptioned using an advanced LLaVA-1.5-LLaMA3-8B model to enhance the alignment and detail of textual descriptions. Dataset Details Dataset Description Our paper aims to bridge this community effort, leveraging the powerful and open-sourced LLaMA-3, a GPT-4 level LLM. Our recaptioning pipeline is simple: first, we fine-tune a LLaMA-3-8B powered… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/Recap-DataComp-1B.imagezero-shot-classification1B<n<10B205 likes9.5k downloads2y agoHugging Face02einrafh /hnm-fashion-recommendations-data Dataset Rekomendasi Fashion H&M Dataset ini berisi data transaksi, atribut pelanggan, dan metadata produk yang telah dianonimkan dari H&M Group. Kumpulan data komprehensif ini memungkinkan pemodelan perilaku pembelian pelanggan secara mendalam. Wawasan yang dihasilkan dapat dimanfaatkan untuk berbagai tujuan bisnis yang strategis, mulai dari meningkatkan personalisasi pengalaman berbelanja, mengoptimalkan manajemen inventaris untuk efisiensi produksi, hingga mendukung inisiatif… See the full description on the dataset page: https://huggingface.co/datasets/einrafh/hnm-fashion-recommendations-data.imagetabular-classification10M<n<100M3 likes3.3k downloads1y agoHugging Face03hbseong /record-pick-and-place-pos5-so101This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 240, "total_frames": 119443, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 30, "splits": { "train": "0:240" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hbseong/record-pick-and-place-pos5-so101.tabularrobotics100K<n<1M0 likes3.1k downloads10mo agoHugging Face04dacorvo /funes-handoff-recall-benchmark handover-vs-recall A long investigation bloats an agent session until each new turn costs more to carry the context than to do the work. Switching to a fresh session avoids that — but the findings have to travel somehow, and the ways of moving them differ in cost. This benchmark measures those ways, as cost per successful task, on tasks that genuinely require the prior investigation: arm channel A branch-only switch, carry nothing — the fresh session re-derives the… See the full description on the dataset page: https://huggingface.co/datasets/dacorvo/funes-handoff-recall-benchmark.tabularn<1K0 likes3.1k downloads23d agoHugging Face05hbseong /record-pick-and-place-so101This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 188, "total_frames": 88161, "total_tasks": 3, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 30, "splits": { "train": "0:188" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hbseong/record-pick-and-place-so101.tabularrobotics10K<n<100K0 likes2.8k downloads10mo agoHugging Face06obadx /mualem-recitations-original المصاحف القرآنية مصاحف مجمعمة من القراء المتقنين لبناء نماذج ذكاء اصطناعي لخدمة القرآن الكريم. أنظر هنا لأكواد بناء قاعدة التلاوات القرآنية البيانات الوصفية للمصاحف ds = load_dataset('obadx/mualem-recitations-original', name='moshaf_metadata')['train'] وصف أوجه حفص Attribute Name Arabic Name Values Default Value More Info rewaya الرواية - hafs (حفص) The type of the quran Rewaya. recitation_speed سرعة التلاوة - mujawad (مجود)-… See the full description on the dataset page: https://huggingface.co/datasets/obadx/mualem-recitations-original.audion<1K0 likes2.6k downloads1y agoHugging Face07cornuHGF /recap-datacomp-12m-wdsimage10M<n<100M0 likes2.3k downloads1y agoHugging Face08techiaith /evals-speech-recognition-cy-en Welsh ASR Model Evaluation Transcription Dataset This resource compiles the output transcriptions from multiple Welsh Automatic Speech Recognition (ASR) models across several test sets. The data is structured hierarchically: Splits delineate the individual test sets. Configs within each split detail the performance (transcriptions) of a specific ASR model and its version on that set. Metrics Results model test task wer cer… See the full description on the dataset page: https://huggingface.co/datasets/techiaith/evals-speech-recognition-cy-en.tabular100K<n<1M0 likes2.1k downloads6d agoHugging Face09obadx /mualem-recitations-annotatedaudio100K<n<1M4 likes2.1k downloads1y agoHugging Face10Zhongzhi1228 /Recursive-Task-Synthesis Recursive Task Synthesis This dataset contains 37,484 validated command-line task instances produced through recursive task synthesis. Public identifiers are opaque and stable. metadata/tasks.parquet: one searchable row per task instance. metadata/shard_manifest.jsonl: TAR sizes and SHA256 checksums. data/tasks-*.tar: sanitized runnable task packages. The searchable task rows include: instruction: contents of instruction.md. task_toml: contents of task.toml. solution:… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhi1228/Recursive-Task-Synthesis.tabularreinforcement-learning10K<n<100K19 likes1.8k downloads2mo agoHugging Face11hbseong /record-pick-and-place-ez2-so101This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 171, "total_frames": 108092, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 30, "splits": { "train": "0:171" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hbseong/record-pick-and-place-ez2-so101.tabularrobotics100K<n<1M0 likes1.7k downloads10mo agoHugging Face12sxiong /ReClor ReClor: A Reading Comprehension Dataset Requiring Logical Reasoning This repository provides the dataset from the paper ReClor: A Reading Comprehension Dataset Requiring Logical Reasoning. We corrected the original format issues to ensure full compatibility with the Hugging Face Datasets library. For more details, please visit the original project page. tabularquestion-answering1K<n<10K1 likes1.5k downloads11mo agoHugging Face13reczoo /Frappe_x1 Frappe_x1 Dataset description: The Frappe dataset contains a context-aware app usage log, which comprises 96203 entries by 957 users for 4082 apps used in various contexts. It has 10 feature fields including user_id, item_id, daytime, weekday, isweekend, homework, cost, weather, country, city. The target value indicates whether the user has used the app under the context. Following the AFN work, we randomly split the data into 7:2:1 as the training set, validation set, and test set… See the full description on the dataset page: https://huggingface.co/datasets/reczoo/Frappe_x1.tabular100K<n<1M1 likes1.4k downloads3y agoHugging Face14hbseong /record-pick-and-place-ez-so101This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 110, "total_frames": 84580, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 30, "splits": { "train": "0:110" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hbseong/record-pick-and-place-ez-so101.tabularrobotics10K<n<100K0 likes1.3k downloads10mo agoHugging Face15originlab /game-recordings-v3gated OriginLab Game Recordings v0.3.0 Human gameplay captured in-engine under per-title licenses at 1080p / 60 FPS CFR on one shared frame clock: every stream starts at frame 0 and frame k matches frame k across pre-HUD and post-HUD RGB, surface normals, metric depth, audio, camera telemetry, keyboard and mouse inputs, in-engine action events and game state, and world telemetry — plus per-frame training tables. Watch full playable previews of every modality, side by side and in… See the full description on the dataset page: https://huggingface.co/datasets/originlab/game-recordings-v3.tabulardepth-estimationn<1K11 likes1.2k downloads4d agoHugging Face16open-index /ccrawl-recrawl-domains Common Crawl Domain Recrawl Live fetches of the home page of every ranked domain in Common Crawl's web graph, rendered to Markdown as they are fetched What is it? Common Crawl's web graph ranks domains by how central they are, but it does not tell you what those domains actually serve today. This dataset walks that ranking from the top and fetches each domain's home page now, storing the response as one Parquet row with the body, the headers, the timing and the… See the full description on the dataset page: https://huggingface.co/datasets/open-index/ccrawl-recrawl-domains.tabulartext-generation1M<n<10M0 likes1.2k downloads1mo agoHugging Face17recoilme /mjnj_flux32tabular100K<n<1M0 likes1.1k downloads7mo agoHugging Face18Zhongzhi1228 /Recursive-Task-Synthesis-Trajectories Recursive Task Synthesis Trajectories This dataset contains 327,189 completed agent trajectories collected on recursively synthesized command-line tasks. Public identifiers are opaque and stable. The trajectory JSON retains messages, actions, observations, and token counts. Token-level log-probability arrays and duplicated debug/session captures are excluded from the public packages. metadata/trajectories.parquet: searchable trajectory metadata. metadata/shard_manifest.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhi1228/Recursive-Task-Synthesis-Trajectories.tabularreinforcement-learning100K<n<1M3 likes1k downloads1mo agoHugging Face19RoboCOIN /alpha_bot_2_recover_after_touching_an_obstaclegated alpha_bot_2_recover_after_touching_an_obstacle 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: alpha_bot_2 | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: home 🤖 Atomic Actions This dataset includes the following atomic actions: pick grasp 📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/alpha_bot_2_recover_after_touching_an_obstacle.tabularrobotics100K<n<1M0 likes825 downloads9mo agoHugging Face20hbseong /record-stacking-so101This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 119, "total_frames": 119446, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 30, "splits": { "train": "0:119" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hbseong/record-stacking-so101.tabularrobotics100K<n<1M0 likes720 downloads10mo agoHugging Face21somosnlp /RecetasDeLaAbuela Motivación inicial Este corpus ha sido creado durante el Hackathon SomosNLP Marzo 2024: #Somos600M (https://somosnlp.org/hackathon). Responde a una de las propuestas somosnlp sobre 'Recetas típicas por país/zona geográfica'. Nombre del Proyecto Este corpus o dataset se llama 'RecetasDeLaAbuel@' y es un homenaje a todas nuestr@s abuel@s que nos han enseñado a cocinar. Se trata de la mayor y más completa colección de recetas open-source en español de países… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp/RecetasDeLaAbuela.tabularquestion-answering10K<n<100K7 likes695 downloads2y agoHugging Face22VibeCuisine /recreate-bug-pre-fix-v1-trimThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 20, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos", "tilt.pos" ]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/recreate-bug-pre-fix-v1-trim.tabularrobotics10K<n<100K0 likes679 downloads2mo agoHugging Face23VibeCuisine /recreate-bug-post-fix-v1-trimThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 20, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos", "tilt.pos" ]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/recreate-bug-post-fix-v1-trim.tabularrobotics10K<n<100K0 likes679 downloads2mo agoHugging Face24jason1966 /aksahaha_crop-recommendation crop recommendation Crop Growth Recommendations: Optimal Conditions for Higher Yields Dataset Info Source: Kaggle Original Size: 0.06 MB Kaggle Downloads: 4,065 Files: 1 Files Crop_recommendation.csv Mirrored from Kaggle tabular1K<n<10K0 likes632 downloads6mo agoHugging Face25LeRobot-worldwide-hackathon /190-RecycloBots-recyclobotsThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 50, "total_frames": 84070, "total_tasks": 1, "total_videos": 100, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LeRobot-worldwide-hackathon/190-RecycloBots-recyclobots.tabularrobotics10K<n<100K0 likes613 downloads1y agoHugging Face26RyanSaklad /ReCITE ReCITE: Real-world CausalIty from Textual Evidence Benchmark ReCITE (Real-world CausalIty from Textual Evidence) is a benchmark for evaluating LLMs on causal graph extraction from real-world scientific text. It contains 292 annotated causal graphs from open-access MDPI and PLOS articles spanning diverse OpenAlex fields. Paper: Can Large Language Models Infer Causal Relationships from Real-World Text? GitHub: ReCITE Repository Dataset Configurations This dataset has… See the full description on the dataset page: https://huggingface.co/datasets/RyanSaklad/ReCITE.tabulartext-generation10K<n<100K0 likes602 downloads8mo agoHugging Face27REBOOT26 /sample_recovery-demonstrationThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "bi_widowxai_follower_robot", "total_episodes": 60, "total_frames": 53886, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:60" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/REBOOT26/sample_recovery-demonstration.tabularrobotics10K<n<100K0 likes573 downloads5mo agoHugging Face28AI4A-lab /RecruitViewgated 🎥 RecruitView: Multimodal Dataset for Personality & Interview Performance for Human Resources Applications Recorded Evaluations of Candidate Responses for Understanding Individual Traits 👋 Welcome to RecruitView We are excited to introduce RecruitView, a robust multimodal dataset designed to push the boundaries of affective computing, automated personality assessment, and soft-skill evaluation. In the realm of Human Resources and psychology, judging a candidate… See the full description on the dataset page: https://huggingface.co/datasets/AI4A-lab/RecruitView.tabularvideo-classification1K<n<10K21 likes563 downloads9d agoHugging Face29randalakab /Crop-recommendationtabular1K<n<10K0 likes553 downloads6mo agoHugging Face30open-index /ccrawl-recrawl-urls Common Crawl URL Recrawl Live refetches of pages from Common Crawl's URL index, with the body inline and the text already extracted What is it? Common Crawl publishes which URLs it saw and when, but the page bodies live in WARC archives that are awkward to query and are as old as the crawl that made them. This dataset takes the URL index for a single monthly crawl and fetches the pages again now, storing each response as one Parquet row with the body, the… See the full description on the dataset page: https://huggingface.co/datasets/open-index/ccrawl-recrawl-urls.tabulartext-generation1M<n<10M0 likes553 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.