CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nexar-ai /nexar_collision_predictiongated Nexar Collision Prediction Dataset This dataset is part of the Nexar Dashcam Crash Prediction Challenge on Kaggle. Dataset The Nexar collision prediction dataset comprises videos from Nexar dashcams. Videos have a resolution of 1280x720 at 30 frames per second and typically have about 40 seconds of duration. The dataset contains 1500 videos where half show events where there was a collision or a collision was eminent (positive cases), and the other half shows… See the full description on the dataset page: https://huggingface.co/datasets/nexar-ai/nexar_collision_prediction.tabularvideo-classification1K<n<10K21 likes15k downloads2h agoHugging Face02Nexusflow /NexusRaven_API_evaluation NexusRaven API Evaluation dataset Please see blog post or NexusRaven Github repo for more information. License The evaluation data in this repository consists primarily of our own curated evaluation data that only uses open source commercializable models. However, we include general domain data from the ToolLLM and ToolAlpaca papers. Since the data in the ToolLLM and ToolAlpaca works use OpenAI's GPT models for the generated content, the data is not commercially… See the full description on the dataset page: https://huggingface.co/datasets/Nexusflow/NexusRaven_API_evaluation.text1K<n<10K17 likes12k downloads3y agoHugging Face03lmms-lab /LLaVA-NeXT-Data Dataset Card for LLaVA-NeXT We provide the whole details of LLaVA-NeXT Dataset. In this dataset, we include the data that was used in the instruction tuning stage for LLaVA-NeXT and LLaVA-NeXT(stronger). Aug 30, 2024: We update the dataset with raw format (de-compress it for json file and images with structured folder), you can directly download them if you are familiar with LLaVA data format. Dataset Sources Compared to the instruction data mixture for LLaVA-1.5… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/LLaVA-NeXT-Data.image100K<n<1M47 likes3.7k downloads2y agoHugging Face04lmms-eval /NExTQAtabular10K<n<100K6 likes3.6k downloads2y agoHugging Face05NexusProjectsAI /Nexus-Agents-ToolCalling Nexus Agents — Tool-Calling Conversations Synthetic, schema-verified tool-calling conversations for training the Nexus Projects agents. This is the exact data behind Nemotron-3-Nano-30B-A3B — Nexus Agents (GGUF), including the verification transcripts that scored it (27/27 on the behavioral interview eval, vs 13/27 for the base model). Links: the fine-tuned model → Nemotron-3-Nano-30B-A3B — Nexus Agents (GGUF) · the generator + seed data + eval harness → Nexus Training Studio ·… See the full description on the dataset page: https://huggingface.co/datasets/NexusProjectsAI/Nexus-Agents-ToolCalling.texttext-generation100K<n<1M1 likes3k downloads3mo agoHugging Face06lmms-lab-encoder /LLaVA-NeXT-Interleave-Bench LLaVA-Interleave Bench Dataset Card Dataset details Dataset type: LLaVA-Interleave Bench is a comprehensive set of multi-image datasets that are collected from public datasets or generated by the GPT-4V API. It is constructed for evaluating the interleaved multi-image reaoning capbilities of LMMs. Dataset date: LLaVA-Interleave Bench was collected in April 2024, and released in June 2024. Paper or resources for more information: Blog:… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/LLaVA-NeXT-Interleave-Bench.imagevisual-question-answering10K<n<100K15 likes2.5k downloads2y agoHugging Face07VLM2Vec /NExTQAtabular10K<n<100K1 likes2.5k downloads2y agoHugging Face08caskcsg /NExtLong-512K-dataset NExtLong: Toward Effective Long-Context Training without Long Documents This repository contains the code ,models and datasets for our paper NExtLong: Toward Effective Long-Context Training without Long Documents. [Github] Quick Links Overview NExtLong Models NExtLong Datasets Datasets list How to use NExtLong datasets Bugs or Questions? Overview Large language models (LLMs) with extended context windows have made significant strides yet remain a… See the full description on the dataset page: https://huggingface.co/datasets/caskcsg/NExtLong-512K-dataset.text10K<n<100K1 likes1.3k downloads1y agoHugging Face09nexacore /solana-dex-datatabular1M<n<10M0 likes1.2k downloads7mo agoHugging Face10Gibrail765 /Nexus_Ulawengtextn<1K2 likes1.2k downloads44m agoHugging Face11dsfsi-anv /za-african-next-voicesgated Swivuriso: ZA-African Next Voices Swivuriso is a large-scale multilingual speech dataset targeting over 3000 hours of audio across 7 South African languages. The dataset is developed to support Automatic Speech Recognition (ASR) and inclusive speech technologies for low-resource African languages. It combines both scripted and unscripted speech, collected through ethical, community-centered processes. Dataset Paper: ArXiv - Work in Progress Language Coverage… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi-anv/za-african-next-voices.audioautomatic-speech-recognition100K<n<1M16 likes1.1k downloads7mo agoHugging Face12OALL /details_Nexusflow__Athene-70B Dataset Card for Evaluation run of Nexusflow/Athene-70B Dataset automatically created during the evaluation run of model Nexusflow/Athene-70B. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Nexusflow__Athene-70B.tabular100K<n<1M0 likes1.1k downloads2y agoHugging Face13Yash514311 /nexus-jobstext10K<n<100K1 likes1k downloads3h agoHugging Face14Ardea /NEXUS-temporal_hierarchical_multi-modal NEXUS: Neural Evolution for eXtensible Universal Semantics Dataset (Temporal Multimodal Slices) This dataset is a multi-modal, hierarchical, temporal representation derived from HuggingFaceFV/finevideo. It is designed for streaming training where the primary unit is a 10 ms "slice" that aggregates upward into moments (100 ms), seconds (1 s), experiences (10 s), and minutes (60 s). It is meant to represent an extensible stream of "experience" as there are… See the full description on the dataset page: https://huggingface.co/datasets/Ardea/NEXUS-temporal_hierarchical_multi-modal.imageautomatic-speech-recognition10M<n<100M5 likes916 downloads4mo agoHugging Face15caskcsg /NExtLong-64K-dataset NExtLong: Toward Effective Long-Context Training without Long Documents This repository contains the code ,models and datasets for our paper NExtLong: Toward Effective Long-Context Training without Long Documents. [Github] Quick Links Overview NExtLong Models NExtLong Datasets Datasets list How to use NExtLong datasets Bugs or Questions? Overview Large language models (LLMs) with extended context windows have made significant strides yet remain a… See the full description on the dataset page: https://huggingface.co/datasets/caskcsg/NExtLong-64K-dataset.text10K<n<100K2 likes840 downloads1y agoHugging Face16caskcsg /NExtLong-128K-dataset NExtLong: Toward Effective Long-Context Training without Long Documents This repository contains the code ,models and datasets for our paper NExtLong: Toward Effective Long-Context Training without Long Documents. [Github] Quick Links Overview NExtLong Models NExtLong Datasets Datasets list How to use NExtLong datasets Bugs or Questions? Overview Large language models (LLMs) with extended context windows have made significant strides yet remain a… See the full description on the dataset page: https://huggingface.co/datasets/caskcsg/NExtLong-128K-dataset.text10K<n<100K1 likes773 downloads1y agoHugging Face17Linear-Next /Linear-Next-Datasets Linear Next Benchmark Linear Next is a comprehensive benchmark designed to fairly compare various efficient transformer architectures. This project evaluates different approaches including linear attention, sparse attention, and other model structures under identical training conditions and datasets. Overview The benchmark aims to provide an unbiased comparison of efficient transformer variants by ensuring all models are trained with the same datasets, hyperparameters… See the full description on the dataset page: https://huggingface.co/datasets/Linear-Next/Linear-Next-Datasets.text100M<n<1B0 likes761 downloads1y agoHugging Face18TIGER-Lab /SWE-Next SWE-Next: Scalable Real-World Software Engineering Tasks for Agents SWE-Next Dataset SWE-Next is an execution-grounded dataset of 2,308 self-verifying software engineering tasks mined from real merged GitHub pull requests. Starting from 3,971 seeded Python repositories and 102,582 executed candidate base/merged commit pairs, SWE-Next retains only instances where the merged commit produces a strict test improvement without regressions. The final release… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/SWE-Next.texttext-generation1K<n<10K1 likes685 downloads5mo agoHugging Face19Na0s /Next_Token_Prediction_datasettext1M<n<10M0 likes612 downloads2y agoHugging Face20NEXTAltair /genai-image-tag-db GenAI Image Tag DB (cc0-1.0) This repository contains the cc0-1.0 build of the tag database. The main artifact is the SQLite database. The parquet_danbooru/ directory is a derived export so the Hugging Face Dataset Viewer can preview a subset of rows (Danbooru-only). Files genai-image-tag-db-cc0.sqlite: SQLite database parquet_danbooru/*.parquet: Parquet export for Dataset Viewer build_manifest.json: Build manifest (revisions and stats) report/: Source effects… See the full description on the dataset page: https://huggingface.co/datasets/NEXTAltair/genai-image-tag-db.tabulartext-retrieval1M<n<10M1 likes569 downloads3d agoHugging Face21AlayaNeW /LLaVA-NeXT-Data Dataset Card for LLaVA-NeXT We provide the whole details of LLaVA-NeXT Dataset. In this dataset, we include the data that was used in the instruction tuning stage for LLaVA-NeXT and LLaVA-NeXT(stronger). Aug 30, 2024: We update the dataset with raw format (de-compress it for json file and images with structured folder), you can directly download them if you are familiar with LLaVA data format. Dataset Sources Compared to the instruction data mixture for LLaVA-1.5… See the full description on the dataset page: https://huggingface.co/datasets/AlayaNeW/LLaVA-NeXT-Data.image100K<n<1M0 likes515 downloads1y agoHugging Face22nexoneAB /swedish-legal-decisions-raw-v1 Swedish Court Decisions — Svenska Domstolsavgöranden 55,096 court decisions spanning 45 years of Swedish case law, purpose-built for LLM training. The most comprehensive open dataset of Swedish appellate court decisions available for AI development. Sourced directly from the official Swedish Courts case law database via their public REST API and preprocessed into three ready-to-use training configurations. Why This Dataset Scale and depth: 55,096 decisions covering… See the full description on the dataset page: https://huggingface.co/datasets/nexoneAB/swedish-legal-decisions-raw-v1.texttext-generation10K<n<100K0 likes485 downloads7mo agoHugging Face23fan-shu /swe-mt-combined-coderforge-hero-lego-nex-swezero fan-shu/swe-mt-combined-coderforge-hero-lego-nex-swezero Concatenated mid-train dataset for Qwen3 Thinking SFT. Each source subset is loaded in order and concatenated into a single config so one training epoch visits every trajectory exactly once (no interleave / no oversampling). Built from fan-shu/swe-instruct-trajectories-empty-think-inserted. Source subsets (7) togethercomputer__CoderForge-Preview nvidia__SWE-Zero-openhands-trajectories nex-agi__agent-sft… See the full description on the dataset page: https://huggingface.co/datasets/fan-shu/swe-mt-combined-coderforge-hero-lego-nex-swezero.text100K<n<1M0 likes471 downloads2mo agoHugging Face24NEXTAltair /genai-image-tag-db-CC4 GenAI Image Tag DB (cc-by-4.0) This repository contains the cc-by-4.0 build of the tag database. The main artifact is the SQLite database. The parquet_danbooru/ directory is a derived export so the Hugging Face Dataset Viewer can preview a subset of rows (Danbooru-only). Files genai-image-tag-db-cc4.sqlite: SQLite database parquet_danbooru/*.parquet: Parquet export for Dataset Viewer build_manifest.json: Build manifest (revisions and stats) report/: Source… See the full description on the dataset page: https://huggingface.co/datasets/NEXTAltair/genai-image-tag-db-CC4.tabulartext-retrieval1M<n<10M0 likes464 downloads3mo agoHugging Face25physicl-community /fast-food-floor-waste-grasping-training-set-next-pack-9f7b7681-1106dcde Fast-Food Cleaning Robot — Floor Mess Dataset Training dataset for a cleaning robot operating in fast-food-style food-service spaces (break areas / dining). Scenes are staged in break-area environments cluttered with food-service furnishings and food items (pizza, grocery food, cups, spoons) so the robot learns to perceive and act on mess. Covers detection, grasping, navigation, obstacle avoidance and pick-and-place. Renders are 1024x1024 with RGB plus albedo, metric depth and… See the full description on the dataset page: https://huggingface.co/datasets/physicl-community/fast-food-floor-waste-grasping-training-set-next-pack-9f7b7681-1106dcde.imagen<1K0 likes436 downloads20d agoHugging Face26ShinoharaHare /LLaVA-NeXT-Data-Reformattedimage100K<n<1M1 likes399 downloads2y agoHugging Face27NEXTAltair /genai-image-tag-db-mit GenAI Image Tag DB (mit) This repository contains the mit build of the tag database. The main artifact is the SQLite database. The parquet_danbooru/ directory is a derived export so the Hugging Face Dataset Viewer can preview a subset of rows (Danbooru-only). Files genai-image-tag-db-mit.sqlite: SQLite database parquet_danbooru/*.parquet: Parquet export for Dataset Viewer build_manifest.json: Build manifest (revisions and stats) report/: Source effects and health… See the full description on the dataset page: https://huggingface.co/datasets/NEXTAltair/genai-image-tag-db-mit.tabulartext-retrieval1M<n<10M0 likes384 downloads3mo agoHugging Face28zelalt /llava-next-data-400k Source & citation This subset is derived from lmms-lab/LLaVA-NeXT-Data. @misc{liu2024llavanext, title={LLaVA-NeXT: Improved reasoning, OCR, and world knowledge}, url={https://llava-vl.github.io/blog/2024-01-30-llava-next/}, author={Liu, Haotian and Li, Chunyuan and Li, Yuheng and Li, Bo and Zhang, Yuanhan and Shen, Sheng and Lee, Yong Jae}, month={January}, year={2024} } image100K<n<1M0 likes382 downloads3mo agoHugging Face29ArkAiLab-Adl /Nexora-music-pd-v1-mediumaudiotext-to-audion<1K3 likes365 downloads9mo agoHugging Face30AtomicChat /Qwen3.8-Flash-Next-GGUF-metricstabularn<1K0 likes346 downloads28d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.