CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lance-format /fineweb-edu FineWeb-Edu (Lance Format) A Lance-formatted version of FineWeb-Edu — over 1.5 billion educational web passages with cleaned text, source metadata, language detection signals, and 384-dim text embeddings — available directly from the Hub at hf://datasets/lance-format/fineweb-edu/data/train.lance. Key features Cleaned passage text in the text column with the source url and title carried alongside. Language detection signals (language, language_probability) for filtered… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/fineweb-edu.tabulartext-retrieval1B<n<10B8 likes8.4k downloads4mo agoHugging Face02jeggers /logiqa2_formatted Dataset Card for "logiqa2_formatted" More Information needed tabular10K<n<100K2 likes6k downloads2y agoHugging Face03EleutherAI /filtering-pretraining-mix-arrow-formattabular100M<n<1B0 likes2.6k downloads2y agoHugging Face04lance-format /Openvid-1M OpenVid Dataset (Lance Format) Lance format version of the OpenVid dataset with 937,957 high-quality videos stored with inline video blobs, embeddings, and rich metadata. Why Lance? Lance is an open-source format designed for multimodal AI data, offering significant advantages over traditional formats for modern AI workloads. Blazing Fast Random Access: Optimized for fetching scattered rows, making it ideal for random sampling, real-time ML serving, and interactive… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/Openvid-1M.tabulartext-to-video100K<n<1M8 likes2.5k downloads8mo agoHugging Face05Merserk /Krea-2-Turbo-Checkpoint-Format-Benchmark Krea 2 Turbo ComfyUI Format Fidelity Benchmark This release is a paired, deterministic comparison of eight Krea 2 Turbo checkpoint formats in ComfyUI: BF16, FP8 Scaled, INT8 ConvRot, MXFP8, NVFP4, INT4 ConvRot W4A4, GGUF Q8_0, and GGUF Q4_K_M. It contains 240 scored 1024×1024 images, saved float32 decoded tensors and final latents, every denoising trajectory, raw metric tables, telemetry, statistical comparisons, and reproduction code. Main result BF16 is the… See the full description on the dataset page: https://huggingface.co/datasets/Merserk/Krea-2-Turbo-Checkpoint-Format-Benchmark.imagetext-to-imagen<1K5 likes1.8k downloads3mo agoHugging Face06lance-format /openvid-lance OpenVid (Lance Format) A Lance-formatted version of the OpenVid-1M corpus — 937,957 high-quality clips with inline MP4 bytes, 1024-dim video embeddings, captions, and rich per-clip quality signals — available directly from the Hub at hf://datasets/lance-format/openvid-lance/data/train.lance. Key features Inline MP4 bytes in the video_blob column, stored in a side blob file and surfaced as lazy BlobFile handles via take_blobs — metadata scans, search, and filtering… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/openvid-lance.tabulartext-to-video100K<n<1M2 likes1.5k downloads4mo agoHugging Face07openforcefield /descent-format-geom Dataset Card for Meta-OMol25 Descent Formatted GEOM v1.0 Dataset Details Dataset Description Meta-OMol25 provides molecular structures, coordinates, energies, and forces, and we derived mapped SMILES for broad OpenFF parameter fitting workflows. This release is designed for general fitting and evaluation of van der Waals and valence terms. Curated by: Jennifer A Clark; jaclark5 Funded by: Open Force Field Initiative Shared by: Open Force Field Initiative, Open… See the full description on the dataset page: https://huggingface.co/datasets/openforcefield/descent-format-geom.tabular10M<n<100M0 likes561 downloads6mo agoHugging Face08allenai /rlvr-code-data-python-r1-format-filteredtabular10K<n<100K4 likes392 downloads1y agoHugging Face09sharkchill-xy /CodeContests_apps_format Dataset Card for "CodeContests_apps_format" More Information needed tabular10K<n<100K0 likes342 downloads3y agoHugging Face10timchen0618 /bcp-traj-ext-formatted-v1 bcp-traj-ext-formatted-v1 Trajectories from seed0 (gpt-oss-120b, Qwen3-Embedding-8B, full split) formatted in the traj_ext style: trajectory_text is the serialized steps ([Reasoning]/[Tool call]/[Tool result]/[Final answer]), and formatted_prompt is the full QUERY_TEMPLATE_GIVEN_TRAJECTORY prompt ready to feed to the next agent. Dataset Info Rows: 830 Columns: 9 Columns Column Type Description query_id Value('string') BrowseComp-Plus query ID… See the full description on the dataset page: https://huggingface.co/datasets/timchen0618/bcp-traj-ext-formatted-v1.tabularn<1K0 likes317 downloads6mo agoHugging Face11binhng /libero_object_mask_depth_IPEC_COMMUNITY_formatThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "franka", "total_episodes": 500, "total_frames": 74507, "total_tasks": 10, "total_videos": 4000, "total_chunks": 0, "chunks_size": 1000, "fps": 20, "splits": { "train": "0:500" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/binhng/libero_object_mask_depth_IPEC_COMMUNITY_format.tabularrobotics10K<n<100K0 likes295 downloads1y agoHugging Face12loubnabnl /kaggle_scripts_new_format_subset Dataset Card for "kaggle_scripts_new_format_subset" More Information needed tabular1M<n<10M0 likes276 downloads3y agoHugging Face13binhng /libero_10_mask_depth_IPEC_COMMUNITY_formatThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "franka", "total_episodes": 500, "total_frames": 138090, "total_tasks": 10, "total_videos": 4000, "total_chunks": 0, "chunks_size": 1000, "fps": 20, "splits": { "train": "0:500" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/binhng/libero_10_mask_depth_IPEC_COMMUNITY_format.tabularrobotics100K<n<1M0 likes253 downloads1y agoHugging Face14jablonkagroup /formation_energies Dataset Details Dataset Description Formation and decomposition energies of inorganic solids mined from the Materials Project database. Curated by: License: CC BY 4.0 Dataset Sources original data source Citation BibTeX: @article{Bartel_2020, doi = {10.1038/s41524-020-00362-y}, url = {https://doi.org/10.1038%2Fs41524-020-00362-y}, year = 2020, month = {jul}, publisher = {Springer Science and Business Media {LLC}}, volume = {6}… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/formation_energies.tabular1M<n<10M0 likes201 downloads1y agoHugging Face15binhng /libero_spatial_mask_depth_IPEC_COMMUNITY_format_version_2This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "franka", "total_episodes": 500, "total_frames": 62250, "total_tasks": 10, "total_videos": 4000, "total_chunks": 0, "chunks_size": 1000, "fps": 20, "splits": { "train": "0:500" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/binhng/libero_spatial_mask_depth_IPEC_COMMUNITY_format_version_2.tabularrobotics10K<n<100K0 likes178 downloads1y agoHugging Face16Bin12345 /screenspot_pro_arrow_formattabular1K<n<10K0 likes163 downloads1y agoHugging Face17jonathanyin /aime_1983_2023_grok-3-mini-high_traces_r1_formattedtabularn<1K0 likes159 downloads1y agoHugging Face18aklein4 /Bitext-SmolLM2-1024-natural-instructions-formattabularn<1K0 likes153 downloads3mo agoHugging Face19binhng /libero_spatial_mask_depth_IPEC_COMMUNITY_formatThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "franka", "total_episodes": 10, "total_frames": 1271, "total_tasks": 10, "total_videos": 80, "total_chunks": 0, "chunks_size": 1000, "fps": 20, "splits": { "train": "0:10" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/binhng/libero_spatial_mask_depth_IPEC_COMMUNITY_format.tabularrobotics10K<n<100K0 likes147 downloads1y agoHugging Face20jensjepsen /danish-icl-schema-format-v3 danish-icl-json-v3 In-context-learning rows derived from jensjepsen/danish-json-grpo-v1. Each row packs 1-5 worked examples into a single user turn, followed by a held-out passage; the assistant turn is the answer for that passage. No instruction is included, so both the schema and the output format have to be inferred from the examples. Two axes vary per row and are held constant within a row: the schema (134 field-sets) and the output format (10 renderers — JSON, key: value… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-icl-schema-format-v3.tabulartext-generation10K<n<100K0 likes139 downloads29d agoHugging Face21openforcefield /descent-format-ani2x Dataset Card for Meta-OMol25 Descent Formatted ANI2X v1.0 Dataset Details Dataset Description Meta-OMol25 provides molecular structures, coordinates, energies, and forces, and we derived mapped SMILES for broad OpenFF parameter fitting workflows. This release is designed for general fitting and evaluation of van der Waals and valence terms. Curated by: Jennifer A Clark; jaclark5 Funded by: Open Force Field Initiative Shared by: Open Force Field… See the full description on the dataset page: https://huggingface.co/datasets/openforcefield/descent-format-ani2x.tabular10M<n<100M0 likes132 downloads6mo agoHugging Face22yuanhezhang /DAG-MATH-Formatted-CoT Benchmark Overview This dataset card contains 2,894 gold-standard DAG-MATH formatted CoT from problems from Omni-MATH. Top‑Level Schema Each JSON file is a list with a single object describing the problem: problem_id: integer identifier of the problem. domain: list of strings describing the topic taxonomy. difficulty: numeric difficulty indicator from 1 (easiest) to 6 (hardest). problem_text: problem statement. sample_id: sample identifier for the solution trace.… See the full description on the dataset page: https://huggingface.co/datasets/yuanhezhang/DAG-MATH-Formatted-CoT.tabular1K<n<10K1 likes125 downloads11mo agoHugging Face23PrimeIntellect /Skywork-OR1-RL-Data-v1-math-prime-rl-formattabular100K<n<1M0 likes117 downloads1y agoHugging Face24binhng /libero_small_debug_mask_depth_IPEC_COMMUNITY_formatThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "franka", "total_episodes": 100, "total_frames": 12194, "total_tasks": 2, "total_videos": 800, "total_chunks": 0, "chunks_size": 1000, "fps": 20, "splits": { "train": "0:100" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/binhng/libero_small_debug_mask_depth_IPEC_COMMUNITY_format.tabularrobotics10K<n<100K0 likes114 downloads1y agoHugging Face25binhng /libero_goal_mask_depth_IPEC_COMMUNITY_formatThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "franka", "total_episodes": 500, "total_frames": 63728, "total_tasks": 10, "total_videos": 4000, "total_chunks": 0, "chunks_size": 1000, "fps": 20, "splits": { "train": "0:500" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/binhng/libero_goal_mask_depth_IPEC_COMMUNITY_format.tabularrobotics10K<n<100K0 likes108 downloads1y agoHugging Face26open-athena /wildchat-glm53-format-completions WildChat format completions 9,975 GLM-5.3 answers across 29 parseable formats. Each answer passed its contract verifier. Failed answers were resampled with the same prompt until one passed; no semantic judge or answer repair was used. This is the final release from a 10,000-prompt run; 25 unfinished prompts were excluded. It stores answers, format instructions, exact contracts, and pinned WildChat-4.8M references—not the source prompts or conversations. All rows are in the train… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/wildchat-glm53-format-completions.tabulartext-generation1K<n<10K1 likes104 downloads10d agoHugging Face27Kyle1668 /mmlu_auxiliary_train_formattedtabular10K<n<100K0 likes103 downloads1y agoHugging Face28lance-format /soarm101-feeding-nuts-lancetabularn<1K0 likes102 downloads1mo agoHugging Face29WebOrganizer /FormatAnnotations-Llama-3.1-8B WebOrganizer/FormatAnnotations-Llama-3.1-8B [Paper] [Website] [GitHub] This dataset contains 1M web pages annotated with format/type labels by the Llama-3.1-8B model. The web pages are a sample of the DCLM RefinedWeb reproduction. It is used as first-stage training data for the WebOrganizer/FormatClassifier. Dataset Structure Each example contains the following fields: text: The text content of the web page url: The URL of the web page top_choice_index: Index of the… See the full description on the dataset page: https://huggingface.co/datasets/WebOrganizer/FormatAnnotations-Llama-3.1-8B.tabular1M<n<10M1 likes96 downloads2y agoHugging Face30Nhanvi282 /openvivqa-formating-vlm OpenViVQA Formatting Dataset for VLM A Vietnamese multimodal instruction-format dataset for training Vision Language Models (VLMs) on Visual Question Answering (VQA) tasks. This dataset reformats OpenViVQA-style samples into conversational instruction-tuning format compatible with modern VLM training pipelines such as: Qwen2-VL LLaVA InternVL Phi-3 Vision Idefics SmolVLM Dataset Structure Each sample contains: image: input image conversations: multi-turn… See the full description on the dataset page: https://huggingface.co/datasets/Nhanvi282/openvivqa-formating-vlm.imagevisual-question-answering10K<n<100K1 likes84 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.