CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ZomiLearner /English-Zomi-OPUS_Tatoeba_v20230412 English–Zomi Parallel Corpus (1.78M) This dataset contains 1.78 million English–Zomi sentence pairs, created to support machine translation, linguistic research, and large‑scale language model training. It is fully open and permissively licensed for commercial and non‑commercial use. 🌐 Linguistic Background: Zomi, Tedim Chin, and ISO Codes Zomi is the endonym (self‑chosen name) of the people and their language.However, Zomi does not yet have an official ISO 639‑3 code.… See the full description on the dataset page: https://huggingface.co/datasets/ZomiLearner/English-Zomi-OPUS_Tatoeba_v20230412.tabulartranslation1M<n<10M0 likes14k downloads7mo agoHugging Face02SetFit /enron_spamThis is a version of the Enron Spam Email Dataset, containing emails (subject + message) and a label whether it is spam or ham. tabular10K<n<100K21 likes6k downloads5y agoHugging Face03bs-modeling-metadata /c4-en-html-with-metadatatabular10M<n<100M14 likes2.8k downloads4y agoHugging Face04masoudjs /c4-en-html-with-metadata-ppl-cleanFile list: "c4-en-html_cc-main-2019-18_pq00-000.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-001.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-002.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-003.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-004.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-005.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-006.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-007.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-008.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-009.jsonl.gz"… See the full description on the dataset page: https://huggingface.co/datasets/masoudjs/c4-en-html-with-metadata-ppl-clean.tabular10K<n<100K1 likes1.1k downloads4y agoHugging Face05bs-modeling-metadata /c4-en-html-with-training_metadata_alltabular10K<n<100K1 likes996 downloads3y agoHugging Face06tripolskypetr /trading-entries TradingView Analyst Grading via Risk-Management Grid Sweep This dataset answers a single question: does an analyst's call differ from the market spread — and if so, at what cost. The cost here is the investor's risk management: how deep a stop they will tolerate, how many days their money stays frozen, and at what percentage they lock in profit. The basis is TradingView Ideas posts on crypto carrying a LONG / SHORT direction. Every post is swept across 21,280 points of a… See the full description on the dataset page: https://huggingface.co/datasets/tripolskypetr/trading-entries.tabulartime-series-forecasting10K<n<100K0 likes987 downloads2mo agoHugging Face07sasha /co2_energy_datatabularn<1K0 likes965 downloads3y agoHugging Face08marin-dna /gpn-star-p-uniform-v1-enhancer-arm-a marin-dna/gpn-star-p-uniform-v1-enhancer-arm-a Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This draft covers the enhancer region cohort with all species scope and preserves source FASTA/2bit letter case. Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus. Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-enhancer-arm-a.tabular100M<n<1B0 likes873 downloads29d agoHugging Face09marin-dna /zoonomia-v1-v4_ccre_noexon_enhancer bolinas-dna/zoonomia-v1-v4_ccre_noexon_enhancer A curated enhancer training set for issue #326 — a de-contaminated derivation of the v4 ccre_non_promoter arm of bolinas-dna/zoonomia-v1-v1, built by the snakemake/zoonomia_projection_dataset pipeline at commit 6b320c268547. Provenance This subset is v4_ccre_noexon further restricted to enhancer-dominant windows (dELS+pELS basepair coverage ≥ the other non-PLS cCRE classes), population-matching the val_enhancer… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_noexon_enhancer.tabular10M<n<100M0 likes858 downloads3mo agoHugging Face10marin-dna /functional-enhancer marin-dna/functional-enhancer Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and official version-matched UCSC hg38-to-target liftOver chains. This draft covers the enhancer region cohort with all species scope and preserves source FASTA/2bit letter case. Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus. For the 28 non-mammalian targets, the stable… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/functional-enhancer.tabular10M<n<100M0 likes630 downloads1mo agoHugging Face11marin-dna /phylop-uniform-v1-enhancer-arm-a marin-dna/phylop-uniform-v1-enhancer-arm-a Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This draft covers the enhancer region cohort with all species scope and preserves source FASTA/2bit letter case. Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus. Anchor eligibility uses the pipeline's pinned phyloP conservation… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/phylop-uniform-v1-enhancer-arm-a.tabular10M<n<100M0 likes600 downloads27d agoHugging Face12JetBrains-Research /EnvBench 🌱⚙️ EnvBench This repository contains data associated with EnvBench benchmark from EnvBench: A Benchmark for Automated Environment Setup. It contains: statistics about repositories from GitHub Search under ghs/data folder; several data splits under splits folder; Git repositories under repos folder; READMEs under readmes folder (available as readmes config); GitHub Actions workflows under workflows folder (available as workflows config); list of files in the repositories… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/EnvBench.tabular10K<n<100K0 likes542 downloads1y agoHugging Face13spade-rl /SPADE-Environment-Pool-GPT5.5-ToolUse SPARE GPT-5.5 Multi-Turn Tool-Use Games v1 A public static pool of 11,039 validated multi-turn tool-use environments generated by GPT-5.5 for SPARE actor training. Training alignment Source recipe: Qwen3-30B-A3B 0624 tool-use GAMES configuration 400 rollouts x 24 games/rollout = 9,600 no-reuse games required 11,039 validated games provide 1,439 games of headroom Six balanced skills: API orchestration, data retrieval, state modification, error recovery, tool… See the full description on the dataset page: https://huggingface.co/datasets/spade-rl/SPADE-Environment-Pool-GPT5.5-ToolUse.tabularreinforcement-learning10K<n<100K1 likes542 downloads29d agoHugging Face14THULab /so100_base_env SO-100 Base Environment (TsFile) Apache TsFile version of shreyasgite/so100_base_env. Overview A LeRobot teleoperation dataset recorded on an SO-100 arm performing a Lego pick-and-place task: "Grasp a lego block and put it in the bin." Each episode captures the synchronized robot joint state and commanded action at every control step. Robot: so100 (6-DoF arm: shoulder pan/lift, elbow flex, wrist flex/roll, gripper). Episodes: 252 (single train split). Frames: 98… See the full description on the dataset page: https://huggingface.co/datasets/THULab/so100_base_env.tabularroboticsn<1K0 likes518 downloads1mo agoHugging Face15flxclxc /encoded_drug_reviewstabular10K<n<100K10 likes481 downloads5y agoHugging Face16leharris3 /ccrfcd-mrms-hrrr-env-2021-2025 1H gauge accumulation + MRMS/HRRR zarr dataset for the Desert Southwest 50+ MRMS+HRRR variables; 220+ gauges; 400k samples NOTE: work in-progress. This is a dataset for training and evaluating synthetic quantitative precipiation estimation (QPE) models. Given some input context (e.g., radar fields, envionrmental parameters), predict how much rain fell at a rain gauge site over some period of time. Concretely, we've gather and QC'd data from 220 tipping bucket gauges through… See the full description on the dataset page: https://huggingface.co/datasets/leharris3/ccrfcd-mrms-hrrr-env-2021-2025.tabulartabular-regressionn<1K3 likes458 downloads6mo agoHugging Face17myduy /oscar-en-0-200tabular1M<n<10M1 likes408 downloads1y agoHugging Face18EngineeringAI-LAB /CineBoard3D-plus 🎬 CineBoard3D++: Dynamic 3D Story World Dataset 📊 Dataset Summary CineBoard3D++ is a collection of editable, movie-inspired 3D story worlds built with StoryBlender for narrative-grounded camera planning and world visual attention. It brings together story scripts, animated characters, scene geometry, and shot-level configurations in native Blender projects. The benchmark covers 50 stories, 457 scenes, 1,585 shots, and 3,197 3D assets (836 plot-related and 2,361… See the full description on the dataset page: https://huggingface.co/datasets/EngineeringAI-LAB/CineBoard3D-plus.3dn<1K0 likes372 downloads13d agoHugging Face19marin-dna /vertebrate-v1-issue473-fullwindow-ccre-enhancer-centered marin-dna/vertebrate-v1-issue473-fullwindow-ccre-enhancer-centered Review status: draft generated for issue #473 review before upload. Human-anchored 255 bp vertebrate sequences for the ccre_enhancer_centered cohort under the full_window policy. The policy projects the complete 255 bp human window, applies the established 128--512 bp compatible-fragment gate, and resizes around the accepted target-span midpoint. Human anchors come from the fixed exp351 ENCODE dELS/pELS-centered… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-issue473-fullwindow-ccre-enhancer-centered.tabular10M<n<100M0 likes365 downloads1mo agoHugging Face20marin-dna /zoonomia-v1-v4_ccre_noexon_enhancer-order bolinas-dna/zoonomia-v1-v4_ccre_noexon_enhancer-order The bolinas-dna/zoonomia-v1-v4_ccre_noexon_enhancer cross-mammal training set, restricted to a species cohort: one representative species per NCBI order — 19 deeply-diverged placental mammals (every pair separated by ~tens of millions of years), versus the implicit-default 108 family-deduplicated species. A strict subset of the family set, so it reuses the v1 cross-mammal projection unchanged (no re-halLiftover). Same human… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_noexon_enhancer-order.tabular10M<n<100M0 likes363 downloads3mo agoHugging Face21EntVista /metrixel-cmu-mocap-clean CMU MoCap — De-jittered & Foot-locked (BVH) 96 motion-capture clips from the CMU Graphics Lab Motion Capture Database, converted to BVH and processed to remove two artefacts present in the source solve: high-frequency rotational noise, and foot sliding during ground contact. Every clip ships with a .cleanup.json alongside it recording the before/after measurements, and the headline numbers are in metadata.jsonl too — so the processing is auditable rather than asserted. Across a… See the full description on the dataset page: https://huggingface.co/datasets/EntVista/metrixel-cmu-mocap-clean.tabularn<1K0 likes350 downloads5d agoHugging Face22simpleG2023 /chinese-clean-energy-battery-open-intelligence 🔬 Chinese Clean Energy, Battery Chemistry & Smart Grid Open Intelligence Dataset Curated open intelligence dataset tracking authentic Chinese scientific breakthroughs in Solid-State Battery chemistry, Perovskite Solar cells, Ultra-High Voltage (UHV) power grids, and industrial decarbonization. [!IMPORTANT] Data Completeness & Research Authenticity Notice: Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-clean-energy-battery-open-intelligence.tabulartext-retrieval1K<n<10K0 likes294 downloads3h agoHugging Face23OpenVoiceOS /ovos-stt-bench-voxpopuli-en-US OVOS stt bench — voxpopuli-en-US Per-clip transcripts predictions of the registered OVOS Plugin Arena stt fighters over facebook/voxpopuli. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena repo; the arena's assemble workflow turns these rows… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-voxpopuli-en-US.tabularn<1K0 likes269 downloads13d agoHugging Face24sxiong /entailmentbank EntailmentBank EntailmentBank is a dataset of multistep entailment trees for open-domain science question answering. Each example links a question and answer to a structured proof: a tree of multi-premise entailment steps from known facts, through intermediate conclusions, to a hypothesis. This repository contains the EMNLP 2021 v2 release in JSONL format, with four configs: Config Description task1 Generate an entailment-tree proof from gold supporting facts task2… See the full description on the dataset page: https://huggingface.co/datasets/sxiong/entailmentbank.tabularquestion-answering10K<n<100K1 likes240 downloads3mo agoHugging Face25MrJackTung /cs-envi-dual-encoder-60audion<1K0 likes237 downloads4mo agoHugging Face26Logics-MLLM /Logics-SWE-Env-2.5K Logics-SWE-Env-2.5K 2,553 software engineering task instances · 1,771 repositories · 4 programming languages 🤗 Related model: Logics-SWE-Qwen3.6-27B 📄 Paper: One to More, More to One 💻 GitHub: AgenticBigBang Overview What is this dataset? Logics-SWE-Env-2.5K is a collection of repository-level software engineering tasks for research on coding agents and environment-based reinforcement learning. It contains 2,553 unique task instances from 1,771… See the full description on the dataset page: https://huggingface.co/datasets/Logics-MLLM/Logics-SWE-Env-2.5K.tabulartext-generation1K<n<10K3 likes220 downloads15h agoHugging Face27OpenVoiceOS /ovos-stt-bench-ami-en-GB OVOS stt bench — ami-en-GB Per-clip transcripts predictions of the registered OVOS Plugin Arena stt fighters over edinburghcstr/ami. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena repo; the arena's assemble workflow turns these rows into… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-ami-en-GB.tabularn<1K0 likes219 downloads13d agoHugging Face28Intel /WEC-Eng WEC-Eng A large-scale dataset for cross-document event coreference extracted from English Wikipedia. Repository (Code for generating WEC): https://github.com/AlonEirew/extract-wec Paper: https://aclanthology.org/2021.naacl-main.198/ Languages English Load Dataset You can read in WEC-Eng files as follows (using the huggingface_hub library): from huggingface_hub import hf_hub_url, cached_download import json REPO_ID = "datasets/Intel/WEC-Eng" splits_files =… See the full description on the dataset page: https://huggingface.co/datasets/Intel/WEC-Eng.tabular100K<n<1M0 likes213 downloads5y agoHugging Face29marin-dna /vertebrate-v1-issue473-center1-ccre-enhancer-centered marin-dna/vertebrate-v1-issue473-center1-ccre-enhancer-centered Review status: draft generated for issue #473 review before upload. Human-anchored 255 bp vertebrate sequences for the ccre_enhancer_centered cohort under the center_1 policy. The policy projects the exact central human nucleotide, requires one target locus, and emits the 255 bp target window centered on that mapped nucleotide. Human anchors come from the fixed exp351 ENCODE dELS/pELS-centered catalog after… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-issue473-center1-ccre-enhancer-centered.tabular10M<n<100M0 likes194 downloads1mo agoHugging Face30Chess-Nut-Engine /chess-sft-eval Chess SFT Eval & Benchmark Held-out evaluation splits and a frozen benchmark for the Chess SFT training pipeline. Every FEN in these files is excluded from training data via a blocklist to guarantee zero contamination. Eval examples 13,000 Benchmark examples 13,000 Splits 9 (perception, rules, tactics, evaluation, openings, endgames, planning, chess960, mate) Format JSONL Training companion Chess-Nut-Engine/chess-sft-data How eval and benchmark differ… See the full description on the dataset page: https://huggingface.co/datasets/Chess-Nut-Engine/chess-sft-eval.tabulartext-generation10K<n<100K0 likes190 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.