CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01china-ai-law-challenge /cail2018 Dataset Card for CAIL 2018 Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation Rationale [More Information Needed] Source Data… See the full description on the dataset page: https://huggingface.co/datasets/china-ai-law-challenge/cail2018.tabularother1M<n<10M31 likes1.2k downloads3y agoHugging Face02Cainiao-AI /LaDe-D 1. About Dataset LaDe is a publicly available last-mile delivery dataset with millions of packages from industry. It has three unique characteristics: (1) Large-scale. It involves 10,677k packages of 21k couriers over 6 months of real-world operation. (2) Comprehensive information, it offers original package information, such as its location and time requirements, as well as task-event information, which records when and where the courier is while events such as task-accept and… See the full description on the dataset page: https://huggingface.co/datasets/Cainiao-AI/LaDe-D.tabular1M<n<10M4 likes678 downloads3y agoHugging Face03lerobot /columbia_cairlab_pusht_realThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "unknown", "total_episodes": 136, "total_frames": 27808, "total_tasks": 1, "total_videos": 272, "total_chunks": 1, "chunks_size": 1000, "fps": 10, "splits": { "train": "0:136" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/columbia_cairlab_pusht_real.tabularrobotics10K<n<100K1 likes421 downloads1y agoHugging Face04cairocode /IEMO_Audio_Text_Mergedtabular1K<n<10K0 likes338 downloads11mo agoHugging Face05cairocode /MSPI_Audio_Text_Mergedtabular1K<n<10K0 likes293 downloads11mo agoHugging Face06Cainiao-AI /LaDe-P 1. About Dataset LaDe is a publicly available last-mile delivery dataset with millions of packages from industry. It has three unique characteristics: (1) Large-scale. It involves 10,677k packages of 21k couriers over 6 months of real-world operation. (2) Comprehensive information, it offers original package information, such as its location and time requirements, as well as task-event information, which records when and where the courier is while events such as task-accept and… See the full description on the dataset page: https://huggingface.co/datasets/Cainiao-AI/LaDe-P.tabular1M<n<10M4 likes281 downloads3y agoHugging Face07CAiRE /belief_r Belief Revision: The Adaptability of Large Language Models Reasoning This is the official dataset for the paper "Belief Revision: The Adaptability of Large Language Models Reasoning", published in the main conference of EMNLP 2024. 📚 Data | 📃 Paper Overview The capability to reason from text is crucial for real-world NLP applications. Real-world scenarios often involve incomplete or evolving data. In response, individuals update their beliefs and… See the full description on the dataset page: https://huggingface.co/datasets/CAiRE/belief_r.tabular1K<n<10K1 likes275 downloads2y agoHugging Face08caiotheodoro /recon-eval ReconEval — Financial Reconciliation Benchmark Reading results from this benchmark. Four properties of ReconEval shape what a score on it means. Anyone comparing models here should know them. One class can dominate a margin. PARTIAL_MATCH is the highest-variance class between models, and its 32 evaluation items are generated from 9 abbreviation pairs — all of which also appear in the training split, overlap fraction 1.0. On this set, "learned the concept" and "memorised nine… See the full description on the dataset page: https://huggingface.co/datasets/caiotheodoro/recon-eval.tabulartext-generation1K<n<10K0 likes246 downloads25d agoHugging Face09starknet-ai /cairo-security-audits Cairo Security Audits A source-traceable corpus of public Cairo and Starknet security-audit metadata and normalized finding annotations. Version 0.3.0 packages every entry in the audit inventory frozen at keep-starknet-strange/starknet-skills@17a76e8. It covers 32 accessible reports from 10 auditing firms and 286 normalized finding annotations. Eleven records are checked against rendered reports and two link to exact vulnerable/fixed commits. The release does not redistribute… See the full description on the dataset page: https://huggingface.co/datasets/starknet-ai/cairo-security-audits.tabulartext-retrievaln<1K1 likes218 downloads1mo agoHugging Face10caiotheodoro /vernier vernier Error bars on a dataset vendor's quality claim. Build AI publishes hand-visibility and active-manipulation rates for Egocentric-10K / Egocentric-100K, judged once by gemini-2.5-flash with no human gold, no interval, and no test that the judge scores a factory floor and a home kitchen on the same scale. This release is the data behind an independent, pre-registered measurement of that claim: human labels against a written rubric, a live open-weights judge on the same… See the full description on the dataset page: https://huggingface.co/datasets/caiotheodoro/vernier.tabularimage-classification10K<n<100K0 likes180 downloads18d agoHugging Face11caiiii /so100_test_0109This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "so100", "total_episodes": 50, "total_frames": 20327, "total_tasks": 1, "total_videos": 100, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/caiiii/so100_test_0109.tabularrobotics10K<n<100K0 likes151 downloads2y agoHugging Face12amu-cai /llmzszl-dataset Dataset description This dataset is a collection of multiple-choice questions, each with one correct answer, sourced from exams conducted within the Polish education system renging from years 2002 up to 2024. The tests are published annually by the Polish Central Examination Board. Questions in the dataset are categorized into the following categories: math, natural sciences, biology, physics, and the Polish language for pre-high school questions, and arts, mechanics (including… See the full description on the dataset page: https://huggingface.co/datasets/amu-cai/llmzszl-dataset.tabular10K<n<100K1 likes138 downloads2y agoHugging Face13polinaeterna /cail2018tabular1M<n<10M0 likes131 downloads3y agoHugging Face14caiotheodoro /lossbench-finance-v1 LossBench finance-v1 Severity-weighted expected-loss evaluation for agents that touch money. Three finance back-office domains, mechanical ground truth, and a contamination certificate. Models are ranked by what their mistakes cost, not by raw accuracy. Overview Task count 2400 Domains reconciliation, payment_repair, settlement License cc-by-4.0 Tasks Each task is an agentic back-office scenario with a deterministic seed, an… See the full description on the dataset page: https://huggingface.co/datasets/caiotheodoro/lossbench-finance-v1.tabulartext-generation1K<n<10K0 likes127 downloads1mo agoHugging Face15caiotheodoro /plumb Plumb Gold tasks and the held-out eval. The study is on the collection. Config n What benchmark 1000 Held-out eval, seed 777. Never in train. train_handseeded 223 Mix matched to the eval, including PASS. train_ornith 58 Ornith-1.5 proposals that passed the oracle. train_blended 281 Both of the above. Leakprobe vs benchmark: exact signature overlap 0. curriculum train n sw-recall precision exact hand-seeded 223 0.318 [0.290, 0.347] 0.308 [0.279… See the full description on the dataset page: https://huggingface.co/datasets/caiotheodoro/plumb.tabulartext-generation1K<n<10K0 likes123 downloads1mo agoHugging Face16caiotheodoro /titer-edgar-officers titer · EDGAR officer corpus 4,206,080 attested person–company–role–date tuples from SEC Forms 3/4/5, published as pointers rather than records, alongside the frozen pre-registrations that were hash-published before any measurement ran. edgar_officers.parquet: 4.2M rows, 230,405 distinct people, 20,266 issuers, 2006q1–2026q2. Column Meaning accession SEC accession number, the pointer that reconstructs the row person_cik Reporting-owner CIK. Never recycled by the… See the full description on the dataset page: https://huggingface.co/datasets/caiotheodoro/titer-edgar-officers.tabular1M<n<10M0 likes110 downloads20d agoHugging Face17Vision-CAIR /InfiniBench_challlenge Download git lfs install git clone https://huggingface.co/datasets/Vision-CAIR/InfiniBench_challlenge cd InfiniBench_challlenge rm -rf .git/ How to Decompress the Videos To extract the videos from the compressed files, follow these steps: Open a terminal and navigate to the videos directory: cd videos Combine all the split archive parts into a single .tar.gz file: cat test_videos.tar.gz.part_* > test_videos.tar.gz Extract the contents of the archive: tar -xvf… See the full description on the dataset page: https://huggingface.co/datasets/Vision-CAIR/InfiniBench_challlenge.tabular1K<n<10K0 likes106 downloads1y agoHugging Face18Arimancy /caiso-demand-weather CAISO Hourly Electricity Demand + Population-Weighted Weather (2015 to 2026) A model ready hourly time series for California grid load forecasting: CAISO balancing authority demand (MWh) pre joined to population weighted weather, with calendar, holiday, degree hour, irradiance, and cloud cover features. 98,080 hours from 2015-07-01 to 2026-09-07, no missing hours, every imputed or preliminary value explicitly flagged, in Parquet and CSV with an easy chronological split. The… See the full description on the dataset page: https://huggingface.co/datasets/Arimancy/caiso-demand-weather.tabular10K<n<100K2 likes106 downloads5d agoHugging Face19caiotheodoro /cyclegraph-flow-gain cyclegraph flow gain How much of a hand's motion a dense optical flow estimator actually recovers, and what it reports instead once it stops. Rendered under the corpus's own fisheye, where the true flow field is known exactly, which is the only reason a gain is measurable at all. The finding, in one line: past a displacement knee the estimator reports the background, at a gain equal to hand distance over background distance (0.45 / 2.5 = 0.18) — and the ego-motion correction… See the full description on the dataset page: https://huggingface.co/datasets/caiotheodoro/cyclegraph-flow-gain.tabularn<1K0 likes84 downloads16d agoHugging Face20justicedao /ipfs_turks_caicos_laws_ir Turks Caicos legislation IR (CID-keyed sparse GraphRAG) Research retrieval release of endomorphosis/ipfs_turks_caicos_laws (revision aaafe4d3bed911181c42e6eec64bc350147c5c4a) packaged as country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir). Not legal advice. This is a research snapshot. The official gazette / authentic source of Turks Caicos prevails over this corpus. Retrieved documents and graph edges are retrieval evidence only. No… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_turks_caicos_laws_ir.tabulartext-retrieval100K<n<1M0 likes81 downloads1d agoHugging Face21CAiRE /prosocial-dialog-jpn_Jpantabular100K<n<1M0 likes75 downloads3y agoHugging Face22caiotheodoro /titer-expertise-claims titer · attested expertise claims 99,984 expertise claims attested by publication record, for 20,000 researchers with an ORCID iD. Built to measure whether a people-search provider can tell a real expert from a claimed one. Sources: OpenAlex (disambiguated authors, works, topics), ORCID (the identity spine), Crossref DOIs (the attestation chain). All open. A researcher cannot self-assert a DOI into existence, which is why authorship is treated as attested where a self-reported… See the full description on the dataset page: https://huggingface.co/datasets/caiotheodoro/titer-expertise-claims.tabular10K<n<100K0 likes74 downloads19d agoHugging Face23caiotheodoro /suture Suture Gold tasks for policy-issuance QC: a structured underwriting binder, a structured issued policy, and the exact discrepancy set. Images are not stored. Re-render with suture_forge.generate.render_doc from the GitHub repo. Paper and adapter: caiotheodoro/suture · caiotheodoro/suture-8b. Predictions on this gold: caiotheodoro/suture-evals. Splits Config Seed n What benchmark 777 1000 Contracted held-out set. Never in train. validation 7 holdout 80… See the full description on the dataset page: https://huggingface.co/datasets/caiotheodoro/suture.tabularimage-text-to-text1K<n<10K0 likes72 downloads1mo agoHugging Face24caiyuhu /MCiteBench MCiteBench Dataset MCiteBench is a benchmark for evaluating the ability of Multimodal Large Language Models (MLLMs) to generate text with citations in multimodal contexts. Websites: https://caiyuhu.github.io/MCiteBench Paper: https://arxiv.org/abs/2503.02589 Code: https://github.com/caiyuhu/MCiteBench Data Download Please download the MCiteBench_full_dataset.zip. It contains the data.jsonl file and the visual_resources folder. Data Statistics… See the full description on the dataset page: https://huggingface.co/datasets/caiyuhu/MCiteBench.imagetext-generation1K<n<10K0 likes70 downloads1y agoHugging Face25CAiRE /prosocial-dialog-eng_Latn-Mistral-7B-Instruct-v0.2tabular100K<n<1M0 likes56 downloads3y agoHugging Face26caiiii /so100_test_0102_1909This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "so100", "total_episodes": 5, "total_frames": 3681, "total_tasks": 1, "total_videos": 10, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:5" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/caiiii/so100_test_0102_1909.tabularrobotics1K<n<10K0 likes54 downloads2y agoHugging Face27caiyuchen /cryptoalpha-paneltabular1M<n<10M0 likes54 downloads3mo agoHugging Face28CAIR-M3LLM /OsteosarcomaTumorAssessment Osteosarcoma data from UT Southwestern/UT Dallas for Viable and Necrotic Tumor Assessment (Osteosarcoma-Tumor-Assessment) Unofficial fork. Folder Structure ML_Features_1144.csv # Contains 1144 rows for all the image tiles and 69 columns for filename, classification, and 65 machine learning features. OsteosarcomaTumorAssessment.tar.zst |-- Osteosarcoma-UT.sums |-- Training-Set-1 # 11 folders with 547 images. Each folder contains 48~50 image tiles and 1 csv for… See the full description on the dataset page: https://huggingface.co/datasets/CAIR-M3LLM/OsteosarcomaTumorAssessment.tabularimage-classification1K<n<10K0 likes53 downloads2mo agoHugging Face29regicid /cairn-metadata Cairn.info — Revues SHS & métadonnées d'articles Ce dataset regroupe deux fichiers, construits à partir de Cairn.info (portail de revues en sciences humaines et sociales). Ils sont déclarés comme deux configurations distinctes (revues et articles, cf. bandeau ci-dessus / sélecteur du viewer) — ne pas les charger ensemble comme un seul dataset, leurs colonnes ne correspondent pas : cairn_revues.csv Catalogue croisé des 672 revues indexées sur Cairn (onglet Revues)… See the full description on the dataset page: https://huggingface.co/datasets/regicid/cairn-metadata.tabular100K<n<1M0 likes53 downloads5d agoHugging Face30caioloures /personal_dictionary OpenGloss Dictionary (Word-Level) Dataset Summary OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological histories, and semantic relationships in a unified resource. This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression). Key Statistics 150,101 lexemes across 150,101 English… See the full description on the dataset page: https://huggingface.co/datasets/caioloures/personal_dictionary.tabulartext-generation100K<n<1M1 likes52 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.