CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01idleengine /financial-news-multisource Multi-Source Financial & General News 🚀 57.1 MILLION ROWS OF NEWS CONTENT — one unified corpus for market-aware AI/ML I combined 24 public news datasets (many small on their own) into one consistent, ready-to-use layer so you don’t have to wrangle them yourself. Everything is normalized to a minimal schema (date, text, extra_fields) and shipped as Parquet shards per subset—streamable, DuckDB-friendly, and built with a trading date policy (this can be edited if folks see other… See the full description on the dataset page: https://huggingface.co/datasets/idleengine/financial-news-multisource.texttext-classification10M<n<100M0 likes2.7k downloads1mo agoHugging Face02Brianferrell787 /financial-news-multisourcegated Multi-Source Financial & General News 🚀 57.1 MILLION ROWS OF NEWS CONTENT — one unified corpus for market-aware AI/ML I combined 24 public news datasets (many small on their own) into one consistent, ready-to-use layer so you don’t have to wrangle them yourself. Everything is normalized to a minimal schema (date, text, extra_fields) and shipped as Parquet shards per subset—streamable, DuckDB-friendly, and built with a trading date policy (this can be edited if folks see other use… See the full description on the dataset page: https://huggingface.co/datasets/Brianferrell787/financial-news-multisource.texttext-classification10M<n<100M102 likes947 downloads10mo agoHugging Face03tobisns /multisource-eeg-motorimagery0 likes525 downloads1y agoHugging Face04DAMO-NLP-SG /Multi-Source-Video-Captioning Multi-source Video Captioning (MSVC) Dataset Card Dataset details Dataset type: MSVC is a set of collected video captioning data. It is constructed to ensure a robust and thorough evaluation of Video-LLMs' video-captioning capabilities. Dataset detail: MSVC is introduced to address limitations in existing video caption benchmarks, MSVC samples a total of 1,500 videos with human-annotated captions from MSVD, MSRVTT, and VATEX, ensuring diverse scenarios and domains.… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/Multi-Source-Video-Captioning.textvisual-question-answering1K<n<10K7 likes323 downloads2y agoHugging Face05ytc1997 /multisource-membench Multi-Source Memory Benchmark Status — public release. A diagnostic testbed for selective question-answering (ANSWER / SKIP) over conflicting multi-source personal memory. Each persona has five evidence streams projected from a single latent event table with known, controlled per-source distortions (bias direction, dropout rate, granularity), allowing methods to be measured against the latent ground truth rather than against any single source. The benchmark accompanies the… See the full description on the dataset page: https://huggingface.co/datasets/ytc1997/multisource-membench.textquestion-answeringn<1K0 likes180 downloads4mo agoHugging Face06anon-neuripsed26 /multisource-memory-benchmark Multi-Source Memory Benchmark Status — anonymous artefact for double-blind review (NeurIPS 2026 Evaluations & Datasets Track). Author identities, organisations, and funders are intentionally withheld until the review period concludes. A diagnostic testbed for selective question-answering (ANSWER / SKIP) over conflicting multi-source personal memory. Each persona has five evidence streams projected from a single latent event table with known, controlled per-source distortions… See the full description on the dataset page: https://huggingface.co/datasets/anon-neuripsed26/multisource-memory-benchmark.textquestion-answeringn<1K0 likes100 downloads2mo agoHugging Face07manueldeprada /multisource_texttext10M<n<100M0 likes86 downloads2y agoHugging Face08MichaelAnthony /lemonseed-multisource-cogen lemonseed-multisource-cogen LemonSeed — multi-source teacher-co-gen training. Contents intelligent_multisource_train.jsonl (464 rows) Format JSON Lines (.jsonl), one example per line. Provenance LLM-teacher co-generated instruction/chat data for LemonSeed fine-tuning. tabulartext-generationn<1K0 likes47 downloads27d agoHugging Face09Farmaanaa /iran_cpi_and_inflation_multisource شاخص بهای مصرف‌کننده و تورم ایران (چندمرجعی) سری‌های رسمیِ شاخص قیمت مصرف‌کننده (CPI) و نرخ تورم ایران از دو مرجعِ رسمی: مرکز آمار ایران (ماهانه) و بانک مرکزی (سالانه، سریِ تاریخیِ بلند). هر مرجع سریِ مستقلِ خود را دارد و در کنارشان یک نمای یکپارچه ارائه می‌شود. پوشش: بانک مرکزی سالانه ۱۳۱۵–۱۴۰۱ · مرکز آمار ماهانه ۱۳۹۰–۱۴۰۵ سطح: ملی · واحد: شاخص و درصد شاخص‌ها شاخص مرجع تناوب شاخص قیمت مصرف‌کننده مرکز آمار (پایهٔ ۱۴۰۰) · بانک مرکزی (پایهٔ ۱۳۹۵)… See the full description on the dataset page: https://huggingface.co/datasets/Farmaanaa/iran_cpi_and_inflation_multisource.tabular1K<n<10K0 likes45 downloads2mo agoHugging Face10KhalidAlharbi377 /pii-detection-multisource-en-saudi-arabic PII Detection Multisource EN + Saudi/Arabic 284,619 English examples. 2,088,335 labelled spans. 31 entity types. One label space. Four public PII datasets, merged into a single schema, plus Saudi and Arabic coverage that none of them had, plus material for two failure modes that matter when you run redaction in production. Built for OnKith, a privacy first voice assistant that transcribes speech and strips personal information on the device itself, before anything is allowed to… See the full description on the dataset page: https://huggingface.co/datasets/KhalidAlharbi377/pii-detection-multisource-en-saudi-arabic.texttoken-classification100K<n<1M0 likes41 downloads1mo agoHugging Face11MichaelAnthony /lemonseed-multisource-cogen-corrective lemonseed-multisource-cogen-corrective LemonSeed — multi-source corrective teacher-co-gen training (v3). Contents intelligent_multisource_corrective_train_v3.jsonl (522 rows) Format JSON Lines (.jsonl), one example per line. Provenance LLM-teacher co-generated instruction/chat data for LemonSeed fine-tuning. tabulartext-generationn<1K0 likes41 downloads27d agoHugging Face12MichaelAnthony /lemonseed-multisource-cogen-diversified lemonseed-multisource-cogen-diversified LemonSeed — multi-source diversified teacher-co-gen training (v2). Contents intelligent_multisource_diversified_train_v2.jsonl (464 rows) Format JSON Lines (.jsonl), one example per line. Provenance LLM-teacher co-generated instruction/chat data for LemonSeed fine-tuning. tabulartext-generationn<1K0 likes36 downloads27d agoHugging Face13Boanerges /multisource-esco-set MultiSource-ESCO-Skills: A Unified Dataset for Skill Extraction This dataset aggregates data from multiple sources—course descriptions, CV content, and job descriptions—all linked to ESCO skills. It is designed to help researchers and practitioners develop and fine-tune NLP models (e.g., BERT or SentenceTransformer-based models) for automated skill extraction. Dataset Overview Name: MultiSource-ESCO-Skills Sources: Course Content: Educational course materials CV Content:… See the full description on the dataset page: https://huggingface.co/datasets/Boanerges/multisource-esco-set.texttext-classification1M<n<10M0 likes34 downloads2y agoHugging Face14sourcegraph /code-multi-line-infilling-benchmark Dataset Summary This dataset is used to evaluate Multi-Line fill in the middle code completion capabilities of a system. The dataset is derived from SWE-Bench dataset. Evaluation is performed by stiching the generated middle portion, with the other patch and passing into the SWE Evaluation harness, which runs unit test verification and calculate Pass@1. Data Instances In addition to the fields already calculated by SWE-Bench dataset, this dataset contains five… See the full description on the dataset page: https://huggingface.co/datasets/sourcegraph/code-multi-line-infilling-benchmark.textn<1K4 likes33 downloads2y agoHugging Face15professorsynapse /eh-j-space-layer-contrast-rep2-multisource j-space-layer-contrast-rep2-multisource -- aggregate exhaust Aggregate-only: every file committed under this experiment's analysis-committed/ tree (dose-response tables, direction fits, gate AUROCs, manifests, and any other analysis artifact), copied byte-for-byte. No source question text, aliases, or per-row generation text -- analysis-committed/ never carries those. HF repo: professorsynapse/eh-j-space-layer-contrast-rep2-multisource Provenance Experiment:… See the full description on the dataset page: https://huggingface.co/datasets/professorsynapse/eh-j-space-layer-contrast-rep2-multisource.text-classification0 likes33 downloads21d agoHugging Face16THULab /greatplains_multisource_2000_2024 Great Plains Multisource NDVI–Climate Time Series (TsFile) Apache TsFile version of TaoDerong/GreatPlains-Multisource-2000-2024. Overview 8-day aligned multivariate time series for the southern Great Plains grassland region of the United States (~105°W–95°W, 32°N–40°N) over 2000–2024. The region is drought-sensitive, grassland-dominated, and highly responsive to precipitation anomalies. The dataset integrates multi-source observations: MODIS Terra NDVI (8-day… See the full description on the dataset page: https://huggingface.co/datasets/THULab/greatplains_multisource_2000_2024.timeseriestime-series-forecasting1K<n<10K0 likes32 downloads19d agoHugging Face17FronyAI /ko-multisource-retrieval-dataset 구성 여러 공개 출처를 하나의 스키마로 통합한 한국어 query–passage 페어 데이터셋입니다. stage1, stage2 두 subset. 각각 train 100,000 / valid 10,000행 (총 220,000행, parquet, 151MB) 필드 설명 dataset 원 출처 식별자 (18종) query 질의문 (2~2,020자) passage 대응 문서 (10~2,050자) 출처 AI 허브 — 기계독해 지식검색 일반상식 도서자료 행정문서 뉴스기사 금융법률 숫자연산 표정보 추상요약 이벤트 (AI허브_ 접두사) KLUE — klue-mrc klue-nli klue-sts 기타 — kakao-nli 공공데이터포털-deepqa LGNLP 참고 출처별로 페어 성격이 다르므로 용도에 맞게 필터링을 권장합니다. klue-nli klue-sts… See the full description on the dataset page: https://huggingface.co/datasets/FronyAI/ko-multisource-retrieval-dataset.texttext-retrieval100K<n<1M0 likes31 downloads2mo agoHugging Face18DylanJHJ /crux-mds-multi_news-sourcetext10K<n<100K0 likes29 downloads1y agoHugging Face19TaoDerong /GreatPlains-Multisource-2000-2024 Great Plains 8-day Multisource NDVI–Climate Time Series (2000–2024) 1. 数据集概述 (Dataset Summary) 本数据集以 美国南部大平原草原区 为研究区域,范围约为: 经度:105°W–95°W 纬度:32°N–40°N 该区域为典型干旱敏感区,植被以草原为主,对降水异常和干旱事件高度敏感。 本数据集整合了 2000–2024 年间的多源观测,包含: MODIS Terra NDVI(8 日合成) CHIRPS 日尺度降水 ERA5-Land 日尺度表层土壤含水量、2 m 气温、潜在蒸散 以 NDVI 时间步为主轴构建的 8 日对齐多变量时间序列(统一建模输入) 适合于: 干旱监测和评估 NDVI 与气候驱动因子的时滞/响应分析 时序预测任务:ARIMA、多变量 LSTM、Encoder–Decoder 等 多源气象–遥感数据融合研究 2. 数据来源 (Data Sources)… See the full description on the dataset page: https://huggingface.co/datasets/TaoDerong/GreatPlains-Multisource-2000-2024.geospatialtime-series-forecasting1K<n<10K0 likes29 downloads10mo agoHugging Face20Ashfaq2000 /artificial-intelligence-multi-source-datasettextn<1K1 likes28 downloads3d agoHugging Face21hw-jessica /CMTDD_Cross-domain_Multi-source_Tunnel_Defect_Dataset0 likes27 downloads1mo agoHugging Face22Cool-EdwardH /heterogeneous-multisource-kg-benchmark Heterogeneous Multi-Source Knowledge Graph Benchmark A benchmark for evaluating multi-source knowledge graph construction and cross-source question answering systems. The benchmark comprises 50 questions designed to assess systems' ability to reason across heterogeneous data sources (structured SQL, semi-structured JSON, and unstructured text). Dataset Overview Property Value Questions 50 Question Categories Source Attribution, Entity Integration, Conflict… See the full description on the dataset page: https://huggingface.co/datasets/Cool-EdwardH/heterogeneous-multisource-kg-benchmark.textquestion-answeringn<1K0 likes26 downloads9mo agoHugging Face23manueldeprada /multisource_tok_nobooktext10M<n<100M0 likes25 downloads2y agoHugging Face24supergoose /buzz_sources_033_processed_unified_multi_newstext10K<n<100K0 likes23 downloads2y agoHugging Face25Dream000 /high_quality_data_10k_multisourcedocument1K<n<10K0 likes21 downloads1y agoHugging Face26huaXiaKyrie /deltalora-memory-multisource-seed42 Delta-LoRA Memory Multisource Seed42 This repository contains a mixed memory-training corpus built for Delta-LoRA long-context and online-memory experiments. Files: deltalora_memory_multisource_large_seed42.jsonl: full mixed corpus deltalora_memory_multisource_large_seed42.jsonl.summary.json: summary for the full corpus deltalora_memory_multisource_6k_seed42.jsonl: stratified 6k subset for quick training deltalora_memory_multisource_6k_seed42.jsonl.summary.json: summary for the 6k… See the full description on the dataset page: https://huggingface.co/datasets/huaXiaKyrie/deltalora-memory-multisource-seed42.100K<n<1M0 likes21 downloads6mo agoHugging Face27yeganehmohammadi98 /persian-multi-source-corpus license: mit --- 📚 Persian Multi-Source Dataset Collection 🌟 Dataset Overview A comprehensive Persian language dataset combining diverse sources to create a rich training corpus for language models. 📊 Data Sources 🌐 Public Datasets MaralGPT Collections Persian blogs Persian quotes Story Collections TinyStories-Farsi FarsiTinyStories News & Media Farsi news corpus Persian daily news Persian blog posts Conversation & QA Persian QA pairs Telegram channel content Specialized Content… See the full description on the dataset page: https://huggingface.co/datasets/yeganehmohammadi98/persian-multi-source-corpus.text10M<n<100M4 likes16 downloads2y agoHugging Face28supergoose /flan_source_task611_mutual_multi_turn_dialogue_291text10K<n<100K0 likes15 downloads2y agoHugging Face29pankajrajdeo /biomedical-multi-source-finetunetext1K<n<10K0 likes13 downloads1y agoHugging Face30CaoHaiNam /summarization_wikilingua_multisource_chatgpt_segmentationtext10K<n<100K0 likes11 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.