CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01vectara /hhem_leaderboard_datasetstext1K<n<10K0 likes227 downloads1y agoHugging Face02binhtran23 /vecombot-dataset VECOM — Bộ dữ liệu thị trường Thương mại điện tử Việt Nam (VEComBot) Bộ dữ liệu phụ lục cho đồ án tốt nghiệp VEComBot — hệ thống Đa tác tử (Multi-Agent System) phân tích và tổng hợp thị trường Thương mại điện tử Việt Nam (VECOM). Đây là kho tài liệu nguồn và corpus đã qua xử lý (figure-aware) được nạp vào PostgreSQL/pgvector để phục vụ cả nhánh MAS lẫn nhánh baseline naive RAG. Mục đích: dùng cho nghiên cứu học thuật và tái lập kết quả đồ án. Các báo cáo gốc là ấn phẩm công… See the full description on the dataset page: https://huggingface.co/datasets/binhtran23/vecombot-dataset.imagequestion-answeringn<1K0 likes162 downloads3mo agoHugging Face03wudq /VECBenchimage100K<n<1M0 likes136 downloads9mo agoHugging Face04vectorJoseph /frequent-stock-fccb17 frequent-stock-fccb17 Synthetic sensors test data: 39 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/vectorJoseph/frequent-stock-fccb17.tabularn<1K0 likes57 downloads13d agoHugging Face05ScoutieAutoML /scoutieDataset_russian_language_grammar_and_rules_vectorized Description in English: A dataset collected from 30 Russian-language Telegram channels on the topic of learning the Russian language. This dataset contains grammar, syntax, spelling and punctuation rules. The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link. Dataset fields: taskId - task identifier in the Scouti service. text - main text. url -… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_russian_language_grammar_and_rules_vectorized.tabulartext-classification10K<n<100K2 likes48 downloads2y agoHugging Face06vectorDawn /international-delivery-68f942 international-delivery-68f942 Synthetic sensors test data: 34 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting… See the full description on the dataset page: https://huggingface.co/datasets/vectorDawn/international-delivery-68f942.tabularn<1K0 likes46 downloads13d agoHugging Face07vectorRemy /dirty-occasion-d21fee dirty-occasion-d21fee Synthetic sensors test data: 34 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/vectorRemy/dirty-occasion-d21fee.tabularn<1K0 likes43 downloads13d agoHugging Face08vectara /leaderboard_resultstext100K<n<1M5 likes38 downloads1y agoHugging Face09vectorridge /significant-table-7d9bce significant-table-7d9bce Synthetic sensors test data: 37 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/vectorridge/significant-table-7d9bce.tabularn<1K0 likes37 downloads13d agoHugging Face10ScoutieAutoML /it_vacancies_vectorisation Description in English: The dataset is collected from Russian-language Telegram channels with information about various IT vacancies for ML, Data Science, Front-end, Back-end development,The dataset was collected and tagged automatically using the data collection and tagging service Scoutie.Try Scoutie and collect the same or another dataset using link for FREE. Dataset fields: taskId - task identifier in the Scouti service. text - main text. url - link to the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/it_vacancies_vectorisation.tabulartext-classification10K<n<100K0 likes31 downloads2y agoHugging Face11ClarusC64 /clinical-nbdm-intervention-vector-selection-v0.1What this dataset tests Given discordance and polaritychoose an intervention vector that should restore coherence. Vectors inside_out outside_in hybrid monitor Inside-out means narrative-first. Outside-in means biology-first. Hybrid means both in parallel. Monitor means low risk and unclear locus. Typical errors recommending therapy first when biomarkers imply organ risk recommending only meds when narrative drives instability forcing action when monitoring is safer… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-nbdm-intervention-vector-selection-v0.1.texttext-classificationn<1K0 likes31 downloads8mo agoHugging Face12Vector-Xueyong /unhappy-discipline-7f43f3 unhappy-discipline-7f43f3 Synthetic weather test data: 58 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Vector-Xueyong/unhappy-discipline-7f43f3.tabularn<1K0 likes29 downloads13d agoHugging Face13ScoutieAutoML /recipes_for_dishes_and_food_with_vectors_sentiment_ners Description in English: The dataset is collected from Russian-language Telegram channels with various food recipes,The dataset was collected and tagged automatically using the data collection and tagging service Scoutie.Try Scoutie and collect the same or another dataset using link for FREE. Dataset fields: taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/recipes_for_dishes_and_food_with_vectors_sentiment_ners.tabulartext-classification10K<n<100K2 likes28 downloads2y agoHugging Face14ScoutieAutoML /russian_jokes_with_vectors Description in English: The dataset is collected from Russian-language Telegram channels with jokes and anecdotes,The dataset was collected and tagged automatically using the data collection and tagging service Scoutie.Try Scoutie and collect the same or another dataset using link for FREE. Dataset fields: taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/russian_jokes_with_vectors.tabulartext-classification10K<n<100K5 likes25 downloads2y agoHugging Face15ClarusC64 /clinical-personal-deviation-vector-detection-v0.1What this dataset tests Whether a model can detect deviation from a person's own coherent basinusing baseline envelope and coupling structure. Required outputs deviation_vector deviation_severity_score_0_100 first_system_departing Deviation vector fields direction magnitude velocity coupling_loss onset_time cross_modal_consensus First system labels sleep_circadian autonomic immune_inflammatory metabolic neurocognitive gut_microbiome behavior_load… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-personal-deviation-vector-detection-v0.1.texttext-classificationn<1K0 likes25 downloads8mo agoHugging Face16letscom /Vector_Database_With_Open-Sourcetabularn<1K0 likes24 downloads3y agoHugging Face17vectara /hcm-examples-aug-2024Dataset of some examples with hallucinations before and after passing through Vectara's Hallucination Correction Model. See our blogpost for details. tabularn<1K1 likes22 downloads2y agoHugging Face18kmanticore /skills_network_vector_databasetabularn<1K0 likes21 downloads3y agoHugging Face19ScoutieAutoML /russian_events_vectors Description in English: The dataset is collected from Russian-language Telegram channels with information about various events and happenings in Russian regions,The dataset was collected and tagged automatically using the data collection and tagging service Scoutie.Try Scoutie and collect the same or another dataset using link for FREE. Dataset fields: taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink -… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/russian_events_vectors.tabulartext-classification10K<n<100K0 likes21 downloads2y agoHugging Face20ScoutieAutoML /scoutieDataset_chinese_russian_dictionary_grammar_spelling_vectorized Description in English: A dataset collected from 30 Russian-language Telegram channels on the topic of learning Chinese, this dataset contains grammar, syntax, spelling and punctuation rules, as well as Chinese words with Russian translations. The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link. Dataset fields: taskId - task identifier in the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_chinese_russian_dictionary_grammar_spelling_vectorized.tabulartext-classification1K<n<10K0 likes16 downloads2y agoHugging Face21KnutJaegersberg /direction_vectors_ftq_enThese direction vectors of antonyms can be used to calculate fasttext interpretable embeddings on the fly, solving the OOV problem of other interpretable embeddings. Simply calculate cosine similarity for each row. tabular10K<n<100K0 likes12 downloads4y agoHugging Face22sudiptabasak /alpaca-vectorstabularn<1K0 likes10 downloads3y agoHugging Face23ScoutieAutoML /scoutieDataset_english_russian_dictionary_grammar_spelling_vectorized Description in English: A dataset collected from 30 Russian-language Telegram channels on the topic of learning English, this dataset contains grammar, syntax, spelling and punctuation rules, as well as English words with Russian translations. The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link. Dataset fields: taskId - task identifier in the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_english_russian_dictionary_grammar_spelling_vectorized.tabulartext-classification10K<n<100K0 likes10 downloads2y agoHugging Face24ScoutieAutoML /scoutieDataset_chemical_terms_with_definition_vectorized Description in English: Dataset collected from 30 Russian-language Telegram channels on the topic of Chemistry, collected and marked up automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link. Dataset fields: taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the channel. views - text… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_chemical_terms_with_definition_vectorized.tabulartext-classification1K<n<10K0 likes9 downloads2y agoHugging Face25sudiptabasak /expressions-vectorstabularn<1K0 likes8 downloads3y agoHugging Face26nataliecastro /climate-bills-lemmed-count-vectorizertabular1K<n<10K0 likes7 downloads1y agoHugging Face27vmadhav /vectordbdemotabularn<1K0 likes6 downloads3y agoHugging Face28vector-institute /MultiHopRAG-syn-data-ctx_len-4096-100textn<1K0 likes6 downloads2y agoHugging Face29Gurashish1994 /vector_agentic_team3_2tabular1M<n<10M0 likes6 downloads1y agoHugging Face30Gurashish1994 /vector_agentic_team3_3tabular10K<n<100K0 likes6 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.