datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hhem_leaderboard_datasetsvecombot-dataset
VECOM — Bộ dữ liệu thị trường Thương mại điện tử Việt Nam (VEComBot)
Bộ dữ liệu phụ lục cho đồ án tốt nghiệp VEComBot — hệ thống Đa tác tử (Multi-Agent
System) phân tích và tổng hợp thị trường Thương mại điện tử Việt Nam (VECOM). Đây là
kho tài liệu nguồn và corpus đã qua xử lý (figure-aware) được nạp vào PostgreSQL/pgvector
để phục vụ cả nhánh MAS lẫn nhánh baseline naive RAG.
Mục đích: dùng cho nghiên cứu học thuật và tái lập kết quả đồ án. Các báo cáo gốc là
ấn phẩm công… See the full description on the dataset page: https://huggingface.co/datasets/binhtran23/vecombot-dataset.VECBenchfrequent-stock-fccb17
frequent-stock-fccb17
Synthetic sensors test data: 39 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/vectorJoseph/frequent-stock-fccb17.scoutieDataset_russian_language_grammar_and_rules_vectorized
Description in English:
A dataset collected from 30 Russian-language Telegram channels on the topic of learning the Russian language. This dataset contains grammar, syntax, spelling and punctuation rules.
The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url -… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_russian_language_grammar_and_rules_vectorized.international-delivery-68f942
international-delivery-68f942
Synthetic sensors test data: 34 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting… See the full description on the dataset page: https://huggingface.co/datasets/vectorDawn/international-delivery-68f942.dirty-occasion-d21fee
dirty-occasion-d21fee
Synthetic sensors test data: 34 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/vectorRemy/dirty-occasion-d21fee.leaderboard_resultssignificant-table-7d9bce
significant-table-7d9bce
Synthetic sensors test data: 37 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/vectorridge/significant-table-7d9bce.it_vacancies_vectorisation
Description in English:
The dataset is collected from Russian-language Telegram channels with information about various IT vacancies for ML, Data Science, Front-end, Back-end development,The dataset was collected and tagged automatically using the data collection and tagging service Scoutie.Try Scoutie and collect the same or another dataset using link for FREE.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url - link to the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/it_vacancies_vectorisation.clinical-nbdm-intervention-vector-selection-v0.1What this dataset tests
Given discordance and polaritychoose an intervention vector that should restore coherence.
Vectors
inside_out
outside_in
hybrid
monitor
Inside-out means narrative-first.
Outside-in means biology-first.
Hybrid means both in parallel.
Monitor means low risk and unclear locus.
Typical errors
recommending therapy first when biomarkers imply organ risk
recommending only meds when narrative drives instability
forcing action when monitoring is safer… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-nbdm-intervention-vector-selection-v0.1.unhappy-discipline-7f43f3
unhappy-discipline-7f43f3
Synthetic weather test data: 58 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Vector-Xueyong/unhappy-discipline-7f43f3.recipes_for_dishes_and_food_with_vectors_sentiment_ners
Description in English:
The dataset is collected from Russian-language Telegram channels with various food recipes,The dataset was collected and tagged automatically using the data collection and tagging service Scoutie.Try Scoutie and collect the same or another dataset using link for FREE.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/recipes_for_dishes_and_food_with_vectors_sentiment_ners.russian_jokes_with_vectors
Description in English:
The dataset is collected from Russian-language Telegram channels with jokes and anecdotes,The dataset was collected and tagged automatically using the data collection and tagging service Scoutie.Try Scoutie and collect the same or another dataset using link for FREE.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/russian_jokes_with_vectors.clinical-personal-deviation-vector-detection-v0.1What this dataset tests
Whether a model can detect deviation from a person's own coherent basinusing baseline envelope and coupling structure.
Required outputs
deviation_vector
deviation_severity_score_0_100
first_system_departing
Deviation vector fields
direction
magnitude
velocity
coupling_loss
onset_time
cross_modal_consensus
First system labels
sleep_circadian
autonomic
immune_inflammatory
metabolic
neurocognitive
gut_microbiome
behavior_load… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-personal-deviation-vector-detection-v0.1.Vector_Database_With_Open-Sourcehcm-examples-aug-2024Dataset of some examples with hallucinations before and after passing through Vectara's Hallucination Correction Model. See our blogpost for details.
skills_network_vector_databaserussian_events_vectors
Description in English:
The dataset is collected from Russian-language Telegram channels with information about various events and happenings in Russian regions,The dataset was collected and tagged automatically using the data collection and tagging service Scoutie.Try Scoutie and collect the same or another dataset using link for FREE.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink -… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/russian_events_vectors.scoutieDataset_chinese_russian_dictionary_grammar_spelling_vectorized
Description in English:
A dataset collected from 30 Russian-language Telegram channels on the topic of learning Chinese, this dataset contains grammar, syntax, spelling and punctuation rules, as well as Chinese words with Russian translations.
The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_chinese_russian_dictionary_grammar_spelling_vectorized.direction_vectors_ftq_enThese direction vectors of antonyms can be used to calculate fasttext interpretable embeddings on the fly, solving the OOV problem of other interpretable embeddings.
Simply calculate cosine similarity for each row.
alpaca-vectorsscoutieDataset_english_russian_dictionary_grammar_spelling_vectorized
Description in English:
A dataset collected from 30 Russian-language Telegram channels on the topic of learning English, this dataset contains grammar, syntax, spelling and punctuation rules, as well as English words with Russian translations.
The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_english_russian_dictionary_grammar_spelling_vectorized.scoutieDataset_chemical_terms_with_definition_vectorized
Description in English:
Dataset collected from 30 Russian-language Telegram channels on the topic of Chemistry,
collected and marked up automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the channel. views - text… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_chemical_terms_with_definition_vectorized.expressions-vectorsclimate-bills-lemmed-count-vectorizervectordbdemoMultiHopRAG-syn-data-ctx_len-4096-100vector_agentic_team3_2vector_agentic_team3_3
