datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
camara-proposicoes-clustering
CamaraProposicoesClustering
Cluster the summaries (ementas) of bills from the Brazilian Chamber of Deputies into legislative themes from the Chamber's official taxonomy (Economia, Educação, Saúde, Meio Ambiente, Direitos Humanos, Administração Pública, etc.). Native PT-BR legislative text; public-domain government open data.
Part of MTEB-BR — the native Brazilian-Portuguese MTEB sub-benchmark. Task type: Clustering · Language: Brazilian Portuguese (mined from real-world sources)… See the full description on the dataset page: https://huggingface.co/datasets/MTEB-BR/camara-proposicoes-clustering.MetaQAFreebase-WebQSP-CWQ-Subgraph
Freebase QA-Oriented Subgraphs
This dataset contains reconstructed Freebase subgraphs for Knowledge Graph Question Answering (KGQA) and multi-hop reasoning tasks.
The subgraphs were reconstructed from the serialized graph structures provided in the preprocessing pipeline of the RoG (Reasoning on Graphs) framework:
https://github.com/RManLuo/reasoning-on-graphs
The dataset was reconstructed using:
WebQSP
ComplexWebQuestions (CWQ)
Files
webqsp_subgraph.tsv… See the full description on the dataset page: https://huggingface.co/datasets/camazlucas/Freebase-WebQSP-CWQ-Subgraph.ementas_camarabr_1934_2024Collected at 26 Sept 2024
camara_audio_sentencescam_assesscamarao_dorme_onda_leva1camarao_dorme_onda_leva4camarao_dorme_onda_leva2camarao_dorme_onda_leva3camarao_dorme_onda_leva5cam_assess_phonemescam_assess_sampleguanaco-llama2-200cam_assess
