datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TextToText_copamisterije-stjenovitog-otoka-hdTextToText_boolqTextToText_record_seqiojurisprudencias_stjFonte: https://scon.stj.jus.br/SCON/
Obs.: Teses geradas por LLM
STJ-SumBR
STJ-SumBR
Dataset de sumarização jurídica construído a partir de acórdãos públicos do Superior Tribunal de Justiça do Brasil (STJ). Cada exemplo alinha o inteiro teor de um acórdão com a ementa oficial produzida pela Secretaria de Jurisprudência do STJ.
Resumo
O STJ disponibiliza publicamente dois conjuntos de dados via STJ Dados Abertos: os espelhos de acórdãos (metadados estruturados e ementa) e as íntegras (texto completo em PDF convertido para TXT). Este… See the full description on the dataset page: https://huggingface.co/datasets/walmeidadf/STJ-SumBR.portuguese-legal-sentences-v0
Work developed as part of Project IRIS.
Thesis: A Semantic Search System for Supremo Tribunal de Justiça
Portuguese Legal Sentences
Collection of Legal Sentences from the Portuguese Supreme Court of Justice
The goal of this dataset was to be used for MLM and TSDAE
Contributions
@rufimelo99
If you use this work, please cite:
@InProceedings{MeloSemantic,
author="Melo, Rui
and Santos, Pedro A.
and Dias, Jo{\~a}o",
editor="Moniz, Nuno
and Vale, Zita
and… See the full description on the dataset page: https://huggingface.co/datasets/stjiris/portuguese-legal-sentences-v0.TextToText_wsc_seqioDescriptors_STJ
Work developed as part of [IRIS] (https://www.inesc-id.pt/projects/PR07005/)
Extreme Multi-Label Classification of Descriptors
The goal of this dataset is to train an Extreme Multi-Label classifier that, given a judgment from the Supreme Court of Justice of Portugal (STJ), can associate relevant descriptors to the judgment.
Dataset Contents:
Judgment ID: Unique identifier for each judgment.
STJ Section: The section of the STJ to which the judgment belongs.Judgment Text: Full… See the full description on the dataset page: https://huggingface.co/datasets/MartimZanatti/Descriptors_STJ.Segmentation_judgments_STJ
Segmentation Dataset for Judgments of the Supreme Court of Justice of Portugal
The goal of this dataset is to train a segmentation model that, given a judgment from the Supreme Court of Justice of Portugal (STJ), can divide its paragraphs into sections of the judgment itself.
Dataset Contents
JSON Files:
Judgment Text:
Contains the judgment text divided into paragraphs, with each paragraph associated with a unique ID.
Denotations:
A list of dictionaries where each… See the full description on the dataset page: https://huggingface.co/datasets/MartimZanatti/Segmentation_judgments_STJ.TextToText_wic_seqioTextToText_squad_seqiosquad_v010_allanswers in T5 paper https://github.com/google-research/text-to-text-transfer-transformer/blob/main/t5/data/tasks.py
DatasetDict({
squad: DatasetDict({
train: Dataset({
features: ['idx', 'inputs', 'targets'],
num_rows: 87599
})
validation: Dataset({
features: ['idx', 'inputs', 'targets'],
num_rows: 10570
})
})
})
qwen3-8b_ultrafeedback_pairrmIRIS_sts
Work developed as part of Project IRIS.
Thesis: A Semantic Search System for Supremo Tribunal de Justiça
Portuguese Legal Sentences
Collection of Legal Sentences pairs from the Portuguese Supreme Court of Justice
The goal of this dataset was to be used for Semantic Textual Similarity
Values from 0-1: random sentences across documents
Values from 2-4: sentences from the same summary (implying some level of entailment)
Values from 4-5: sentences pairs generated through OpenAi'… See the full description on the dataset page: https://huggingface.co/datasets/stjiris/IRIS_sts.TextToText_copa_seqioTextToText_mnli_seqioTextToText_rte_seqioTextToText_DocNLI_seqiotext to text implementation basing on https://github.com/salesforce/DocNLI
DatasetDict({
train: Dataset({
features: ['idx', 'inputs', 'targets'],
num_rows: 942314
})
validation: Dataset({
features: ['idx', 'inputs', 'targets'],
num_rows: 234258
})
test: Dataset({
features: ['idx', 'inputs', 'targets'],
num_rows: 267086
})
})
CT-RATE-Dataset-cleanedsumulas_stjqwen3-4b-no-think_ultrafeedback_pairrmquestion_to_sql_with_ddl
Dataset Card for "question_to_sql_with_ddl"
More Information needed
deepseek-llm-7b-chat_ultrafeedback_pairrmjurisprudencia_stj_pt
Dataset de Jurisprudência do STJ de Portugal (2025)
Descrição do Dataset
Este dataset contém uma amostra de acórdãos proferidos pelo Supremo Tribunal de Justiça (STJ) de Portugal durante o ano de 2025. Cada registo no dataset corresponde a um acórdão completo, incluindo o seu texto integral, o sumário e um conjunto de metadados ricos.
Os dados representam uma amostra aleatória de 5% do total de acórdãos de 2025 disponíveis na base de dados de origem, filtrados para… See the full description on the dataset page: https://huggingface.co/datasets/ffantini/jurisprudencia_stj_pt.TextToText_multirc_seqioTextToText_boolq_seqioTextToText_axg_seqio
text-to-text format from superglue axg
Note that RTE train and val set has been added
axg: DatasetDict({
test: Dataset({
features: ['idx', 'inputs', 'targets'],
num_rows: 356
})
train: Dataset({
features: ['idx', 'inputs', 'targets'],
num_rows: 2490
})
validation: Dataset({
features: ['idx', 'inputs', 'targets'],
num_rows: 277
})
})
TextToText_rteTextToText_axb_seqioaxb: DatasetDict({
test: Dataset({
features: ['idx', 'inputs', 'targets'],
num_rows: 1104
})
train: Dataset({
features: ['idx', 'inputs', 'targets'],
num_rows: 2490
})
validation: Dataset({
features: ['idx', 'inputs', 'targets'],
num_rows: 277
})
})
Text to text implemantion of T5
note that RTE train and validation set has been added
stj_docket_generation
Dataset: Legal Documents from STJ for Jurimetrics Research
Dataset Overview
This dataset contains legal documents from the Superior Tribunal de Justiça (STJ), designed for research in jurimetrics, automatic text summarization, and retrieval-augmented generation (RAG). The dataset focuses on the challenges posed by hierarchical structures, legal vocabulary, ambiguity, and citations in legal texts.
Contents
The dataset includes:
Ementas (Summaries): Concise… See the full description on the dataset page: https://huggingface.co/datasets/bfunicheli/stj_docket_generation.
