datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
time_expressions_dataset
Dataset Card for Time Expressions Dataset
Dataset Summary
The Time Expressions Dataset is a collection of synthetic data designed for training and evaluating natural language processing (NLP) models on temporal expression recognition and resolution tasks. It contains 378 unique data points, each consisting of a natural language sentence (input_text) and a corresponding JSON-structured output (target_output) that resolves a specific time expression to a standardized date… See the full description on the dataset page: https://huggingface.co/datasets/namesarnav/time_expressions_dataset.time-entries-and-phases
Time Entry Dataset at a Glance
31 litigation matters
≈13k unique time entries
≈20k hours of billed time
4 phases labeled: Pleading, Discovery, Pretrial, Trial
Law firm invoices were OCRed with tersseract 3.0, LLMs extracted time entries, with manual data cleaning.
Source PDF documents available on request.
Sample: New York Commercial Contract Case
Time entries for an example matter are shown for a New York commercial contract dispute:
Title
Cowen and Company… See the full description on the dataset page: https://huggingface.co/datasets/LexPipe/time-entries-and-phases.egowalk_sampletime_esi_1
