timee
Datasets
All datasets matching “timee”time-embed-korean-temporal-inventory-v2
Time-Embed Korean Temporal Inventory v2
한국어 시간 표현 임베딩 학습·평가 데이터셋입니다.
C1, C3, C5: 문맥 허용 범위가 다른 학습 조건이며 각 314개 query를 포함합니다.
각 학습 query는 승인된 모든 동의 positive(최소 3개)와 정확히 7개 negative를 가집니다.
legacy_validation, legacy_test: 기존 동결 dev/test bytes를 그대로 보존합니다.
inventory_eval: 승인된 희귀·격식·경계 시간 표현 118개입니다.
Inventory Test 원문과 라벨은 공개하지 않으며 sealed/inventory_test_handle.json만 제공합니다.
승인 방식은 개별 행 검수로 위장하지 않은 owner_policy_waiver입니다. 정확한 승인·감사
해시와 산출물은 evidence/, audit/, dataset_manifest.json에… See the full description on the dataset page: https://huggingface.co/datasets/kev-KOH/time-embed-korean-temporal-inventory-v2.time_expressions_dataset
Dataset Card for Time Expressions Dataset
Dataset Summary
The Time Expressions Dataset is a collection of synthetic data designed for training and evaluating natural language processing (NLP) models on temporal expression recognition and resolution tasks. It contains 378 unique data points, each consisting of a natural language sentence (input_text) and a corresponding JSON-structured output (target_output) that resolves a specific time expression to a standardized date… See the full description on the dataset page: https://huggingface.co/datasets/namesarnav/time_expressions_dataset.time-entries-and-phases
Time Entry Dataset at a Glance
31 litigation matters
≈13k unique time entries
≈20k hours of billed time
4 phases labeled: Pleading, Discovery, Pretrial, Trial
Law firm invoices were OCRed with tersseract 3.0, LLMs extracted time entries, with manual data cleaning.
Source PDF documents available on request.
Sample: New York Commercial Contract Case
Time entries for an example matter are shown for a New York commercial contract dispute:
Title
Cowen and Company… See the full description on the dataset page: https://huggingface.co/datasets/LexPipe/time-entries-and-phases.egowalk_sampletime-embed-bge-m3
Korean Temporal Query Embedding Data for BGE-M3 - v1.9 Semantic Retention
This dataset is a FlagEmbedding/BGE-M3 fine-tuning dataset for Korean LMS temporal retrieval.
The objective is to keep semantically equivalent Korean temporal expressions close in embedding space while separating Korean expressions that look similar but mean different time ranges. The root files now point to the v1.9 training dataset.
What v1.9 Adds
v1.9 keeps the v1.7 calendar-focused… See the full description on the dataset page: https://huggingface.co/datasets/kev-KOH/time-embed-bge-m3.TimeExtractor
Time Extractor Training Dataset
Author: JioNLP
Link: JioNLP
This dataset is designed for fine-tuning LLMs to extract time entities from the text, which is aimed to get the standard time string in json format.
It is divided into two parts:
general.json: Samples extracted from various news sources.
smartspeaker.json: Samples obtained from voice assistants.
The process involves:
First, extract the original time entity strings, which are then analyzed by a large model to… See the full description on the dataset page: https://huggingface.co/datasets/dongrixinyu/TimeExtractor.
