extraction
wikidata-extraction
Wikidata Extraction
This dataset contains all RDF triples extracted from the latest Wikidata, converted from the N-Triples format to Parquet.
The data originates from Wikidata, a free and open knowledge base that acts as central storage for structured data used by Wikipedia and other Wikimedia projects. The source file is the "truthy" N-Triples dump (latest-truthy.nt.bz2), which contains only the current, non-deprecated statements.
The code to extract this data is available at… See the full description on the dataset page: https://huggingface.co/datasets/piebro/wikidata-extraction.vi-books_extractionll books_text_final/ | wc -l
=> 13968 text
After filter (not vi, too small)
wc -l vi-books.jsonl
=> 12832 text
=> !!! Cần lọc truyện từ dtv ra !!!
xzcat dtv_ebooks_details.jsonl.xz | wc -l
# => 13486
ll books_text_final/ | grep dtv_ebooks_ | wc -l
# => 3535 từ dtv
lấy các whitelist categories tại dtv_categories_whitelist.jsonl
lần theo từng ebook một trong dtv_ebooks_details.jsonl.xz trường {"cat": "Self Help - Khởi nghiệp", "cat_url":… See the full description on the dataset page: https://huggingface.co/datasets/tiendung/vi-books_extraction.tweet_sentiment_extraction
TweetSentimentExtractionClassification
An MTEB dataset
Massive Text Embedding Benchmark
Task category
t2c
Domains
Social, Written
Reference
https://www.kaggle.com/competitions/tweet-sentiment-extraction/overview
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["TweetSentimentExtractionClassification"])
evaluator = mteb.MTEB(task)
model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/tweet_sentiment_extraction.sec-contracts-financial-extraction-instructions
S&P 500 SEC Financial Extraction Instructions
Dataset Summary
7,683 instruction-tuning examples for training LLMs to extract structured financial data from SEC filings. Covers two filing types across S&P 500 companies:
Split
Examples
Filing Type
Description
train
3,430
Exhibit 10 + DEF 14A
Positive examples with validated outputs
corrective
4,253
Exhibit 10 + DEF 14A
Corrective, rescued, and negative examples
Exhibit 10 — Material Contracts (2… See the full description on the dataset page: https://huggingface.co/datasets/TheTokenFactory/sec-contracts-financial-extraction-instructions.tool-output-extraction-swebench
Tool Output Extraction Dataset
Paper | Code
Training data for squeez — a small model that prunes verbose coding agent tool output to only the evidence the agent needs next.
Task
Task-conditioned context pruning of a single tool observation for coding agents.
Given a focused extraction query and one verbose tool output, return the smallest verbatim evidence block(s) the agent should read next.
The model copies lines from the tool output — it never rewrites, summarizes, or… See the full description on the dataset page: https://huggingface.co/datasets/KRLabsOrg/tool-output-extraction-swebench.affirming-review-extraction
