datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mckinsey_state_of_ai_doc_understanding
Mckinsey State Of Ai Doc Understanding
This dataset was generated using YourBench (v0.3.1), an open-source framework for generating domain-specific benchmarks from document collections.
Pipeline Steps
ingestion: Read raw source documents, convert them to normalized markdown and save for downstream steps
summarization: Perform hierarchical summarization: chunk-level LLM summaries followed by combine-stage reduction
chunking: Split texts into token-based single-hop and… See the full description on the dataset page: https://huggingface.co/datasets/yourbench/mckinsey_state_of_ai_doc_understanding.FedE4OODRAG4FINhf_doc_qa_eval_best_answers
