AmazonScience/document-haystack
Document Haystack Dataset This repository contains the dataset for the paper “Document Haystack: A Long Context Multimodal Image/Document Understanding Vision LLM Benchmark”. 📑 Abstract Paper The proliferation of multimodal Large Language Models has significantly advanced the ability to analyze and understand complex data inputs from different modalities. However, the processing of long documents remains under-explored, largely due to a lack of suitable… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/document-haystack.
2090k
1We face risks resulting from the extensive use of models, AI, and data.2We rely on quantitative models and the use of AI, as well as our ability to manage and aggregate data in an accurate and timely 3manner, to assess and manage our various risk exposures, create estimates and forecasts, and manage compliance with 4regulatory capital requirements. We continue to invest in building new capabilities that employ new AI technologies such as 5generative AI, and we expect our use of these technologies to increase over time. However, there are significant risks involved 6in utilizing models and AI and no assurance can be provided that our use will produce only intended or beneficial results. AI 7may subject us to new or heightened legal, regulatory, ethical, or other challenges; and negative public opinion of AI could 8impair the acceptance of AI solutions. If the models or AI solutions that we create or use are deficient, inaccurate or 9controversial, we could incur operational inefficiencies, competitive harm, legal liability, brand or reputational harm, or other 10adverse impacts on our business and financial results. We also may incur liability through the violation of applicable laws and 11regulations, third-party intellectual property, privacy or other rights, or contracts to which we are a party.12We may use models and AI in processes such as determining the pricing of various products, identifying potentially fraudulent 13transactions, grading loans and extending credit, measuring interest rate and other market risks, predicting deposit levels or loan 14losses, assessing capital adequacy, calculating managerial and regulatory capital levels, estimating the value of financial 15instruments and balance sheet items, and other operational functions. Development and implementation of some of these 16models , such as the models for credit loss accounting under CECL, require us to make difficult, subjective and complex 17judgments. Our risk reporting and management, including business decisions based on information incorporating models and 18the use of AI, depend on the effectiveness of our models and AI and our policies, programs, processes and practices governing 19how data, models and AI, as applicable, are acquired, validated, stored, protected, processed and analyzed. Any issues with the 20quality or effectiveness of our data aggregation and validation procedures, as well as the quality and integrity of data inputs, 21formulas or algorithms, could result in inaccurate forecasts, ineffective risk management practices or inaccurate risk reporting. 22In addition, models and AI based on historical data sets might not be accurate predictors of future outcomes and their ability to 23appropriately predict future outcomes may degrade over time due to limited historical patterns, extreme or unanticipated market 24movements or customer behavior and liquidity, especially during severe market downturns or stress events (e.g., geopolitical or 25pandemic events).26While we continuously update our policies, programs, processes and practices, many of our data management, modeling, AI, 27aggregation and implementation processes are manual and may be subject to human error, data limitations, process delays or 28system failure. Failure to manage data effectively and to aggregate data in an accurate and timely manner may limit our ability 29to manage current and emerging risk, to produce accurate financial, regulatory and operational reporting as well as to manage 30changing business needs. If our Framework is ineffective, we could suffer unexpected losses which could materially adversely 31affect our results of operation or financial condition. Also, any information we provide to the public or to our regulators based 32on incorrectly designed or implemented models or AI could be inaccurate or misleading. Some of the decisions that our 33regulators make, including those related to capital distribution to our stockholders, could be affected adversely due to the 34perception that the quality of the data, models and AI used to generate the relevant information is insufficient. In addition, 35regulation of AI is rapidly evolving worldwide as legislators and regulators are increasingly focused on these powerful 36emerging technologies. The technologies underlying AI and its uses are subject to a variety of laws and regulations, including 37intellectual property, privacy, data protection and information security, consumer protection, competition, and equal 38opportunity laws, and are expected to be subject to increased regulation and new laws or new applications of existing laws and 39regulations. AI is the subject of ongoing review by various U.S. governmental and regulatory agencies, and various U.S. states 40and other foreign jurisdictions are applying, or are considering applying, their platform moderation, privacy, data protection and 41data security laws and regulations to AI or are considering general legal frameworks for AI. We may not be able to anticipate 42how to respond to these rapidly evolving frameworks, and we may need to expend resources to adjust our offerings in certain 43jurisdictions if the legal frameworks are inconsistent across jurisdictions. Furthermore, because AI technology itself is highly 44complex and rapidly developing, it is not possible to predict all of the legal, operational or technological risks that may arise 45relating to the use of AI.46Legal and Regulatory Risk47Compliance with new and existing domestic and foreign laws, regulations and regulatory expectations is costly and complex.48A wide array of laws and regulations, including banking and consumer lending laws and regulations, apply to every aspect of 49our business and these laws can be uncertain and evolving. We and our subsidiaries are also subject to supervision and 50examination by multiple regulators both in the U.S. and abroad, and the manner in which our regulators interpret applicable 51laws and regulations may affect how we comply with them. Failure to comply with these laws and regulations, even if the 52failure was inadvertent or reflects a difference in interpretation or conflicting legal requirements, could subject us to restrictions 5333 Capital One Financial Corporation (COF)