Law
Datasheet Y ubo Ma
Q1: F or what purpose was the dataset created? As stated in Section 1, most previous datasets on DU focus on single-page DU. Our benchmark is constructed to bridge such a gap. Q2: Who created the dataset (e.g., which team, research group) and on behalf of which entity (e.g., Q3: What support was needed to make this dataset? Q1: What do the instances that comprise the dataset represent (e.g., documents, photos, people, countries)?
A Additional Results
The acronym dataset is a QA task that requires models to decode financial acronyms. The FinMA7B-full model achieved the highest ROUGE-1 score of 0.12 and the B.1 Why was the datasheet created? B.2 Has the dataset been used already? If so, where are the results so others can compare (e.g., links to published papers)? Y es, the dataset has already been used. It was employed in the FinLLM Share Task during the FinNLP-AgentScen Workshop at IJCAI 2024, known as the FinLLM Challenge.
A Appendix
The complete list may be seen in Table 8. Here are a few general notes about these strings: 1. Based on their recommendations, we did the following: 1. zh, zh_Latn: This resulted in the special filters described below. URLs) the corpora were in languages different from the LangID predictions. This is mainly mis-rendered PDFs and may have practical applications for denoising, or for decoding such garbled PDFs.