Unsupervised Parsing by Searching for Frequent Word Sequences among Sentences with Equivalent Predicate-Argument Structures



Chen, Junjie, He, Xiangheng, Bollegala, Danushka ORCID: 0000-0003-4476-7003 and Miyao, Yusuke
(2024) Unsupervised Parsing by Searching for Frequent Word Sequences among Sentences with Equivalent Predicate-Argument Structures In: Findings of the Association for Computational Linguistics ACL 2024, 2024-8 - 2024-8.

[thumbnail of Improving_Unsupervised_Constituency_Parsing_via_Maximizing_Semantic_Information.pdf] Text
Improving_Unsupervised_Constituency_Parsing_via_Maximizing_Semantic_Information.pdf - Author Accepted Manuscript
Available under License Creative Commons Attribution.

Download (4MB) | Preview

Abstract

Unsupervised constituency parsing focuses on identifying word sequences that form a syntactic unit (i.e., constituents) in target sentences. Linguists identify the constituent by evaluating a set of Predicate-Argument Structure (PAS) equivalent sentences where we find the constituent appears more frequently than non-constituents (i.e., the constituent corresponds to a frequent word sequence within the sentence set). However, such frequency information is unavailable in previous parsing methods that identify the constituent by observing sentences with diverse PAS. In this study, we empirically show that constituents correspond to frequent word sequences in the PAS-equivalent sentence set. We propose a frequency-based parser span-overlap that (1) computes the span-overlap score as the word sequence's frequency in the PAS-equivalent sentence set and (2) identifies the constituent structure by finding a constituent tree with the maximum span-overlap score. The parser achieves state-of-the-art level parsing accuracy, outperforming existing unsupervised parsers in eight out of ten languages. Additionally, we discover a multilingual phenomenon: participant-denoting constituents tend to have higher span-overlap scores than equal-length event-denoting constituents, meaning that the former tend to appear more frequently in the PAS-equivalent sentence set than the latter. The phenomenon indicates a statistical difference between the two constituent types, laying the foundation for future labeled unsupervised parsing research.

Item Type: Conference Item (Unspecified)
Uncontrolled Keywords: 47 Language, Communication and Culture, 4704 Linguistics
Divisions: Faculty of Science & Engineering
Faculty of Science & Engineering > School of Electrical Engineering, Electronics and Computer Science
Depositing User: Symplectic Admin
Date Deposited: 19 May 2025 08:02
Last Modified: 23 May 2026 09:17
DOI: 10.18653/v1/2024.findings-acl.225
Related Websites:
URI: https://livrepository.liverpool.ac.uk/id/eprint/3187245
Disclaimer: The University of Liverpool is not responsible for content contained on other websites from links within repository metadata. Please contact us if you notice anything that appears incorrect or inappropriate.