{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,20]],"date-time":"2025-10-20T17:06:07Z","timestamp":1760979967343},"reference-count":29,"publisher":"Wiley","issue":"3","license":[{"start":{"date-parts":[[2004,12,2]],"date-time":"2004-12-02T00:00:00Z","timestamp":1101945600000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/onlinelibrary.wiley.com\/termsAndConditions#vor"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J. Am. Soc. Inf. Sci."],"published-print":{"date-parts":[[2005,2]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>For the sake of national security, very large volumes of data and information are generated and gathered daily. Much of this data and information is written in different languages, stored in different locations, and may be seemingly unconnected. Crosslingual semantic interoperability is a major challenge to generate an overview of this disparate data and information so that it can be analyzed, shared, searched, and summarized. The recent terrorist attacks and the tragic events of September 11, 2001 have prompted increased attention on national security and criminal analysis. Many Asian countries and cities, such as Japan, Taiwan, and Singapore, have been advised that they may become the next targets of terrorist attacks. Semantic interoperability has been a focus in digital library research. Traditional information retrieval (IR) approaches normally require a document to share some common keywords with the query. Generating the associations for the related terms between the two term spaces of users and documents is an important issue. The problem can be viewed as the creation of a thesaurus. Apart from this, terrorists and criminals may communicate through letters, e\u2010mails, and faxes in languages other than English. The translation ambiguity significantly exacerbates the retrieval problem. The problem is expanded to crosslingual semantic interoperability. In this paper, we focus on the English\/Chinese crosslingual semantic interoperability problem. However, the developed techniques are not limited to English and Chinese languages but can be applied to many other languages. English and Chinese are popular languages in the Asian region. Much information about national security or crime is communicated in these languages. An efficient automatically generated thesaurus between these languages is important to crosslingual information retrieval between English and Chinese languages. To facilitate crosslingual information retrieval, a corpus\u2010based approach uses the term co\u2010occurrence statistics in parallel or comparable corpora to construct a statistical translation model to cross the language boundary. In this paper, the text\u2010based approach to align English\/Chinese Hong Kong Police press release documents from the Web is first presented. We also introduce an algorithmic approach to generate a robust knowledge base based on statistical correlation analysis of the semantics (knowledge) embedded in the bilingual press release corpus. The research output consisted of a thesaurus\u2010like, semantic network knowledge base, which can aid in semantics\u2010based crosslingual information management and retrieval.<\/jats:p>","DOI":"10.1002\/asi.20118","type":"journal-article","created":{"date-parts":[[2004,12,2]],"date-time":"2004-12-02T23:47:37Z","timestamp":1102031257000},"page":"272-282","source":"Crossref","is-referenced-by-count":14,"title":["Automatic crosslingual thesaurus generated from the Hong Kong SAR Police Department Web corpus for crime analysis"],"prefix":"10.1002","volume":"56","author":[{"given":"Kar Wing","family":"Li","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Christopher C.","family":"Yang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"311","published-online":{"date-parts":[[2005,1,12]]},"reference":[{"key":"e_1_2_7_2_1","doi-asserted-by":"publisher","DOI":"10.1002\/(SICI)1097-4571(198611)37:6<357::AID-ASI1>3.0.CO;2-H"},{"key":"e_1_2_7_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/21.179830"},{"key":"e_1_2_7_4_1","doi-asserted-by":"publisher","DOI":"10.1002\/(SICI)1097-4571(199701)48:1<17::AID-ASI4>3.0.CO;2-4"},{"key":"e_1_2_7_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/34.531798"},{"key":"e_1_2_7_6_1","first-page":"50","volume-title":"Proceedings of ACM SIGIR","author":"Chien L.F.","year":"1997"},{"key":"e_1_2_7_7_1","doi-asserted-by":"publisher","DOI":"10.1177\/016555158701300203"},{"key":"e_1_2_7_8_1","doi-asserted-by":"publisher","DOI":"10.7202\/002692ar"},{"key":"e_1_2_7_9_1","doi-asserted-by":"publisher","DOI":"10.1177\/016555159201800208"},{"key":"e_1_2_7_10_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1007974605290"},{"key":"e_1_2_7_11_1","doi-asserted-by":"crossref","unstructured":"Fung P.(1995 June).A pattern matching method for finding noun and proper noun translations from noisy parallel corpora. Paper presented at the 33rd Annual Meeting of the Association for Computational Linguistics Boston MA.","DOI":"10.3115\/981658.981690"},{"key":"e_1_2_7_12_1","doi-asserted-by":"publisher","DOI":"10.1002\/1097-4571(2000)9999:9999<::AID-ASI1006>3.0.CO;2-#"},{"key":"e_1_2_7_13_1","volume-title":"Meaning\u2010based translation: A guide to cross\u2010language equivalence","author":"Larson M.L.","year":"1998"},{"issue":"4","key":"e_1_2_7_14_1","article-title":"Equivalence in translation: Between myth and reality","volume":"4","author":"Leonardi V.","year":"2000","journal-title":"Translation Journal"},{"key":"e_1_2_7_15_1","doi-asserted-by":"publisher","DOI":"10.1002\/asi.4630200106"},{"key":"e_1_2_7_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/3477.484439"},{"key":"e_1_2_7_17_1","unstructured":"Ma X. &Liberman M.(1999 September).BITS: A method for bilingual text search over the web. Paper presented at the Machine Translation Summit VII Kent Ridge Digital Labs National University of Singapore."},{"key":"e_1_2_7_18_1","unstructured":"Macklovitch E. &Hannan M.\u2010L.(1996).Line 'em up: Advances in alignment technology and their impact on translation support tools. Paper presented at the Second Conference of the Association for Machine Translation in the Americas (AMTA\u201096) Montr\u00e9al Qu\u00e9bec."},{"key":"e_1_2_7_19_1","first-page":"131","volume-title":"Proceedings of the AAAI Symposium in Cross\u2010Language Text and Speech Retrieval","author":"Oard D.W.","year":"1997"},{"key":"e_1_2_7_20_1","doi-asserted-by":"crossref","unstructured":"ResnikP.(1999 June).Mining the web for bilingual text. Paper presented at the 37th Annual Meeting of the Association for Computational Linguistics (ACL'99) College Park MD.","DOI":"10.3115\/1034678.1034757"},{"key":"e_1_2_7_21_1","first-page":"31","volume-title":"Translation spectrum: Essays in theory and practice","author":"Rose M.G.","year":"1981"},{"key":"e_1_2_7_22_1","volume-title":"Automatic text processing","author":"Salton G.","year":"1989"},{"key":"e_1_2_7_23_1","unstructured":"Simard M.(1999 June).Text\u2010translation alignment: Three languages are better than two. Paper presented at the Joint SIGDAT Conference on Empirical Methods in Natural Language Processing and Very Large Corpora College Park MD."},{"key":"e_1_2_7_24_1","unstructured":"Simard M. Foster G. &Isabelle P.(1992 June).Using cognates to align sentences in bilingual corpora. Paper presented at the Fourth International Conference on Theoretical and Methodological Issues in Machine Translation (TMI\u201092) Montreal Canada."},{"key":"e_1_2_7_25_1","doi-asserted-by":"crossref","unstructured":"Wu D.(1994 June).Aligning a parallel English\u2010Chinese corpus statistically with lexical criteria. Paper presented at the 32nd Annual Conference of the Association for Computational Linguistics Las Cruces New Mexico.","DOI":"10.3115\/981732.981744"},{"key":"e_1_2_7_26_1","doi-asserted-by":"publisher","DOI":"10.1002\/asi.10261"},{"key":"e_1_2_7_27_1","unstructured":"Yang C.C. &Li K.W.(2003b May).Generating cross\u2010lingual concept space from parallel corpora on the web. Paper presented at the International World Wide Web Conference Budapest Hungary."},{"key":"e_1_2_7_28_1","doi-asserted-by":"crossref","unstructured":"Yang C.C. &Li K.W.(2003c December).Segmenting Chinese unknown words by heuristic method. Paper presented at the International Conference on Asia Digital Libraries Malaysia.","DOI":"10.1007\/978-3-540-24594-0_52"},{"key":"e_1_2_7_29_1","doi-asserted-by":"publisher","DOI":"10.1002\/(SICI)1097-4571(2000)51:4<340::AID-ASI4>3.0.CO;2-I"},{"key":"e_1_2_7_30_1","doi-asserted-by":"publisher","DOI":"10.7202\/004638ar"}],"container-title":["Journal of the American Society for Information Science and Technology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/api.wiley.com\/onlinelibrary\/tdm\/v1\/articles\/10.1002%2Fasi.20118","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1002\/asi.20118","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,9,11]],"date-time":"2023-09-11T17:27:32Z","timestamp":1694453252000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1002\/asi.20118"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2005,1,12]]},"references-count":29,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2005,2]]}},"alternative-id":["10.1002\/asi.20118"],"URL":"https:\/\/doi.org\/10.1002\/asi.20118","archive":["Portico"],"relation":{},"ISSN":["1532-2882","1532-2890"],"issn-type":[{"value":"1532-2882","type":"print"},{"value":"1532-2890","type":"electronic"}],"subject":[],"published":{"date-parts":[[2005,1,12]]}}}