Detecting and classifying online health misinformation with 'Content Similarity Measure (CSM)' algorithm : an automated fact-checking-based approach

© The Author(s), under exclusive licence to Springer Science+Business Media, LLC, part of Springer Nature 2023, Springer Nature or its licensor (e.g. a society or other partner) holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author se...

Ausführliche Beschreibung

Bibliographische Detailangaben
Veröffentlicht in:The Journal of supercomputing. - 1998. - 79(2023), 8 vom: 16., Seite 9127-9156
1. Verfasser: Barve, Yashoda (VerfasserIn)
Weitere Verfasser: Saini, Jatinderkumar R
Format: Online-Aufsatz
Sprache:English
Veröffentlicht: 2023
Zugriff auf das übergeordnete Werk:The Journal of supercomputing
Schlagworte:Journal Article Content Similarity Measure Content similarity score Fact-checking Healthcare Misinformation detection Natural language processing Similarity measures
LEADER 01000caa a22002652 4500
001 NLM351554009
003 DE-627
005 20240916231949.0
007 cr uuu---uuuuu
008 231226s2023 xx |||||o 00| ||eng c
024 7 |a 10.1007/s11227-022-05032-y  |2 doi 
028 5 2 |a pubmed24n1535.xml 
035 |a (DE-627)NLM351554009 
035 |a (NLM)36644509 
040 |a DE-627  |b ger  |c DE-627  |e rakwb 
041 |a eng 
100 1 |a Barve, Yashoda  |e verfasserin  |4 aut 
245 1 0 |a Detecting and classifying online health misinformation with 'Content Similarity Measure (CSM)' algorithm  |b an automated fact-checking-based approach 
264 1 |c 2023 
336 |a Text  |b txt  |2 rdacontent 
337 |a ƒaComputermedien  |b c  |2 rdamedia 
338 |a ƒa Online-Ressource  |b cr  |2 rdacarrier 
500 |a Date Revised 16.09.2024 
500 |a published: Print-Electronic 
500 |a Citation Status PubMed-not-MEDLINE 
520 |a © The Author(s), under exclusive licence to Springer Science+Business Media, LLC, part of Springer Nature 2023, Springer Nature or its licensor (e.g. a society or other partner) holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law. 
520 |a Information dissemination occurs through the 'word of media' in the digital world. Fraudulent and deceitful content, such as misinformation, has detrimental effects on people. An implicit fact-based automated fact-checking technique comprising information retrieval, natural language processing, and machine learning techniques assist in assessing the credibility of content and detecting misinformation. Previous studies focused on linguistic and textual features and similarity measures-based approaches. However, these studies need to gain knowledge of facts, and similarity measures are less accurate when dealing with sparse or zero data. To fill these gaps, we propose a 'Content Similarity Measure (CSM)' algorithm that can perform automated fact-checking of URLs in the healthcare domain. Authors have introduced a novel set of content similarity, domain-specific, and sentiment polarity score features to achieve journalistic fact-checking. An extensive analysis of the proposed algorithm compared with standard similarity measures and machine learning classifiers showed that the 'content similarity score' feature outperformed other features with an accuracy of 88.26%. In the algorithmic approach, CSM showed improved accuracy of 91.06% compared to the Jaccard similarity measure with 74.26% accuracy. Another observation is that the algorithmic approach outperformed the feature-based method. To check the robustness of the algorithms, authors have tested the model on three state-of-the-art datasets, viz. CoAID, FakeHealth, and ReCOVery. With the algorithmic approach, CSM showed the highest accuracy of 87.30%, 89.30%, 85.26%, and 88.83% on CoAID, ReCOVery, FakeHealth (Story), and FakeHealth (Release) datasets, respectively. With a feature-based approach, the proposed CSM showed the highest accuracy of 85.93%, 87.97%, 83.92%, and 86.80%, respectively 
650 4 |a Journal Article 
650 4 |a Content Similarity Measure 
650 4 |a Content similarity score 
650 4 |a Fact-checking 
650 4 |a Healthcare 
650 4 |a Misinformation detection 
650 4 |a Natural language processing 
650 4 |a Similarity measures 
700 1 |a Saini, Jatinderkumar R  |e verfasserin  |4 aut 
773 0 8 |i Enthalten in  |t The Journal of supercomputing  |d 1998  |g 79(2023), 8 vom: 16., Seite 9127-9156  |w (DE-627)NLM098252410  |x 0920-8542  |7 nnns 
773 1 8 |g volume:79  |g year:2023  |g number:8  |g day:16  |g pages:9127-9156 
856 4 0 |u http://dx.doi.org/10.1007/s11227-022-05032-y  |3 Volltext 
912 |a GBV_USEFLAG_A 
912 |a SYSFLAG_A 
912 |a GBV_NLM 
912 |a GBV_ILN_350 
951 |a AR 
952 |d 79  |j 2023  |e 8  |b 16  |h 9127-9156