Jump to ContentJump to Main Navigation
Show Summary Details
More options …

Bio-Algorithms and Med-Systems

Editor-in-Chief: Roterman-Konieczna , Irena

4 Issues per year


CiteScore 2017: 0.43

SCImago Journal Rank (SJR) 2017: 0.160
Source Normalized Impact per Paper (SNIP) 2017: 0.223

Online
ISSN
1896-530X
See all formats and pricing
More options …

Automatic Protein Abbreviations Discovery and Resolution from Full-Text Scientific Papers: The PRAISED Framework

Daniele Toti / Paolo Atzeni / Fabio Polticelli
  • Department of Biology, University of Roma Tre, 00146 Rome, Italy
  • National Institute of Nuclear Physics, Roma Tre Section, 00146 Rome, Italy
  • Email
  • Other articles by this author:
  • De Gruyter OnlineGoogle Scholar

Abstract

This paper describes a methodology for discovering and resolving protein names abbreviations from the full-text versions of scientific articles, implemented in the PRAISED framework with the ultimate purpose of building up a publicly available abbreviation repository. Three processing steps lie at the core of the framework: i) an abbreviation identification phase, carried out via domain-independent metrics, whose purpose is to identify all possible abbreviations within a scientific text; ii) an abbreviation resolution phase, which takes into account a number of syntactical and semantic criteria in order to match an abbreviation with its potential explanation; and iii) a dictionary-based protein name identification, which is meant to select only those abbreviations belonging to the protein science domain. A local copy of the UniProt database is used as a source repository for all the known proteins. The PRAISED implementation has been tested against several known annotated corpora, such as the Medstract Gold Standard Corpus, the AB3P Corpus, the BioText Corpus and the Ao and Takagi Corpus, obtaining significantly high levels of recall and extremely fast performance, while also keeping promising levels of precision and overall f-measure, in comparison to the most relevant similar methods. This comparison has been carried out up to Phase 2, since those methods stop at expanding abbreviations, without performing any entity recognition. Instead, the entity recognition performed in the last phase provides PRAISED with an effective strategy for protein discovery, thus moving further from existing context-free techniques. Furthermore, this implementation also addresses the complexity of full-text papers, instead of the simpler abstracts more generally used. As such, the whole PRAISED process (Phase 1, 2 and 3) has been also tested against a manually annotated subset of full-text papers retrieved from the PubMed repository, with significant results as well.

Keywords: proteins; abbreviations; data mining; extraction; resolution

About the article

Received: 2011-07-27

Revised: 2011-11-22

Accepted: 2011-12-05

Published in Print:


Citation Information: Bio-Algorithms and Med-Systems BAMS, Volume 8, Pages 13–, ISSN (Online) 1896-530X, ISSN (Print) 1895-9091, DOI: https://doi.org/10.2478/bams-2012-0002.

Export Citation

©Jagiellonian University, Medical College, Kraków, Poland, 2012.

Citing Articles

Here you can find all Crossref-listed publications in which this article is cited. If you would like to receive automatic email messages as soon as this article is cited in other publications, simply activate the “Citation Alert” on the top of this page.

[1]
Daniele Toti and Andrea Longhi
Journal of Ambient Intelligence and Humanized Computing, 2017
[2]
Daniele Toti and Marco Rinelli
Journal of Ambient Intelligence and Humanized Computing, 2017, Volume 8, Number 2, Page 291

Comments (0)

Please log in or register to comment.
Log in