Please use this identifier to cite or link to this item:
Title: Compressing population DNA sequences using multiple reference sequences
Author(s): Siu, Wan Chi 
Author(s): Cheng, K.-O.
Law, N.-F.
Issue Date: 2017
Publisher: IEEE
Related Publication(s): 2017 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) Proceedings
Start page: 760
End page: 764
Compressing population DNA sequences often relies on the use of a reference sequence so that only the differences between the target DNA sequences to be compressed and the reference sequence are encoded. Despite the importance of the choice of the reference sequence, state-of-the-art algorithms in population sequence compression often selected one of the population sequences as a reference sequence in an ad hoc manner. In this paper, we investigated issues about the choice of the reference sequence. In particular, population sequences are first clustered into a number of groups. A reference sequence is then obtained for each group so that substructures within each group can be characterized by this reference sequence. Afterwards, the reference sequence is used to compress sequences within that group. In this way, the multiple reference sequences framework can optimize the overall compression performance on the set of population sequences. Results show that our proposed method reduces the compressed size by up to 91% as compared to state-of-the-art reference- based approaches.
DOI: 10.1109/APSIPA.2017.8282136
CIHE Affiliated Publication: No
Appears in Collections:CIS Publication

SFX Query Show full item record

Google ScholarTM




Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.