https://doi.org/10.4081/jbr.2026.14607
Trinucleotide Repeat (TNR) expansions mining algorithm based on the Apriori algorithm
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.
Published: 7 August 2026
The Apriori algorithm is a data mining algorithm used exclusively for identifying associations between itemsets in a series of business transactions. In this paper, we modify this algorithm to identify associations between trinucleotide repeat expansions in DNA, which are known for their associations with specific genetic diseases. The proposed modification here is to treat all possible forms of DNA trinucleotides as item sets and the list of all DNA segments as a series of transactions. In the first iteration of the algorithm, the number of occurrences of every single trinucleotide is counted across all DNA segments. Then, the results are filtered out according to a preset threshold. In the second iteration, the number of occurrences of every pair of trinucleotides is counted, and similarly, the results are filtered out according to a threshold. In the third iteration, the number of occurrences of every trio of trinucleotides is handled similarly. The algorithm continues until no results are produced. Lastly, the association rules are formed, yielding the conclusion of the existence of a certain genetic disease will also mean an existence of another genetic disease.
Downloads
[1] Mitas M. Trinucleotide repeats associated with human disease. Nucleic Acids Res 1997;25:2245–53. DOI: https://doi.org/10.1093/nar/25.12.2245
[2] Kovtun I, McMurray C. Features of trinucleotide repeat in-stability in vivo. Cell Res 2008;18:198–213. DOI: https://doi.org/10.1038/cr.2008.5
[3] Samadash-wily G, Raca G, Mirkin S. Trinucleotide repeats affect DNA replication in vivo. Nat. Genet 1997;17:298–304. DOI: https://doi.org/10.1038/ng1197-298
[4] Lieberman A, Shakkottai V, Albin R. Polyglutamine repeats in neurodegenerative diseases. Annu Rev Pathol 2019;14:1–27. DOI: https://doi.org/10.1146/annurev-pathmechdis-012418-012857
[5] Everett C, Wood N. Trinucleotide repeats and neurodegenerative disease. Brain 2004;127:2385–405. DOI: https://doi.org/10.1093/brain/awh278
[6] La Spada A, Paulson H, Fischbeck K. Trinucleotide repeat expansion in neurological disease. Ann Neurol 1994;36:814–22. DOI: https://doi.org/10.1002/ana.410360604
[7] Pearson C, Edamura K, Cleary J. Repeat in-stability: mechanisms of dynamic mutations. Nat Rev Genet 2005;6:729–42. DOI: https://doi.org/10.1038/nrg1689
[8] Richards R. Dynamic mutations: a decade of unstable expanded repeats in human genetic disease. Hum Mol Genet 2001;10:2187–94. DOI: https://doi.org/10.1093/hmg/10.20.2187
[9] Sarkar S. Decoding “coding”: Information and DNA. BioScience 1996;46:857–64. DOI: https://doi.org/10.2307/1312971
[10] Nogueira F, De Rosa J, Vicente E, et al. RNA expression profiles and data mining of sugarcane response to low temperature. Plant Physiol 2003;132:1811–24. DOI: https://doi.org/10.1104/pp.102.017483
[11] Fu X. Non-coding RNA: a new frontier in regulatory biology. Natl Sci Rev 2014;1:190–204. DOI: https://doi.org/10.1093/nsr/nwu008
[12] Gusella J, MacDonald M, Ambrose C, Duyao M. Molecular genetics of Huntington’s disease. Arch Neurol 1993;50:1157–63. DOI: https://doi.org/10.1001/archneur.1993.00540110037003
[13] Imarisio S, Carmichael J, Korolchuk V, et al. Huntington’s disease: from pathology and genetics to potential therapies. Biochem J 2008;412:191–209. DOI: https://doi.org/10.1042/BJ20071619
[14] Santoro M, Bray S, Warren S. Molecular mecha-nisms of fragile x syndrome: A twenty-year perspective. Annu Rev Pathol 2012;7:219–45. DOI: https://doi.org/10.1146/annurev-pathol-011811-132457
[15] Jin P, Warren S. Understanding the molecular basis of fragile x syndrome. Hum Mol Genet 2000;9:901–8. DOI: https://doi.org/10.1093/hmg/9.6.901
[16] Turner C, Hilton-Jones D. The myotonic dystrophies: diagnosis and management. J Neurol Neurosurg Psychiatry 2010;81:358–67. DOI: https://doi.org/10.1136/jnnp.2008.158261
[17] Lee J, Cooper T. Pathogenic mechanisms of myotonic dystrophy. Biochem Soc Trans 2009;37:1281–6. DOI: https://doi.org/10.1042/BST0371281
[18] Manto M. The wide spectrum of spinocerebellar ataxias (scas). Cerebellum 2005;4:2–6. DOI: https://doi.org/10.1080/14734220510007914
[19] Koeppen A. The pathogenesis of spinocerebellar ataxia. Cerebellum 2005;4:62–73. DOI: https://doi.org/10.1080/14734220510007950
[20] Illarioshkin S, Igarashi S, Onodera O, et al. Trinucleotide re-peat length and rate of progression of Huntington’s disease. Ann Neurol 1994;36:630–5. DOI: https://doi.org/10.1002/ana.410360412
[21] Kieburtz K, MacDonald M, Shih C, et al. Trinucleotide repeat length and progression of illness in Huntington’s disease. J Med Genet 1994;31:872–4. DOI: https://doi.org/10.1136/jmg.31.11.872
[22] Lee J, Conrad A, Epping E, et al. Effect of trinucleotide repeats in the Huntington’s gene on intelligence. eBioMedicine 2018;31:47–53. DOI: https://doi.org/10.1016/j.ebiom.2018.03.031
[23] Andrew S, Goldberg Y, Kremer B, et al. The relationship between trinucleotide (cag) repeat length and clinical features of Huntington’s disease. Nat Genet 1993;4:398–403. DOI: https://doi.org/10.1038/ng0893-398
[24] Rosenblatt A, Kumar B, Mo A, et al. Age, cag repeat length, and clinical progression in Huntington’s disease. Mov Disord 2012;27:272–6. DOI: https://doi.org/10.1002/mds.24024
[25] Bañez-Coronel M, Porta S, Kagerbauer B, et al. A pathogenic mechanism in Huntington’s disease involves small cag-repeated RNAs with neurotoxic activity. PLoS Genet 2012;8:1–15, 02. DOI: https://doi.org/10.1371/journal.pgen.1002481
[26] Agrawal R, Imieliński T, Swami A. Mining association rules between sets of items in large databases. SIGMOD Rec 1993;22:207–16. DOI: https://doi.org/10.1145/170036.170072
[27] Dongre J, Prajapati G, Tokekar S. The role of apriori algorithm for finding the association rules in data mining. Proceedings of the 2014 International Conference on Issues and Challenges in Intelligent Computing Techniques (ICICT) 2014; pp. 657–60. DOI: https://doi.org/10.1109/ICICICT.2014.6781357
[28] Yuan X. An improved apriori algorithm for mining association rules. AIP Conference Proceedings 2017;1820:080005.
[29] Larose D. Introduction to Data Mining, chapter 1, John Wiley & Sons, Ltd, 2004; pp. 1–26.
[30] Jackson J. Data mining; a conceptual overview. Commun Assoc Inf Syst 2002;8:19. DOI: https://doi.org/10.17705/1CAIS.00819
[31] Zhang S, Zhang C, Yang Q. Data preparation for data mining. Appl Artif Intell 2003;17:375–81. DOI: https://doi.org/10.1080/713827180
[32] Yuan X. An improved apriori algorithm for mining association rules. AIP Conference Proceedings 2017;1820:080005. DOI: https://doi.org/10.1063/1.4977361
[33] Du J, Zhang X, Zhang H, Chen L. Research and improvement of apriori algorithm. Proceedings of the 2016 International Conference on Issues and Challenges in Intelligent Computing Techniques (ICICT) 2016; pp. 117–21. DOI: https://doi.org/10.1109/ICIST.2016.7483396
[34] Perego R, Orlando S, Palmerini P. Enhancing the apriori algorithm for frequent set counting. In Kambayashi Y, Winiwarter W, Arikawa M, eds. Data Warehousing and Knowledge Discovery. Berlin Heidelberg, Springer Nature. 2001; pp. 71–82. DOI: https://doi.org/10.1007/3-540-44801-2_8
[35] Aflori C, Craus M. Grid implementation of the apriori algorithm. Adv Eng Softw 2007;38:295–300. DOI: https://doi.org/10.1016/j.advengsoft.2006.08.011
[36] Al-Turaiki I, Badr G, Mathkour H. Trieamd: a scalable and efficient apriori motif discovery approach. Int J Data Min Bioinform 2015;13:13–30. DOI: https://doi.org/10.1504/IJDMB.2015.070833
[37] Alaqeeli O. A comparison of dropout rate of three commonly used sin-gle cell RNA-sequencing protocols. Biotechnol Biotechnol Equip 2024;38:237837. DOI: https://doi.org/10.1080/13102818.2024.2379837
[38] Alaqeeli O,Alturki R. Evaluating the performance of the generalized linear model (glm) r package using single-cell RNA-sequencing data. Appl Sci 2023;13:11512. DOI: https://doi.org/10.3390/app132011512
[39] Schalling M, Hudson T, Buetow K, Housman D. Direct detection of novel expanded trinucleotide repeats in the human genome. Nat Genet 1993;4:135–9. DOI: https://doi.org/10.1038/ng0693-135
[40] Richards R, Sutherland G. Dynamic mutation: possible mechanisms and significance in human disease. Trends Biochem Sci 1997;22:432–6. DOI: https://doi.org/10.1016/S0968-0004(97)01108-0
Ethics Approval
Supporting Agencies
Data Availability Statement
Datasets used in this research can be found at: https://www.ncbi.nlm.nih.gov/Traces/study/?acc=PRJNA701282 The result of processing these datasets can be found at: https://github.com/Omar-Alaqeeli/TNR-AprioriAlgorithm
How to Cite

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
PAGEPress has chosen to apply the Creative Commons Attribution NonCommercial 4.0 International License (CC BY-NC 4.0) to all manuscripts to be published.