Document Type : Research Paper
Author
Department of Animal sciences, Animal breeding and genetic branch, Faculty of Agriculture, Lorestan University, Khorramabad, Iran
Abstract
Keywords
Main Subjects
Extended Abstract
Introduction
When large amount of genomic markers became available in the last decade, livestock breeding programs have evolved rapidly through genomic selection (GS). GS refers to selection of individuals based on the aggregate genetic merit of mentioned markers. Briefly, GS involves application of a genomic statistical method in a reference population which has both phenotypic and genotypic data, and predicting genomic breeding value (GBV) in a candidate population which has only genotypic data. In context of GS in livestock, several statistical methods have been proposed to predict marker effects which are consequently used to estimate GBV of individuals. However, as the number of markers is very large relative to the phenotypes, genomic selection is faced to the challenge of the “curse of dimensionality”. Bayesian methods are usually proposed to handle these large sets. Bayesian LASSO is a penalized regression model with the potential of variable selection using the regularization parameter λ, which has a more prominent role than the other Bayesian methods. In a fully Bayesian framework, it is possible to treat λ as a random parameter and therefore different priors can be assigned to it. Despite the theoretical foundations of developing different priors and hyper parameters in Bayesian LASSO, the potential of applying different priors to this model has been less evaluated in genomic prediction studies. Therefore, the aim of the present study was to evaluate the application of different priors and hyper parameters to Bayesian LASSO regularization parameter on the genomic prediction of traits with different genetic architectures using real and simulated data of a mouse population. Variance component estimation and predictability of the models measured by prediction accuracy and bias were targeted.
Material and methods
In general, two datasets were analyzed in the present study. First, a real dataset consisted of phenotypic and genotypic data from a mice study. Second, a simulated dataset using mentioned real mice genome. The used phenotypes and genotypes and how they were collected have been fully described in previous studies, and these datasets are freely available (http://gscan.well.ox.ac.uk) and can be found in popular libraries in R such as BGLR.
Bayesian Least Absolute Shrinkage and Selection Operator (LASSO) was used to investigate its predictability under different priors for regularization parameter. Bayesian LASSO is a hierarchical model in which a double exponential distribution is considered as the marginal distribution of marker effect. This distribution is implemented as mixture of scaled normal densities by assigning independent normal densities to marker effects in the first level of the hierarchy. Then, IID exponential densities were assigned to the marker-specific scale parameters at the second hierarchy. Different priors were used for rate parameter at the final hierarchy. The effect of assigning three different priors including gamma, beta, and fixed distributions was investigated using three real traits of mice data and four simulated traits based on the real mice genome. These traits are selected to represent a range of different genetic architectures. Markov chain Monte Carlo (MCMC) was used to estimate the posterior distribution of variances component and the model parameters. The capability of the different models was compared by considering the correlation between the observed phenotypes and predicted GBV, and the regression slope of observed phenotypes onto predicted GBV from a fivefold cross validation. Estimated variance components and heritability of traits were also compared.
Results
Estimation of variance components and heritability are strongly affected by the assigned prior. The fixed λ model estimated the highest genetic variance and heritability and the lowest residual variance for different real traits. Regardless of the type of simulated trait, three priors underestimated the variance component and the heritability of traits. The average accuracy obtained for each trait using different priors was relatively similar, but standard deviation of the accuracies by the fixed λ model was always greater than the other two priors. Comparison of the minimum and maximum values of the models accuracies via fivefold cross validation revealed that the highest accuracy values were obtained using the assignment of gamma and beta distributions to the regularization parameter. By increasing the heritability and the number of background QTL of the traits, the accuracy of the models increased and the difference between the three different priors decreased. The difference between the models was greater in terms of prediction bias, but it generally followed the same trend of accuracy, i.e. the models with the highest accuracy also had the lowest prediction bias.
Conclusion
Based on the present results, the effect of prior assignment to λ is more critical for traits that are influenced by a small number of QTL or have low heritability.
Analyzed data in the present work are freely available at (http://gscan.well.ox.ac.uk) and can be found in popular libraries in R such as BGLR.
This study did not involve data collection, and pre-existing data were used. The author avoided data fabrication, falsification, plagiarism, and misconduct.
The authors declare that they have no competing interests.