![]()
- Accuracy of genomic prediction is a measure of how reliably genomic information can be used to estimate the genetic merit of animals. In animal breeding, it is essential for determining whether genomic predictions can distinguish superior breeding candidates from those with lower genetic potential. High prediction accuracy enables breeders to make more confident selection decisions, improve genetic gain, and develop sustainable livestock populations. Accuracy is particularly important when selecting young animals that have limited personal performance records or no progeny information.
- Genomic prediction uses genome-wide genetic markers, phenotypic records, and statistical models to estimate an animal’s genomic estimated breeding value (GEBV). The reliability of these estimates depends on how well the prediction model captures the genetic factors influencing the target trait. Accuracy is influenced by the quality of the reference population, the heritability of the trait, genetic relationships between reference animals and selection candidates, marker information, and the statistical method used. Understanding these factors helps breeding organizations evaluate the strengths and limitations of genomic selection.
- In quantitative genetics, the accuracy of an estimated breeding value is commonly defined as the correlation between the estimated breeding value and the animal’s true breeding value. It is generally represented by the symbol r and ranges from 0 to 1 when expressed as a non-negative accuracy measure. An accuracy close to 1 indicates that the prediction closely reflects the true genetic merit, while a lower value indicates greater uncertainty. The true breeding value is not directly observable, so accuracy must be estimated using suitable validation methods or statistical theory rather than measured directly in routine breeding programs.
- Accuracy should be distinguished from reliability, which is commonly defined as the squared correlation between the estimated and true breeding values. Under the standard definition, reliability is the square of accuracy, so an accuracy of 0.80 corresponds to a reliability of 0.64. Some genetic evaluation systems report reliability or an equivalent measure rather than accuracy; therefore, the reported statistic should always be checked before results are compared. High reliability generally indicates greater confidence in an estimate, but it does not guarantee that every selection decision will produce the expected outcome.
- One of the most important factors affecting genomic prediction accuracy is the size and quality of the reference population. A reference population consists of animals with genomic information and reliable phenotypic records or appropriate genetic evaluations. These animals provide the training data needed to establish relationships between genetic markers and the trait being predicted. Larger reference populations often improve accuracy by providing more information about genetic effects, particularly for complex traits. However, increasing the number of animals does not automatically guarantee better predictions if their records are unreliable or their genetic backgrounds differ substantially from those of the selection candidates.
- The heritability of a trait also influences prediction accuracy. Heritability describes the proportion of phenotypic variance attributable to additive genetic variance in a defined population and environment. Traits with higher heritability generally provide stronger information about genetic differences among animals, which can support more accurate genomic predictions when other factors are comparable. Traits with low heritability, such as some fertility, health, and disease-resistance characteristics, may require larger reference populations and more precise phenotypic records to achieve useful accuracy. However, heritability alone does not determine prediction accuracy; reference population size, genetic relationships, trait architecture, and data quality also matter.
- Genetic relationships between reference animals and selection candidates are another major influence. Genomic prediction can be especially accurate when candidates share recent ancestry with animals in the reference population because genomic similarities provide information about inherited genetic effects. A genomic relationship matrix can quantify the genetic similarity among animals using genome-wide marker data. As the genetic distance between the reference population and future candidates increases, prediction accuracy may decline, particularly when predictions depend heavily on close relationships. Regularly updating the reference population can help maintain prediction performance across generations.
- The genetic architecture of the trait also affects accuracy. Many livestock traits are polygenic, meaning that they are influenced by numerous genetic variants with small effects. Genomic prediction is designed to combine information across the genome rather than relying exclusively on a few individually significant markers. However, traits influenced by rare variants, complex interactions, or genetic effects that are poorly represented in the reference population may be more difficult to predict. The choice of statistical model should reflect the available data and the likely distribution of genetic effects.
- Different statistical methods can produce different levels of genomic prediction accuracy. Genomic best linear unbiased prediction (GBLUP) uses a genomic relationship matrix to estimate genetic merit, while Bayesian methods estimate marker effects under different assumptions about their distribution. Other approaches may incorporate haplotypes, sequence variants, or machine-learning methods. No single method is best for every trait or population. The most appropriate model depends on the size and structure of the reference population, the genetic architecture of the trait, marker density, and the quality of phenotypic and genomic information. Model performance should be assessed through appropriate validation rather than assumed from model complexity alone.
- The quality of phenotypic records is critical because these records provide much of the information used to train genomic prediction models. Inaccurate measurements, inconsistent trait definitions, missing data, or inadequate adjustment for environmental effects can weaken the relationship between genotypes and observed performance. For example, differences in milk yield may reflect feeding conditions, herd management, animal age, or health as well as genetic merit. Appropriate statistical models must account for relevant environmental and management effects so that the reference data provide a more reliable estimate of genetic differences.
- Genomic data quality also affects prediction accuracy. Errors in sample identification, missing genotypes, low-quality markers, and inconsistent genomic processing can introduce noise into the prediction model. Genotyping arrays, sequencing technologies, and genotype imputation methods may provide different levels of marker information and data quality. Quality control procedures should ensure that the genomic data for reference animals and selection candidates are compatible and sufficiently reliable for the intended analysis. More markers do not necessarily improve prediction if the additional data are inaccurate or contribute little useful information.
- Validation is essential for assessing genomic prediction accuracy. Researchers commonly evaluate prediction models using animals or records that were not used to train the model. Depending on the breeding objective, validation may involve later-born animals, separate herds, independent populations, or other data that reflect the intended application. Validation designs must avoid data leakage, in which information from the validation animals indirectly enters the training process. Randomly dividing closely related animals between training and validation groups may give an overly optimistic impression of how well predictions will work in unrelated animals or future generations.
- Prediction accuracy can be estimated in different ways, and the method used affects how results should be interpreted. In research, accuracy may be assessed by comparing genomic predictions with an independent estimate of genetic merit, such as a reliable progeny-based evaluation. When only adjusted phenotypes are available, their correlation with genomic predictions is not automatically equal to the true accuracy because phenotypes contain environmental variation and may have limited reliability as measures of genetic merit. Researchers therefore need to account for the properties of the validation data and the assumptions of the chosen evaluation method.
- Accurate genomic prediction is particularly valuable for traits that are expensive, difficult, or impractical to measure directly on every breeding candidate. These include feed efficiency, carcass characteristics, methane emissions, disease resistance, fertility, and longevity. If reliable reference data are available, genomic predictions can help identify promising animals before these traits are directly measured on them. However, accuracy for each trait should be evaluated separately because the amount and quality of information can differ substantially between traits.
- Multi-trait genomic prediction can sometimes improve accuracy by using information from genetically correlated traits. For example, a trait that is easier to measure may provide useful information about a more expensive or difficult-to-measure trait if the genetic correlation between them is sufficiently strong. The improvement depends on the strength and stability of the genetic correlation, the reliability of the records, and the population structure. If the traits are weakly correlated or the supporting data are poor, combining them may provide little benefit or even reduce prediction performance.
- Prediction accuracy also affects the expected response to selection. When the selection intensity and additive genetic variation remain comparable, more accurate estimates generally help breeders identify animals with greater genetic merit. Genomic prediction can also enable selection at a younger age, shortening the generation interval. The expected rate of genetic gain per year is commonly expressed as:
- ΔG/year = (i × r × σ_A) / L
- Here, ΔG/year is the expected genetic gain per year, i is selection intensity, r is the accuracy of selection, σ_A is the additive genetic standard deviation, and L is the generation interval. This relationship illustrates why accuracy is important, but it also shows that genetic improvement depends on several factors. Maximizing prediction accuracy alone may not maximize long-term progress if doing so increases costs, delays selection, or reduces genetic diversity.
- In practical breeding programs, accuracy must be considered alongside genetic diversity, animal health, fertility, welfare, and long-term breeding objectives. Selecting only the animals with the highest genomic predictions may concentrate genetic contributions in a small number of families and increase inbreeding. Genomic relationships and methods such as optimal contribution selection can help balance expected genetic gain with the maintenance of genetic diversity. Breeding organizations should also monitor realized performance and update reference populations as new generations and records become available.
- Overall, accuracy of genomic prediction is a central factor determining the usefulness of genomic selection in animal breeding. It depends on reference population size and relevance, heritability, genetic relationships, phenotypic and genomic data quality, statistical methods, and validation design. By understanding and monitoring these influences, breeding organizations can make more reliable selection decisions, improve genetic evaluation, and increase the effectiveness of livestock improvement programs.