![]()
- Reference populations are groups of animals with both genomic information and reliable phenotypic records, or suitable genetic evaluation data, that are used to develop and improve genomic prediction models. In modern animal breeding, reference populations are essential for estimating genomic breeding values and identifying animals with high genetic merit. They provide the information needed to establish relationships between genetic markers and economically important traits, helping breeders make more accurate selection decisions and accelerate genetic improvement in livestock.
- The main purpose of a reference population is to connect an animal’s genome-wide genetic marker information with its measured performance or estimated genetic merit. Animals in the population are typically genotyped using single-nucleotide polymorphism (SNP) markers and evaluated for traits such as growth rate, milk production, feed efficiency, fertility, disease resistance, carcass quality, and longevity. Statistical models use these data to estimate marker effects or genomic relationships and develop predictions that can be applied to other genotyped animals, known as selection candidates.
- A reference population is a fundamental component of genomic selection because the accuracy of genomic predictions depends strongly on the information available for training the prediction model. A well-designed reference population allows the model to learn how genetic variation relates to differences in performance. The resulting model can then estimate genomic estimated breeding values (GEBVs) for young animals or other candidates whose genetic merit needs to be assessed. These predictions may be available before the candidates have accumulated their own performance records or produced offspring.
- The quality of a reference population depends on several factors, including the number of animals, the accuracy of their phenotypic records, the reliability of their genetic evaluations, and the quality of their genomic data. A larger population often improves prediction accuracy because it provides more information about genetic effects and relationships. However, increasing population size alone does not guarantee better predictions. Animals must have relevant records, reliable genotypes, and sufficient genetic similarity to the target population for the additional data to provide meaningful benefits.
- Phenotypic recording is particularly important when developing reference populations. Traits such as milk yield, body weight, feed intake, fertility, and disease resistance must be measured consistently using appropriate definitions and recording procedures. Researchers and breeding organizations should account for relevant environmental and management factors, including herd, age, sex, feeding conditions, and production system. If phenotypic records contain substantial errors or systematic bias, the genomic prediction model may learn misleading relationships and produce less reliable breeding values.
- The genetic composition of a reference population also affects its usefulness. A population containing animals closely related to the selection candidates often provides more accurate genomic predictions because shared genetic relationships help predict inherited effects. However, a reference population that represents only a narrow group of families may perform poorly when applied to genetically different animals. Including sufficient genetic diversity and relevant families can improve the usefulness of the data and help maintain prediction performance across a wider range of candidates.
- Reference populations may be developed within a single breed, across multiple breeds, or through collaboration among breeding organizations. A within-breed reference population is commonly used when the selection candidates belong to the same breed as the training animals. Multi-breed reference populations may increase the amount of available information and improve prediction in some circumstances, particularly when breeds share relevant genetic variants. However, differences in allele frequencies, linkage disequilibrium, genetic backgrounds, and trait expression can reduce the benefits of combining breeds. Multi-breed prediction therefore requires suitable statistical models and careful validation.
- Another important consideration is the genetic relationship between the reference animals and the candidates being evaluated. Genomic predictions can be particularly informative when reference animals and candidates share recent ancestry. As generations pass, genetic relationships change and the prediction model may lose some of its predictive power. Regularly adding newly recorded and genotyped animals can help update the reference population, maintain prediction accuracy, and reflect changes in the breeding population. The need for updating depends on the species, trait, breeding structure, and rate of genetic turnover.
- The choice of genomic technology is another part of reference population design. Many breeding programs use SNP genotyping arrays because they provide marker information across the genome at a practical cost. Some programs also use whole-genome sequencing, sequence-derived variants, or genotype imputation to increase the amount of genomic information available. Whatever technology is used, quality control is necessary to identify unreliable samples, missing genotypes, inconsistent animal identities, and low-quality markers. Consistent genomic data processing helps ensure that the reference animals and selection candidates can be evaluated using compatible information.
- Reference populations support different statistical methods for genomic prediction, including genomic best linear unbiased prediction (GBLUP) and Bayesian approaches. GBLUP uses a genomic relationship matrix to estimate genetic similarity among animals, while Bayesian methods estimate marker effects under specified statistical assumptions. Both approaches depend on the quality and relevance of the reference data. The appropriate method is determined by the characteristics of the trait, the genetic architecture, the available records, and the design of the breeding program.
- The size of a reference population is particularly important for traits with low heritability or complex genetic architecture. Such traits may require more animals and higher-quality phenotypic records to achieve useful prediction accuracy. Traits that are expensive to measure, such as feed efficiency, methane emissions, some disease-resistance characteristics, or carcass traits measured after slaughter, may benefit from carefully designed reference populations that collect specialized records from representative animals. Once the relationship between genomic information and these traits has been established, predictions may be extended to candidates that have not been measured directly.
- Reference populations can also be used for traits expressed in only one sex or later in life. For example, male selection candidates can receive genomic predictions for female reproductive or milk-production traits when suitable reference data exist. Similarly, young animals may be evaluated for longevity or disease-related traits before sufficient direct records are available. The reliability of these predictions depends on the quality of the reference information, the genetic relationships between traits, and the relevance of the population to the animals being evaluated.
- One challenge in developing reference populations is the cost of collecting genomic and phenotypic information. Genotyping large numbers of animals and maintaining accurate long-term performance records require investment, technical expertise, and coordinated data management. Breeding organizations can reduce these challenges through collaboration, shared reference populations, standardized trait definitions, and efficient recording systems. However, data sharing must preserve data quality, animal identification, ownership agreements, and appropriate privacy or commercial restrictions.
- Reference populations are also important for managing genetic diversity. If the population used to train a genomic prediction model represents only a small number of related families, predictions may be less transferable and breeding decisions may concentrate genetic contributions in a limited group of animals. Including appropriate genetic diversity and monitoring genomic relationships can help improve robustness. Reference population design should therefore consider both short-term prediction accuracy and the long-term sustainability of the breeding program.
- The effectiveness of a reference population should be evaluated through genomic prediction validation. Researchers may compare predicted breeding values with reliable later-life performance records, progeny information, or independent genetic evaluations. Validation should avoid data leakage, in which information from the same animals or closely related records unintentionally influences both model training and testing. Suitable validation strategies depend on the intended use of the model. For example, predicting younger generations requires validation procedures that reflect the genetic distance between the reference animals and future selection candidates.
- Reference populations should be viewed as dynamic resources rather than fixed datasets. As new animals are born, new traits become important, production environments change, and genetic improvement progresses, breeding organizations may need to update the data and retrain prediction models. Monitoring prediction accuracy, genetic trends, trait recording quality, and population representation helps ensure that genomic evaluations remain useful over time. Maintaining well-documented data and reproducible analytical procedures also supports consistency between evaluation cycles.
- Overall, reference populations form the foundation of reliable genomic prediction in modern animal breeding. By combining high-quality genomic data with accurate phenotypic and genetic evaluation records, they enable prediction models to estimate the genetic merit of selection candidates, improve genomic estimated breeding values, and support faster genetic progress. Carefully designed and regularly updated reference populations can also improve the usefulness of genomic selection across generations while supporting genetic diversity and sustainable livestock improvement.