TL;DR: A simple new approach to variable selection in linear regression, with a particular focus on quantifying uncertainty in which variables should be selected, is introduced, based on a new model based on the ‘sum of single effects’ model, called ‘SuSiE’.
Abstract: We introduce a simple new approach to variable selection in linear regression, with a particular focus on quantifying uncertainty in which variables should be selected. The approach is based on a new model—the ‘sum of single effects’ model, called ‘SuSiE’—which comes from writing the sparse vector of regression coefficients as a sum of ‘single‐effect’ vectors, each with one non‐zero element. We also introduce a corresponding new fitting procedure—iterative Bayesian stepwise selection (IBSS)—which is a Bayesian analogue of stepwise selection methods. IBSS shares the computational simplicity and speed of traditional stepwise methods but, instead of selecting a single variable at each step, IBSS computes a distribution on variables that captures uncertainty in which variable to select. We provide a formal justification of this intuitive algorithm by showing that it optimizes a variational approximation to the posterior distribution under SuSiE. Further, this approximate posterior distribution naturally yields convenient novel summaries of uncertainty in variable selection, providing a credible set of variables for each selection. Our methods are particularly well suited to settings where variables are highly correlated and detectable effects are sparse, both of which are characteristics of genetic fine mapping applications. We demonstrate through numerical experiments that our methods outperform existing methods for this task, and we illustrate their application to fine mapping genetic variants influencing alternative splicing in human cell lines. We also discuss the potential and challenges for applying these methods to generic variable‐selection problems.
TL;DR: The importance of including appropriate variables, following the proper steps, and adopting the proper methods when selecting variables for prediction models is focused on.
Abstract: Clinical prediction models are used frequently in clinical practice to identify patients who are at risk of developing an adverse outcome so that preventive measures can be initiated. A prediction model can be developed in a number of ways; however, an appropriate variable selection strategy needs to be followed in all cases. Our purpose is to introduce readers to the concept of variable selection in prediction modelling, including the importance of variable selection and variable reduction strategies. We will discuss the various variable selection techniques that can be applied during prediction model building (backward elimination, forward selection, stepwise selection and all possible subset selection), and the stopping rule/selection criteria in variable selection (p values, Akaike information criterion, Bayesian information criterion and Mallows’ Cp statistic). This paper focuses on the importance of including appropriate variables, following the proper steps, and adopting the proper methods when selecting variables for prediction models.
TL;DR: In recent years, the terms machine learning, data science, and predictive modeling have become ubiquitous in nearly every discipline in which data analysis plays a central role as discussed by the authors. But when first diving into the field of data analysis, it was not a new topic.
Abstract: In recent years, the terms machine learning, data science, and predictive modeling have become ubiquitous in nearly every discipline in which data analysis plays a central role. When first diving i...
TL;DR: A new method to binarize a continuous pigeon inspired optimizer is proposed and compared to the traditional way for binarizing continuous swarm intelligent algorithms.
Abstract: Feature selection plays a vital role in building machine learning models. Irrelevant features in data affect the accuracy of the model and increase the training time needed to build the model. Feature selection is an important process to build Intrusion Detection System (IDS). In this paper, a wrapper feature selection algorithm for IDS is proposed. This algorithm uses the pigeon inspired optimizer to utilize the selection process. A new method to binarize a continuous pigeon inspired optimizer is proposed and compared to the traditional way for binarizing continuous swarm intelligent algorithms. The proposed algorithm was evaluated using three popular datasets: KDDCUP99, NLS-KDD and UNSW-NB15. The proposed algorithm outperformed several feature selection algorithms from state-of-the-art related works in terms of TPR, FPR, accuracy, and F-score. Also, the proposed cosine similarity method for binarizing the algorithm has a faster convergence than the sigmoid method.
TL;DR: This paper focuses only on selection hyper-heuristics and presents critical discussion, current research trends and directions for future research, and the existing classification of selectionhyper- heuristics is extended, in order to reflect the nature of the challenges faced in contemporary research.
TL;DR: It is proven that the fourth measure, called relative neighborhood self-information, is better for feature selection than the other measures, because not only does it consider both the lower and the upper approximations but also the change of its magnitude is largest with the variation of feature subsets.
Abstract: The concept of dependency in a neighborhood rough set model is an important evaluation function for the feature selection. This function considers only the classification information contained in the lower approximation of the decision while ignoring the upper approximation. In this paper, we construct a class of uncertainty measures: decision self-information for the feature selection. These measures take into account the uncertainty information in the lower and the upper approximations. The relationships between these measures and their properties are discussed in detail. It is proven that the fourth measure, called relative neighborhood self-information, is better for feature selection than the other measures, because not only does it consider both the lower and the upper approximations but also the change of its magnitude is largest with the variation of feature subsets. This helps to facilitate the selection of optimal feature subsets. Finally, a greedy algorithm for feature selection has been designed and a series of numerical experiments was carried out to verify the effectiveness of the proposed algorithm. The experimental results show that the proposed algorithm often chooses fewer features and improves the classification accuracy in most cases.
TL;DR: This work conducted a comprehensive analysis of the genomic and phenotypic changes associated with modern maize breeding through chronological sampling of 350 elite inbred lines representing multiple eras of germplasm from both China and the United States to demonstrate the use of the breeding-era approach for identifying breeding signatures.
Abstract: Since the development of single-hybrid maize breeding programs in the first half of the twentieth century1, maize yields have increased over sevenfold, and much of that increase can be attributed to tolerance of increased planting density2-4. To explore the genomic basis underlying the dramatic yield increase in maize, we conducted a comprehensive analysis of the genomic and phenotypic changes associated with modern maize breeding through chronological sampling of 350 elite inbred lines representing multiple eras of germplasm from both China and the United States. We document several convergent phenotypic changes in both countries. Using genome-wide association and selection scan methods, we identify 160 loci underlying adaptive agronomic phenotypes and more than 1,800 genomic regions representing the targets of selection during modern breeding. This work demonstrates the use of the breeding-era approach for identifying breeding signatures and lays the foundation for future genomics-enabled maize breeding.
TL;DR: Re-sequencing of elite cultivars from the historical series of wheat breeding in China demonstrates the impact of " founder genotypes" on the output of breeding efforts over multiple decades, and suggests "founder genotype" perspectives are in fact more dynamic when applied in the context of modern genomics-informed breeding.
TL;DR: The hybrid method resulting from combining the Intuitionistic Fuzzy Set and TOPSIS is very effective to select which supplier is more suitable among the alternatives and also this method can be integrated to similar problems.
Abstract: One of the most important functions of supply chain management is to enhance competitive pressure. Competition conditions and customer perception have changed in favor of environmentalist attitude. Therefore, green supplier selection (GSS) has become an important issue. In this study, the problem of GSS aiming for lean, agile, environmentally sensitive, sustainability, and durability is addressed. The environmental criteria considered in GSS and classical supplier selection are different from each other in terms of carbon footprint, water consumption, environmental applications, and recycling applications. The Technique for Order Preference by Similarity to Ideal Solution (TOPSIS) method has been used in the problem of GSS by considering the multi-criteria decision-making (MCDM) method since MCDM is very effective in many aspects such as evaluating and selecting the classical and environmental criteria. Due to linguistic criteria and no possibility to measure all criteria, it is needed to consolidate the fuzzy approach with the TOPSIS method to reduce the effects of ambiguity and instability. The Intuitionistic Fuzzy TOPSIS method is used because this method makes evaluating decision-makers and criteria convenient. According to the criteria determined by the order of importance, the hybrid method resulting from combining the Intuitionistic Fuzzy Set and TOPSIS is very effective to select which supplier is more suitable among the alternatives and also this method can be integrated to similar problems.
TL;DR: This chapter presents the most fundamental concepts, operators, and mathematical models of this algorithm, which mimics Darwinian theory of survival of the fittest in nature.
Abstract: Genetic Algorithm (GA) is one of the most well-regarded evolutionary algorithms in the history. This algorithm mimics Darwinian theory of survival of the fittest in nature. This chapter presents the most fundamental concepts, operators, and mathematical models of this algorithm. The most popular improvements in the main component of this algorithm (selection, crossover, and mutation) are given too. The chapter also investigates the application of this technique in the field of image processing. In fact, the GA algorithm is employed to reconstruct a binary image from a completely random image.
TL;DR: An introductory-level presentation of its most popular estimator is started, highlighting the latter's temporal dependency, and suggesting how it might be correctly used to inform model selection, and elaborating a simplified framework that may enable an easier interpretation and quantification of C-index improvements or deteriorations.
TL;DR: The results suggest that different types of travellers present differences in hotel key factors, criterion importance and selection results, however, families and friends have similar hotel selection results.
TL;DR: A comprehensive survey on feature selection approaches for clustering is introduced by reflecting the advantages/disadvantages of current approaches from different perspectives and identifying promising trends for future research.
Abstract: The massive growth of data in recent years has led challenges in data mining and machine learning tasks. One of the major challenges is the selection of relevant features from the original set of available features that maximally improves the learning performance over that of the original feature set. This issue attracts researchers’ attention resulting in a variety of successful feature selection approaches in the literature. Although there exist several surveys on unsupervised learning (e.g., clustering), lots of works concerning unsupervised feature selection are missing in these surveys (e.g., evolutionary computation based feature selection for clustering) for identifying the strengths and weakness of those approaches. In this paper, we introduce a comprehensive survey on feature selection approaches for clustering by reflecting the advantages/disadvantages of current approaches from different perspectives and identifying promising trends for future research.
TL;DR: The GABONST algorithm has the capability of producing good quality solutions and it also has better control of the exploitation and exploration as compared to the conventional GA, EATLBO, Bat, and Bee algorithms in terms of the statistical assessment.
Abstract: The metaheuristic genetic algorithm (GA) is based on the natural selection process that falls under the umbrella category of evolutionary algorithms (EA). Genetic algorithms are typically utilized for generating high-quality solutions for search and optimization problems by depending on bio-oriented operators such as selection, crossover, and mutation. However, the GA still suffers from some downsides and needs to be improved so as to attain greater control of exploitation and exploration concerning creating a new population and randomness involvement happening in the population at the solution initialization. Furthermore, the mutation is imposed upon the new chromosomes and hence prevents the achievement of an optimal solution. Therefore, this study presents a new GA that is centered on the natural selection theory and it aims to improve the control of exploitation and exploration. The proposed algorithm is called genetic algorithm based on natural selection theory (GABONST). Two assessments of the GABONST are carried out via (i) application of fifteen renowned benchmark test functions and the comparison of the results with the conventional GA, enhanced ameliorated teaching learning-based optimization (EATLBO), Bat and Bee algorithms. (ii) Apply the GABONST in language identification (LID) through integrating the GABONST with extreme learning machine (ELM) and named (GABONST-ELM). The ELM is considered as one of the most useful learning models for carrying out classifications and regression analysis. The generation of results is carried out grounded upon the LID dataset, which is derived from eight separate languages. The GABONST algorithm has the capability of producing good quality solutions and it also has better control of the exploitation and exploration as compared to the conventional GA, EATLBO, Bat, and Bee algorithms in terms of the statistical assessment. Additionally, the obtained results indicate that (GABONST-ELM)-LID has an effective performance with accuracy reaching up to 99.38%.
TL;DR: A hyperplane assisted evolutionary algorithm, referred here as hpaEA, is proposed which significantly outperforms the compared algorithms on 20 out of 36 benchmark instances and is compared with five state-of-the-art many- objective evolutionary algorithms on 36 many-objective benchmark instances.
Abstract: In many-objective optimization problems (MaOPs), forming sound tradeoffs between convergence and diversity for the environmental selection of evolutionary algorithms is a laborious task. In particular, strengthening the selection pressure of population toward the Pareto-optimal front becomes more challenging, since the proportion of nondominated solutions in the population scales up sharply with the increase of the number of objectives. To address these issues, this paper first defines the nondominated solutions exhibiting evident tendencies toward the Pareto-optimal front as prominent solutions, using the hyperplane formed by their neighboring solutions, to further distinguish among nondominated solutions. Then, a novel environmental selection strategy is proposed with two criteria in mind: 1) if the number of nondominated solutions is larger than the population size, all the prominent solutions are first identified to strengthen the selection pressure. Subsequently, a part of the other nondominated solutions are selected to balance convergence and diversity and 2) otherwise, all the nondominated solutions are selected; then a part of the dominated solutions are selected according to the predefined reference vectors. Moreover, based on the definition of prominent solutions and the new selection strategy, we propose a hyperplane assisted evolutionary algorithm, referred here as hpaEA , for solving MaOPs. To demonstrate the performance of hpaEA , extensive experiments are conducted to compare it with five state-of-the-art many-objective evolutionary algorithms on 36 many-objective benchmark instances. The experimental results show the superiority of hpaEA which significantly outperforms the compared algorithms on 20 out of 36 benchmark instances.
TL;DR: Experimental analysis shows that the ACO-FCP ensemble model is superior and more robust than its counterparts, and this study strongly recommends that the proposed ACO -FCP model is highly competitive than traditional and other artificial intelligence techniques.
TL;DR: In this article, the authors propose a new network resampling strategy based on splitting node pairs rather than nodes, which is applicable to cross-validation for a wide range of network model selection tasks.
Abstract: Summary While many statistical models and methods are now available for network analysis, resampling of network data remains a challenging problem. Cross-validation is a useful general tool for model selection and parameter tuning, but it is not directly applicable to networks since splitting network nodes into groups requires deleting edges and destroys some of the network structure. In this paper we propose a new network resampling strategy, based on splitting node pairs rather than nodes, that is applicable to cross-validation for a wide range of network model selection tasks. We provide theoretical justification for our method in a general setting and examples of how the method can be used in specific network model selection and parameter tuning tasks. Numerical results on simulated networks and on a statisticians’ citation network show that the proposed cross-validation approach works well for model selection.
TL;DR: This study proposes a new integrated MCDM model combining fuzzy SWARA and CoCoSo methods to select an optimal location for a logistics center in Sivas province, Turkey, considering multiple criteria and evaluating the accuracy of CoCoSo results against other MCDM methods.
Abstract: Logistics centers are home to many and varied facilities, such as storage, transportation of goods, handling, reassembling, clearing, disassembling, quality control, social services and providing accommodation, so on. Providing logistical activities from one location can provide some macro advantages, as well as regional development in developing countries. For the micro level, logistics center selection has an effective role in increasing the operational efficiency and decreasing the costs of the firms. While the wrong location selection for logistics center affects the operations and costs of the companies negatively, the optimal location selection increases the performance, competitiveness, profitability of the firms and reduces the costs of the firms. Since many different qualitative and quantitative criteria are considered in the selection of the logistics center, this selection problem is an MCDM problem. A new integrated MCDM model is proposed to solve this problem for Sivas province in Turkey. This study presents two contributions to the literature. Firstly, the number of studies related to CoCoSo method is limited in the literature, therefore, the CoCoSo method is proposed in this study. Secondly, a new integrated GIS-based MCDM model comprising fuzzy SWARA and CoCoSo is introduced to literature to address the location selection problem for a logistics center. In this study, the results of CoCoSo method and the resulfts of other MCDM methods (COPRAS, VIKOR, ARAS, MOORA, and MABAC) are compared to test the accuracy of results obtained by CoCoSo. Besides, the criteria weights are changed and the possible changes in the results are tracked.
TL;DR: A new ABC is proposed, in which a new selection method based on neighborhood radius is used, and unlike the probability selection in the original ABC, NSABC chooses the best solution in the neighborhood radius to generate offspring.
TL;DR: A practical structure of LIS-based spatial modulation (LIS-SM) is proposed, in order to utilize both transmit and receive antenna indices, and a low-complexity selection algorithm is designed on the basis of minimum squared Euclidian distance and signal-to-leakage-and-noise ratio.
Abstract: Novel communication technology based on large intelligent surface (LIS) [1] has arisen recently, with the aim to enhance the signal quality at the receiver In this paper, a practical structure of LIS-based spatial modulation (LIS-SM) is proposed, in order to utilize both transmit and receive antenna indices Meanwhile, the theoretical average bit error rate (ABER) performance bound of the developed LIS-SM scheme is investigated For the sake of achieving further spatial diversity gain, we extend its employment to the antenna selection (AS) scenario, and a low-complexity selection algorithm is designed on the basis of minimum squared Euclidian distance and signal-to-leakage-and-noise ratio as well as the idea of greedy elimination algorithm Performance analysis shows that AS-aided LIS-SM is more robust in terms of ABER compared with conventional LIS-SM Moreover, complexity analysis also depicts that the proposed fast selection algorithm achieves much lower complexity yet a comparable ABER performance, compared to the traditional exhaustive search
TL;DR: It is shown that branching events in a seagrass clone or genet lead to population bottlenecks of tissue that result in the evolution of genetically differentiated ramets in a process of somatic genetic drift.
Abstract: All multicellular organisms are genetic mosaics owing to somatic mutations. The accumulation of somatic genetic variation in clonal species undergoing asexual (or clonal) reproduction may lead to phenotypic heterogeneity among autonomous modules (termed ramets). However, the abundance and dynamics of somatic genetic variation under clonal reproduction remain poorly understood. Here we show that branching events in a seagrass (Zostera marina) clone or genet lead to population bottlenecks of tissue that result in the evolution of genetically differentiated ramets in a process of somatic genetic drift. By studying inter-ramet somatic genetic variation, we uncovered thousands of single nucleotide polymorphisms that segregated among ramets. Ultra-deep resequencing of single ramets revealed that the strength of purifying selection on mosaic genetic variation was greater within than among ramets. Our study provides evidence for multiple levels of selection during the evolution of seagrass genets. Somatic genetic drift during clonal propagation leads to the emergence of genetically unique modules that constitute an elementary level of selection and individuality in long-lived clonal species. The accumulation of somatic genetic variation in clonal species leads to heterogeneity among autonomous modules (ramets). Ultra-deep resequencing of single ramets in a clonal seagrass shows somatic genetic drift resulting in genetically differentiated ramets that are targets of selection.
TL;DR: Large data computations in ssGBLUP were solved by exploiting limited dimensionality of genomic data due to limited effective population size, and involves new validation procedures that are unaffected by selection, parameter estimation that accounts for all the genomic data used in selection, and strategies to address reduction in genetic variances after genomic selection was implemented.
Abstract: Early application of genomic selection relied on SNP estimation with phenotypes or de-regressed proofs (DRP). Chips of 50k SNP seemed sufficient for an accurate estimation of SNP effects. Genomic estimated breeding values (GEBV) were composed of an index with parent average, direct genomic value, and deduction of a parental index to eliminate double counting. Use of SNP selection or weighting increased accuracy with small data sets but had minimal to no impact with large data sets. Efforts to include potentially causative SNP derived from sequence data or high-density chips showed limited or no gain in accuracy. After the implementation of genomic selection, EBV by BLUP became biased because of genomic preselection and DRP computed based on EBV required adjustments, and the creation of DRP for females is hard and subject to double counting. Genomic selection was greatly simplified by single-step genomic BLUP (ssGBLUP). This method based on combining genomic and pedigree relationships automatically creates an index with all sources of information, can use any combination of male and female genotypes, and accounts for preselection. To avoid biases, especially under strong selection, ssGBLUP requires that pedigree and genomic relationships are compatible. Because the inversion of the genomic relationship matrix (G) becomes costly with more than 100k genotyped animals, large data computations in ssGBLUP were solved by exploiting limited dimensionality of genomic data due to limited effective population size. With such dimensionality ranging from 4k in chickens to about 15k in cattle, the inverse of G can be created directly (e.g., by the algorithm for proven and young) at a linear cost. Due to its simplicity and accuracy, ssGBLUP is routinely used for genomic selection by the major chicken, pig, and beef industries. Single step can be used to derive SNP effects for indirect prediction and for genome-wide association studies, including computations of the P-values. Alternative single-step formulations exist that use SNP effects for genotyped or for all animals. Although genomics is the new standard in breeding and genetics, there are still some problems that need to be solved. This involves new validation procedures that are unaffected by selection, parameter estimation that accounts for all the genomic data used in selection, and strategies to address reduction in genetic variances after genomic selection was implemented.
TL;DR: P phenotypic selection analysis is used to estimate the type and strength of selection that acts on more than 15,000 transcripts in rice ( Oryza sativa), which provides insight into the adaptive evolutionary role of selection on gene expression.
Abstract: Levels of gene expression underpin organismal phenotypes1,2, but the nature of selection that acts on gene expression and its role in adaptive evolution remain unknown1,2. Here we assayed gene expression in rice (Oryza sativa)3, and used phenotypic selection analysis to estimate the type and strength of selection on the levels of more than 15,000 transcripts4,5. Variation in most transcripts appears (nearly) neutral or under very weak stabilizing selection in wet paddy conditions (with median standardized selection differentials near zero), but selection is stronger under drought conditions. Overall, more transcripts are conditionally neutral (2.83%) than are antagonistically pleiotropic6 (0.04%), and transcripts that display lower levels of expression and stochastic noise7–9 and higher levels of plasticity9 are under stronger selection. Selection strength was further weakly negatively associated with levels of cis-regulation and network connectivity9. Our multivariate analysis suggests that selection acts on the expression of photosynthesis genes4,5, but that the efficacy of selection is genetically constrained under drought conditions10. Drought selected for earlier flowering11,12 and a higher expression of OsMADS18 (Os07g0605200), which encodes a MADS-box transcription factor and is a known regulator of early flowering13—marking this gene as a drought-escape gene11,12. The ability to estimate selection strengths provides insights into how selection can shape molecular traits at the core of gene action. Phenotypic selection analysis is used to estimate the type and strength of selection that acts on more than 15,000 transcripts in rice (Oryza sativa), which provides insight into the adaptive evolutionary role of selection on gene expression.
TL;DR: An integrated framework is proposed for HGR using deep neural network and Fuzzy Entropy controlled Skewness (FEcS) approach and the obtained overall recognition results lead to conclude that the proposed system is very promising.
Abstract: Human gait recognition (HGR) shows high importance in the area of video surveillance due to remote access and security threats. HGR is a technique commonly used for the identification of human style in daily life. However, many typical situations like change of clothes condition and variation in view angles degrade the system performance. Lately, different machine learning (ML) techniques have been introduced for video surveillance which gives promising results among which deep learning (DL) shows best performance in complex scenarios. In this article, an integrated framework is proposed for HGR using deep neural network and Fuzzy Entropy controlled Skewness (FEcS) approach. The proposed technique works in two phases: In the first phase, Deep Convolutional Neural Network (DCNN) features are extracted by pre-trained CNN models (VGG19 and AlexNet) and their information is mixed by parallel fusion approach. In the second phase, entropy and skewness vectors are calculated from fused feature vector (FV) to select best subsets of features by suggested FEcS approach. The best subsets of picked features are finally fed to multiple classifiers and finest one is chosen on the basis of accuracy value. The experiments were done on four well-known datasets namely AVAMVG gait, CASIA A, B and C. The achieved accuracy of each dataset was 99.8%, 99.7%, 93.3% and 92.2%, respectively. Therefore, the obtained overall recognition results lead to conclude that the proposed system is very promising.
TL;DR: This Review discusses how genomic technologies are providing a deeper understanding of colour traits, revealing fresh insights into their genetic architecture, evolvability and origins of adaptive variation.
Abstract: Coloration is an easily quantifiable visual trait that has proven to be a highly tractable system for genetic analysis and for studying adaptive evolution. The application of genomic approaches to evolutionary studies of coloration is providing new insight into the genetic architectures underlying colour traits, including the importance of large-effect mutations and supergenes, the role of development in shaping genetic variation and the origins of adaptive variation, which often involves adaptive introgression. Improved knowledge of the genetic basis of traits can facilitate field studies of natural selection and sexual selection, making it possible for strong selection and its influence on the genome to be demonstrated in wild populations.
TL;DR: In this paper, an angle-based selection strategy and a shift-based density estimation strategy are employed in the environmental selection to delete poor individuals one by one, and the results suggest that AnD can achieve highly competitive performance.
TL;DR: A novel band selection method called optimal neighborhood reconstruction (ONR), which exploits a recurrence relation that underlies the optimization target to obtain the optimal solution in an efficient way, is proposed and compared with state-of-the-art methods.
Abstract: Band selection is one of the most important technique in the reduction of hyperspectral image (HSI). Different from traditional feature selection problem, an important characteristic of it is that there is usually strong correlation between neighboring bands, that is, bands with close indexes. Aiming to fully exploit this prior information, a novel band selection method called optimal neighborhood reconstruction (ONR) is proposed. In ONR, band selection is considered as a combinatorial optimization problem. It evaluates a band combination by assessing its ability to reconstruct the original data, and applies a noise reducer to minimize the influence of noisy bands. Instead of using some approximate algorithms, ONR exploits a recurrence relation that underlies the optimization target to obtain the optimal solution in an efficient way. Besides, we develop a parameter selection approach to automatically determine the parameter of ONR, ensuring it is adaptable to different data sets. In experiments, ONR is compared with some state-of-the-art methods on six HSI data sets. The results demonstrate that ONR is more effective and robust than the others in most of the cases.
TL;DR: Three fundamental pillars for future breeding strategies in the framework of Green Systems Biology are proposed, combining genome selection with environment‐dependent PANOMICS analysis and deep learning to improve prediction accuracy for marker‐dependent trait performance and combining genome editing and speed breeding tools to accelerate and enhance large‐scale functional validation of trait‐specific precision breeding.
Abstract: Genotyping-by-sequencing has enabled approaches for genomic selection to improve yield, stress resistance and nutritional value. More and more resource studies are emerging providing 1000 and more genotypes and millions of SNPs for one species covering a hitherto inaccessible
intraspecific genetic variation. The larger the databases are growing, the better statistical approaches for genomic selection will be available. However, there are clear limitations on the statistical but also on the biological part. Intraspecific genetic variation is able to explain a high proportion of the phenotypes, but a large part of phenotypic plasticity also stems from environmentally driven transcriptional, post-transcriptional, ranslational, post-translational, epigenetic and metabolic regulation. Moreover, regulation of the same gene can have different
phenotypic outputs in different environments. Consequently, to explain and understand environment-dependent phenotypic plasticity based on the available genotype variation we have
to integrate the analysis of further molecular levels reflecting the complete information flow from the gene to metabolism to phenotype. Interestingly, metabolomics platforms are already more cost-effective than NGS platforms and are decisive for the prediction of nutritional value or stress resistance. Here, we propose three fundamental pillars for future breeding strategies in the framework of Green Systems Biology: (i) combining genome selection with environment dependent
PANOMICS analysis and deep learning to improve prediction accuracy for marker dependent trait performance; (ii) PANOMICS resolution at subtissue, cellular and subcellular level provides information about fundamental functions of selected markers; (iii) combining PANOMICS with genome editing and speed breeding tools to accelerate and enhance large-scale functional validation of trait-specific precision breeding.
TL;DR: A comprehensive overview of the general concept, various methodologies, and bioinformatics tools currently available for the detection of selective sweeps and the results of recent selection signature studies carried out in various livestock species are given.
TL;DR: An integrated methodology including the Intuitionistic Fuzzy Technique for Order Preference by Similarity to Ideal Solution (IF-TOPSIS) and a modified two-phase fuzzy goal programming model are proposed to better address this selection problem in a multi-item/multi-supplier/ multi-period environment.