TL;DR: The advances of the human genome project and the completion of total genome sequences for yeast and many bacterial species, have enabled investigators to view genetic information in the context of the entire genome and recognize that the mechanisms for some genetic diseases are best understood at a genomic level.
TL;DR: In this paper, the authors used DNA flow cytometry to determine the genome size and GC-percent of 154 vertebrate species and found that the overall distribution of points is not linear but triangular, whereas with an increase in genome size the lower limit for GC- percent is elevated, gradually approaching the upper limit.
Abstract: Genome size and GC-percent were determined by means of a special method of DNA flow cytometry in 154 vertebrate species. For the total dataset, a highly significant positive correlation was found between both parameters. The overall distribution of points is not linear but triangular: a wide range of GC-percent values is observed at the lower end of genome size range, whereas with an increase in genome size the lower limit for GC-percent is elevated, gradually approaching the upper limit (about 46%). In teleost fishes, which occupy the lower part of genome size range, the negative relationship between both parameters was observed. Two positive linear relationships were found between mean genome size and GC-percent of the main vertebrate groups (one includes fishes, amphibians, and mammals, the other consists of reptiles and birds, which show the higher GC-percent for their genome sizes). Distribution of variance between taxonomic levels indicates that GC-percent is more evolutionarily conservative than genome size in anamniotes. Anuran amphibians show the greatest part of genome size variability at the lower taxonomic levels as compared to other vertebrates (with no additional variance already above the genus level). The data obtained with different methods are compared. It is shown that the proposed method can provide useful data for studies on genome evolution and biodiversity.
TL;DR: It is shown that genes on the Borrelia burgdorferi genome have two separate, distinct, and significantly different codon usages, depending on whether the gene is transcribed on the leading or lagging strand of replication, a new paradigm of codon selection in prokaryotes.
Abstract: With more than 10 fully sequenced, publicly available prokaryotic genomes, it is now becoming possible to gain useful insights into genome evolution. Before the genome era, many evolutionary processes were evaluated from limited data sets and evolutionary models were constructed on the basis of small amounts of evidence. In this paper, I show that genes on the Borrelia burgdorferi genome have two separate, distinct, and significantly different codon usages, depending on whether the gene is transcribed on the leading or lagging strand of replication. Asymmetrical replication is the major source of codon usage variation. Replicational selection is responsible for the higher number of genes on the leading strands, and transcriptional selection appears to be responsible for the enrichment of highly expressed genes on these strands. Replicational–transcriptional selection, therefore, has an influence on the codon usage of a gene. This is a new paradigm of codon selection in prokaryotes.
TL;DR: Using molecular and cytological approaches, a range of differentially organized repetitive DNA sequence elements from the genomes of cultivated and wild beet species are characterized, leading to an extensive model of the repetitive DNA, its organization and evolution.
TL;DR: It was shown that cognates (orthologs) of human duplicated genes can be found in other vertebrates, including bony fishes, and that large-scale duplications occurred after the echinoderms/chordates split and before the bony vertebrate radiation.
Abstract: Paralogous genes from several families were found in four human chromosome regions (4p16, 5q33-35, 8p12-21, and 10q24-26), suggesting that their common ancestral region underwent several rounds of large-scale duplication. Searches in the EMBL databases, followed by phylogenetic analyses, showed that cognates (orthologs) of human duplicated genes can be found in other vertebrates, including bony fishes. In contrast, within each family, only one gene showing the same high degree of similarity with all the duplicated mammalian genes was found in nonvertebrates (echinoderms, insects, nematodes). This indicates that large-scale duplications occurred after the echinoderms/chordates split and before the bony vertebrate radiation. It has been suggested that two rounds of gene duplication occurred in the vertebrate lineage after the separation of Amphioxus and craniate (vertebrates + Myxini) ancestors. Before these duplications, the genes that have led to the families of paralogous genes in vertebrates must have been physically linked in the craniate ancestor. Linkage of some of these genes can be found in the Drosophila melanogaster and Caenorhabditis elegans genomes, suggesting that they were linked in the triploblast Metazoa ancestor.
TL;DR: The results suggest that plant mitochondria originate from the same ancestor as other mitochondria and that most genes were lost from the mitochondrial genome at a fairly early stage of the evolution of the plants.
Abstract: The complete nucleotide sequence of the mitochondrial genome of a very primitive unicellular red alga, Cyanidioschyzon merolae , has been determined The mitochondrial genome of Cmerolae contains 34 genes for proteins including unidentified open reading frames (ORFs) (three subunits of cytochrome c oxidase, apocytochrome b protein, three subunits of F1F0-ATPase, seven subunits of NADH ubiquinone oxidoreductase, three subunits of succinate dehydrogenase, four proteins implicated in c-type cytochrome biogenesis, 11 ribosomal subunits and two unidentified open reading frames), three genes for rRNAs and 25 genes for tRNAs The G+C content of this mitochondrial genome is 272% The genes are encoded on both strands The genome size is comparatively small for a plant mitochondrial genome (32 211 bp) The mitochondrial genome resembles those of plants in its gene content because it contains several ribosomal protein genes and ORFs shared by other plant mitochondrial genomes In contrast, it resembles those of animals in the genome organization, because it has very short intergenic regions and no introns The gene set in this mitochondrial genome is a subset of that of Reclinomonas americana , an amoeboid protozoan The results suggest that plant mitochondria originate from the same ancestor as other mitochondria and that most genes were lost from the mitochondrial genome at a fairly early stage of the evolution of the plants
TL;DR: Several eukaryotes, including maize, yeast and Xenopus, are degenerate polyploids formed by relatively recent whole-genome duplications, suggesting that more ancient genome duplications occurred in an ancestor of vertebrates.
TL;DR: The study infers the frequency of functional divergence from the size distribution of gene families produced by two successive genome duplications early in vertebrate evolution and reasons for this unexpectedly high frequency are discussed.
Abstract: Gene duplication events are important sources of novel gene functions. However, more often than not, a duplicate gene may lose its function and become a pseudogene. What is the relative frequency of these two scenarios: functional divergence versus gene loss? Given that most non-neutral mutations are deleterious, gene loss should be far more frequent than divergence. However, a recent empirical study suggests that about 50% of all gene duplications will lead to functional divergence. The study infers the frequency of functional divergence from the size distribution of gene families produced by two successive genome duplications early in vertebrate evolution. Reasons for this unexpectedly high frequency of functional divergence are discussed.
TL;DR: Part One: Genome Structure, Stability, and Evolution; Part Two: Strategies for Genome Analysis; Part Three: Physical Maps of Bacteria and Their Methods for Construction.
Abstract: Part One: Genome Structure, Stability, and Evolution. Genome structure. Genome stability. Genome evolution. Part Two: Strategies for Genome Analysis. Physical mapping strategies. Genomic fingerprinting methods. Genomic sequencing projects. Part Three: Physical Maps of Bacteria and Their Methods for Construction. Index.
TL;DR: This work focuses on the potential of transposons for tagging and identifying fungal genes, and presents structural and functional features for such transposon that have been identified so far in filamentous fungi.
Abstract: Transposons are ubiquitous genetic elements discovered so far in all investigated prokaryotes and eukaryotes. In remarkable contrast to all other genes, transposable elements are able to move to new locations within their host genomes. Transposition of transposons into coding sequences and their initiation of chromosome rearrangements have tremendous impact on gene expression and genome evolution. While transposons have long been known in bacteria, plants, and animals, only in recent years has there been a significant increase in the number of transposable elements discovered in filamentous fungi. Like those of other eukaryotes, each fungal transposable element is either of class or of class II. While class I elements transpose by a RNA intermediate and employ reverse transcriptases, class II elements transpose directly at the DNA level. We present structural and functional features for such transposons that have been identified so far in filamentous fungi. Emphasis is given to specific advantages or unique features when fungal systems are used to study transposable elements, e.g., the evolutionary impact of transposons in coenocytic organisms and possible experimental approaches toward horizontal gene transfer. Finally, we focus on the potential of transposons for tagging and identifying fungal genes.
TL;DR: The powerful combination of genomics and bioinformatics is providing a wealth of information about Mycobacterium tuberculosis, the aetiological agent of human tuberculosis, that will facilitate the conception and development of new therapies.
Abstract: The powerful combination of genomics and bioinformatics is providing a wealth of information about Mycobacterium tuberculosis, the aetiological agent of human tuberculosis, that will facilitate the conception and development of new therapies. The starting point for genome sequencing was the integrated map of the 4.4 Mb circular chromosome of the widely used, virulent reference strain, M. tuberculosis H37Rv. Cosmids and bacterial artificial chromosomes were selected from ordered libraries and subjected to systematic shotgun sequence analysis. This approach simplified sequence assembly as the genome is rich in repetitive DNA. In common with most bacteria, > 90% of the potential coding capacity is used, and probable or tentative functions could be attributed to > 70% of the genes. The potential biological roles of two of the principal driving forces in genome dynamics, insertion sequence elements and polymorphic multigene families are discussed.
TL;DR: It is shown that, as the repeat array of a microsatellite grows, its mutation rate increases and spectrum of mutations changes; these effects have important consequences for the evolution of microsatellites in genomes.
TL;DR: Known examples of genomic duplications present on the human X chromosome and autosomes are reviewed.
Abstract: As large-scale sequencing accumulates momentum, an increasing number of instances are being revealed in which genes or other relatively rare sequences are duplicated, either in tandem or at nearby locations. Such duplications are a source of considerable polymorphism in populations, and also increase the evolutionary possibilities for the coregulation of juxtaposed sequences. As a further consequence, they promote inversions and deletions that are responsible for significant inherited pathology. Here we review known examples of genomic duplications present on the human X chromosome and autosomes.
TL;DR: Genome organisation and phylogenetic ancestor-descendent relationships between extant bacteria of closely related genera and within the same monophyletic genus and species suggest that some strains have undergone transition from two chromosomes to a single replicon.
Abstract: Animal intracellular Proteobacteria of the alpha subclass without plasmids and containing one or more chromosomes are phylogenetically entwined with opportunistic, plant-associated, chemoautotrophic and photosynthetic alpha Proteobacteria possessing one or more chromosomes and plasmids. Local variations in open environments, such as soil, water, manure, gut systems and the external surfaces of plants and animals, may have selected alpha Proteobacteria with extensive metabolic alternatives, broad genetic diversity, and more flexible and larger genomes with ability for horizontal gene flux. On the contrary, the constant and isolated animal cellular milieu selected heterotrophic alpha Proteobacteria with smaller genomes without plasmids and reduced genetic diversity as compared to their plant-associated and phototrophic relatives. The characteristics and genome sizes in the extant species suggest that a second chromosome could have evolved from megaplasmids which acquired housekeeping genes. Consequently, the genomes of the animal cell-associated Proteobacteria evolved through reductions of the larger genomes of chemoautotrophic ancestors and became rich in adenosine and thymidine, as compared to the genomes of their ancestors. Genome organisation and phylogenetic ancestor–descendent relationships between extant bacteria of closely related genera and within the same monophyletic genus and species suggest that some strains have undergone transition from two chromosomes to a single replicon. It is proposed that as long as the essential information is correctly expressed, the presence of one or more chromosomes within the same genus or species is the result of contingency. Genetic drift in clonal bacteria, such as animal cell-associated alpha Proteobacteria, would depend almost exclusively on mutation and internal genetic rearrangement processes. Alternatively, genomic variations in reticulate bacteria, such as many intestinal and plant cell-associated Proteobacteria, will depend not only on these processes, but also on their genetic interactions with other bacterial strains. Common pathogenic domains necessary for the invasion and survival in association with cells have been preserved in the chromosomes of the animal and plant-associated alpha Proteobacteria. These pathogenic domains have been maintained by vertical inherence, extensively ameliorated to match the chromosome G+C content and evolved within chromosomes of alpha Proteobacteria.
TL;DR: The physical sites of 18S-5.8S-25S and 5S rRNA genes and telomeric sequences in the MusaL genome were localized by fluorescentin situhybridization on mitotic chromosomes of selected lines indicating variation in the number of copies.
Abstract: The physical sites of 18S-5.8S-25S and 5S rRNA genes and telomeric sequences in theMusaL. genome were localized by fluorescentin situhybridization on mitotic chromosomes of selected lines. A single major intercalary site of the 18S-5.8S-25S rDNA was observed on the short arm of the nucleolar organizing chromosome in each genome. AA and BB genome diploids had a single pair of sites, triploids had three sites while a tetraploid hybrid had four sites. The probe is useful for quick determination of ploidy, even using interphase nuclei from slowly growing tissue culture material. Variation in the intensity of signals was observed among heterogeneousMusalines indicating variation in the number of copies of the 18S-5.8S-25S rRNA genes. Eight subterminal sites of 5S rDNA were observed in Calcutta 4 (AA) while Butohan 2 (BB) had six sites; some were weaker in both genotypes. Triploid lines showed six to nine major sites of 5S rDNA of widely varying intensity and near the limit of detection. The diploid hybrids had five to nine sites of 5S rDNA while the tetraploid hybrid had 11 sites. The telomeric sequence was detected as pairs of dots at the ends of all the chromosomes analysed but no intercalary sequences were seen. The molecular cytogenetic studies ofMusausing repetitive and single copy DNA probes should yield insight into the genome and its evolution and provide data forMusabreeders, as well as generating genetic markers inMusa.
TL;DR: It is shown how genomic constraints might decide the pattern of distribution of excess DNA within chromosome complements and predetermine future evolutionary patterns in karyotype organization.
TL;DR: By analysing retrotransposon end-sequences, Phillip SanMiguel and colleagues have determined that an eruption of transposon activity over the past six million years has led to the plethora oftransposons that currently litter the maize genome.
Abstract: The sizes of plant genomes vary remarkably the well-studied Arabidopsis genome is 108 base pairs while those of some lilies are over one thousand times this size1, 2. Genome sizes vary even among closely related plants. The plum genome, for example, is three times larger than that of peaches (both are members of the genus Prunus ), suggesting that fluctuations in genome size occur over relatively short periods of evolutionary time. While transposon insertion is recognized as a force underlying genome fluidity, the pace at which it contributes to genome evolution has remained obscure. On page 43, Phillip SanMiguel and colleagues document the rapidity by which transposons can restructure genomes; by analysing retrotransposon end-sequences, they have determined that an eruption of transposon activity over the past six million years has led to the plethora of transposons that currently litter the maize genome3.
TL;DR: Part of the observed variability in genome organization may be explained by the presence or absence, in a given strain, of dispensable genomic regions and/or chromosomes.
Abstract: The genome structure of Colletotrichum lindemuthianum in a set of diverse isolates was investigated using a combination of physical and molecular approaches. Flow cytometric measurement of genome size revealed significant variation between strains, with the smallest genome representing 59% of the largest. Southern-blot profiles of a cloned fungal telomere revealed a total chromosome number varying from 9 to 12. Chromosome separations using pulsed-field gel electrophoresis (PFGE) showed that these chromosomes belong to two distinct size classes: a variable number of small (< 2.5 Mb) polymorphic chromosomes and a set of unresolved chromosomes larger than 7 Mb. Two dispersed repeat elements were shown to cluster on distinct polymorphic minichromosomes. Single-copy flanking sequences from these repeat-containing clones specifically marked distinct small chromosomes. These markers were absent in some strains, indicating that part of the observed variability in genome organization may be explained by the presence or absence, in a given strain, of dispensable genomic regions and/or chromosomes.
TL;DR: Genetic analysis indicates that this function is independent of UV-damage repair and mutation avoidance, supporting the notion that RAD3 and SSL1 together play a novel role in the maintenance of genome integrity.
Abstract: Maintaining genome stability requires that recombination between repetitive sequences be avoided. Because short, repetitive sequences are the most abundant, recombination between sequences that are below a certain length are selectively restricted. Novel alleles of the RAD3 and SSL1 genes, which code for components of a basal transcription and UV-damage-repair complex in Saccharomyces cerevisiae, have been found to stimulate recombination between short, repeated sequences. In double mutants, these effects are suppressed, indicating that the RAD3 and SSL1 gene products work together in influencing genome stability. Genetic analysis indicates that this function is independent of UV-damage repair and mutation avoidance, supporting the notion that RAD3 and SSL1 together play a novel role in the maintenance of genome integrity.
TL;DR: Because certain types of transposable elements are embedded in regulatory regions of plant genes and have become greatly amplified in plant genomes, they could contribute substantially to normal gene expression and to the generation of genomic methylation patterns.
Abstract: Transgenes often become silenced in plants because of repressive influences exerted by flanking plant DNA and/or because of interactions among multiple copies of closely linked transgenes. Repeated transgenes on different chromosomes can also interact in a way that leads to silencing and methylation, suggesting a previously unrecognized ability of unlinked homologous sequences to cross-talk in complex genomes. Non-Mendelian inheritance is a frequent consequence of these interactions because the silenced genes do not fully reactivate or lose methylation after segregating in progeny. Several examples of gene silencing in plants appear to reflect the action of genome defence system that methylates and inactivates foreign or invasive sequences such as transgenes and transposable elements. Because certain types of transposable elements are embedded in regulatory regions of plant genes and have become greatly amplified in plant genomes, they could contribute substantially to normal gene expression and to the generation of genomic methylation patterns. Polyploidy, which has been a major force in plant and vertebrate evolution, might encourage proliferation of transposable elements because genes in polyploids are duplicated and hence less susceptible to the consequences of insertional mutagenesis. Accordingly, the appearance of genome-wide methylation has often coincided with episodes of polyploidization.
TL;DR: It is demonstrated that by using pulsed-field gel electrophoresis it is possible to identify ancestrally related chromosome segments in a complex and duplicated genome, such as the genome of B. nigra, permitting one to draw conclusions as to its origin and evolution.
Abstract: Genetic and physical maps, consisting of a large number of DNA markers for Arabidopsis thaliana chromosomes, represent excellent tools to determine the organization of related genomes such as those of Brassica. In this paper we report the chromosomal localization and physical analysis by pulsed-field gel electrophoresis (PFGE) of a well-defined gene complex of A. thaliana in the Brassica nigra genome (B genome n=8). This complex is approximately 30 kb in length in A. thaliana and contains a cluster of six genes including ABI1 (ABA-responsive), RPS2 (resistance against Pseudomonas syringae, a bacterial disease), CK1 (casein kinase I), NAP (nucleosome-assembly protein), X9 and X14 (both of unknown function). The Arabidopsis chromosomal complex was found to be duplicated and conserved in gene number at different levels in the Brassica genome. Linkage group B1 had the most-conserved arrangement carrying all six genes tightly linked. Group B4 had an almost complete complex except for the absence of RPS2. Other partial complexes of fewer members were found on three other chromosomes. Our studies demonstrate that by this approach it is possible to identify ancestrally related chromosome segments in a complex and duplicated genome, such as the genome of B. nigra, permitting one to draw conclusions as to its origin and evolution.
TL;DR: The localization pattern of the repeats on the vole chromosomes confirms the independent origin of the two repeats and suggests that expansion of the heterochromatic blocks has occurred subsequent to speciation.
Abstract: We have characterized two novel, complex, heterochromatic repeat sequences, MS3 and MS4, isolated from Microtus rossiaemeridionalis genomic DNA Sequence analysis indicates that both repeats consist of unique sequences interrupted by repeat elements of different origin and can be classified as long complex repeat units (LCRUs) A unique feature of both repeat units is the presence of short interspersed repeat elements (SINEs), which are usually characteristic of the euchromatic part of the genome Comparative analysis revealed no significant stretches of homology in the nucleotide sequences between the two repeats, suggesting that the repeats originated independently during the course of vole genome evolution Fluorescence in situ hybridization analysis demonstrates that MS3 and MS4 occupy distinct domains in the heterochromatic regions of the sex chromosomes in M transcaspicus and M arvalis but collocalize in M rossiaemeridionalis and M kirgisorum heterochromatic blocks The localization pattern of the repeats on the vole chromosomes confirms the independent origin of the two repeats and suggests that expansion of the heterochromatic blocks has occurred subsequent to speciation
TL;DR: Recent progress in the compilation and effective application of gene mapping data has been particularly evident in the fascinating group of species that comprise the mammalian infraclass Metatheria, better known as marsupial mammals.
Abstract: Increased interest in the structural characteristics of mammalian genomes, together with rapid advances in mapping technology, have led to the explosive expansion of gene mapping activity in recent years. Species for which gene mapping data were virtually nonexistent just a decade ago now possess substantial and serviceable gene maps that are enabling the localization of economically and biomedically relevant loci and are contributing to a fuller understanding of genome evolution and the relationships between genome structure and gene function. Recent progress in the compilation and effective application of gene mapping data has been particularly evident in the fascinating group of species that comprise the mammalian infraclass Metatheria, better known as marsupial mammals. As late as 1988, only 22 marsupial genes had been reported as mapped by any method in any species (a gene is considered mapped if it adheres to any of the following criteria: (1) assigned to a specific chromosome; (2) a member of a linkage group; (3) autosomal in marsupials but known to be X or Y linked in eutherians). These few gene assignments were achieved by a variety of methods, and were scattered piecemeal among a dozen different species (this number excludes 14 additional species in which 1 or more ribosomal RNA (RNR) genes were the only loci mapped [Hayman and Rofe 1977; Hayman and Sharp 1981; Young and others 1982]). As a consequence, few species had more than 2 or 3 genes mapped, and only 1 species could boast as many as 9 gene assignments. Currently, at least 142 loci have been assigned to physical locations or linkage groups in marsupials, and more than 15 species have at least 3 gene assignments. Most important, 2 distantly related species, Macropus eugenii (tammar wallaby) with 70 loci mapped and Monodelphis domestica (gray, short-tailed opossum) with 69
TL;DR: This introduction discusses the difiiculties inherent in edit-distance formulations of multiple rearrangement, referring to relevant work, and argues for a potentially simpler approach based on “breakpoint analysis”.
Abstract: Multiple alignment of macromolecular sequences, an important topic of algorithmic research for at least 25 years 113, lo], generalizes the compsrison of just two sequences which have diverged through the local processes of insertion, deletion and substitution. Recently there has been much interest in gene-order sequences which diverge through non-local genome rearrangement processes such as inversion (or reversal) and transposition (reviewed in [15] ch.7, [9], [S] ch. 19 and 131). What would be the analog of multiple alignment under these models of divergence? In this introduction we tit review some formulations of multiple alignment and show which have counterparts in multiple rearrangement. We then discuss the difiiculties inherent in edit-distance formulations of multiple rearrangement, referring to relevant work, and argue for a potentially simpler approach based on “breakpoint analysis”.
TL;DR: The year 1997 saw the publication of the complete nucleotide sequence of Helicobacter pylori and Escherichia coli, and the high proportion of open reading frames that have no known function.
TL;DR: The level of ribosomal RNA gene sequence divergence in both mitochondrial and chloroplast genomes is higher in the Chlamydomonas lineage than in land plants and is most likely due to higher rates of nucleotide substitution in ChlamYDomonas organellar DNAs.
Abstract: Chlamydomonas mitochondrial and chloroplast genomes, in contrast to the land plant counterparts, exhibit concerted modes and tempos of evolution. The 1.5-fold variation currently observed in the size of both organelle genomes is mostly accounted for by changes in the spacer DNA and intron number, with less contribution from changes in gene content and amount of repeated DNA. Gene order is highly variable in both mitochondrial and chloroplast genomes of Chlamydomonas, the level of gene rearrangement being correlated with the abundance of short dispersed repeated sequences throughout the genome. Intron-containing-, fragmented-, and fragmented and scrambled coding regions are common features of mitochondrial and chloroplast gene structure and organization within the group. The level of ribosomal RNA gene sequence divergence in both mitochondrial and chloroplast genomes is higher in the Chlamydomonas lineage than in land plants and is most likely due to higher rates of nucleotide substitution in Chlamydomonas organellar DNAs. The mechanisms as well as the selective pressures that shaped the organellar genomes in the Chlamydomonas lineage remain to be explained.
TL;DR: The main strategies describing possible ways to analyse the function of new genes that have been identified by systematic sequencing of Saccharomyces cerevisiae genome are described.
Abstract: The genome of the yeast Saccharomyces cerevisiae was sequenced by an international consortium of laboratories from Europe, Canada, the U.S.A. and Japan. This project is now finished and the complete sequence of the first eukaryotic genome was released to the public data bases in April 1996. An overview and preliminary analysis of the entire genome sequence was presented in a special issue of Nature in May 1997, entitled "The yeast genome directory". At its origin the Yeast Genome Sequencing Project provoked much debate and controversy; however, the final results obtained and the insights this has given us into the organisation and content of a eukaryotic genome have more than justified the expectations of the supporters of the project. The importance of genomic sequencing and analysis, especially of model organisms, is now widely accepted and this has resulted in the birth of the new science of genomics (Botstein & Cherry, 1997, Proc. Natl. Acad. Sci. U.S.A. 94, 5506). The information from gene and protein sequences ultimately lead to functional description of all genes. The main strategies describing possible ways to analyse the function of new genes that have been identified by systematic sequencing of Saccharomyces cerevisiae genome are described.