Metabolic gene clusters or biosynthetic gene clusters are tightly linked sets of mostly non-homologousgenes participating in a common, discrete metabolic pathway. The genes are in physical vicinity to each other on the genome, and their expression is often coregulated.[1][2][3] Metabolic gene clusters are common features of bacterial[4], most fungal[5], and plant[6] genomes (although there is contradicting evidence on plant[7] organisms). They are most widely known for producing secondary metabolites, the source or basis of most pharmaceutical compounds, natural toxins, chemical communication, and chemical warfare between organisms. Metabolic gene clusters are also involved in nutrient acquisition, toxin degradation,[8] antimicrobial resistance, and vitamin biosynthesis.[5] Given all these properties of metabolic gene clusters, they play a key role in shaping microbial ecosystems, including microbiome-host interactions. Thus several computational genomics tools have been developed to predict metabolic gene clusters.
Databases
MIBiG, BiG-FAM
Bioinformatic tools
Tools based on rules
Bioinformatic tools have been developed to predict, and determine the abundance and expression of, this kind of gene cluster in microbiome samples, from metagenomic data.[9] Since the size of metagenomic data is considerable, filtering and clusterization thereof are important parts of these tools. These processes can consist of dimensionality -reduction techniques, such as Minhash,[10] and clusterization algorithms such as k-medoids and affinity propagation. Also several metrics and similarities have been developed to compare them.