
A base pair (bp) is a fundamental unit of double-stranded nucleic acids consisting of two nucleobases bound to each other by hydrogen bonds. They form the building blocks of the DNA double helix and contribute to the folded structure of both DNA and RNA. Dictated by specific hydrogen bonding patterns, "Watson–Crick" (or "Watson–Crick–Franklin") base pairs (guanine–cytosine and adenine–thymine/uracil)[1] allow the DNA helix to maintain a regular helical structure that is subtly dependent on its nucleotide sequence.[2] The complementary nature of this based-paired structure provides a redundant copy of the genetic information encoded within each strand of DNA. The regular structure and data redundancy provided by the DNA double helix make DNA well suited to the storage of genetic information, while base-pairing between DNA and incoming nucleotides provides the mechanism through which DNA polymerase replicates DNA and RNA polymerase transcribes DNA into RNA. Many DNA-binding proteins can recognize specific base-pairing patterns that identify particular regulatory regions of genes.
Intramolecular base pairs can occur within single-stranded nucleic acids. This is particularly important in RNA molecules (e.g., transfer RNA), where Watson–Crick base pairs (guanine–cytosine and adenine-uracil) permit the formation of short double-stranded helices, and a wide variety of non–Watson–Crick interactions (e.g., G–U or A–A) allow RNAs to fold into a vast range of specific three-dimensional structures. In addition, base-pairing between transfer RNA (tRNA) and messenger RNA (mRNA) forms the basis for the molecular recognition events that result in the nucleotide sequence of mRNA becoming translated into the amino acid sequence of proteins via the genetic code.
個々の遺伝子または生物のゲノム全体のサイズは、DNAが通常二本鎖であるため、塩基対で測定されることが多い。したがって、総塩基対数は、一方の鎖のヌクレオチド数に等しい(テロメアの非コード一本鎖領域を除く)。半数体ヒトゲノム(23本の染色体)は、約32億塩基対の長さで、20,000~25,000個の異なるタンパク質コード遺伝子を含むと推定されている。[ 3 ] [ 4 ] [ 5 ] [ 6 ]キロベース(kb)は、分子生物学における測定単位で、DNAまたはRNAの1000塩基対に相当する。[ 7 ]地球上のDNA塩基対の総数は、5.0 × 10と推定されている。37重量は 500 億トンである。 [ 8 ]それに対し、生物圏の総質量は4 TtC (兆トン炭素)にも達すると推定されている。 [ 9 ]
本稿では、IUPACの1970年の勧告[ 10 ]に従い、あらゆる種類の塩基対を含む非共有結合相互作用を記述する際に「•」記号を使用する。N3.4.2
IUPACによれば、「-」は共有結合を暗示するため許容されず、「:」と「/」も比率と誤解される可能性があるため許容されない。記号を一切使用しないことも、(共有結合した)ポリマー配列と混同される可能性があるため許容されない。[ 10 ]: N3.4.2
IUPACは非共有結合の種類を区別するための具体的な推奨事項を定めていない。区別が必要な場合、本稿ではホグスティーンペアを表すのに「*」を用いる。
水素結合は、上述の塩基対形成規則の根底にある化学的相互作用である。水素結合供与体と受容体の適切な幾何学的対応により、「正しい」対のみが安定的に形成される。GC含量の高いDNAは、GC含量の低いDNAよりも安定している。しかしながら、重要なのは、スタッキング相互作用が主に二重らせん構造の安定化に寄与しているということである。ワトソン・クリック塩基対形成の全体的な構造安定性への寄与は最小限であるが、相補性の根底にある特異性におけるその役割は、対照的に、セントラルドグマの鋳型依存プロセス(例えばDNA複製)の根底にあるため、極めて重要である。[ 11 ]
より大きな核酸塩基であるアデニンとグアニンは、プリンと呼ばれる二重環構造のクラスに属します。より小さな核酸塩基であるシトシンとチミン(およびウラシル)は、ピリミジンと呼ばれる単環構造のクラスに属します。プリンはピリミジンとのみ相補的です。ピリミジン同士のペアリングは、分子が離れすぎていて水素結合が確立されないため、エネルギー的に不利です。プリン同士のペアリングは、分子が近すぎて重なり反発が生じるため、エネルギー的に不利です。AT、GC、またはUA(RNA中)のプリン-ピリミジン塩基対形成は、適切な二重らせん構造をもたらします。他に考えられるプリン-ピリミジンペアリングは、AC、GT、およびUG(RNA中)のみです。これらのペアリングは、水素供与体と受容体のパターンが一致しないため、ミスマッチとなります。 2つの水素結合を持つGUペアリングは、RNAではかなり頻繁に発生します(ウォブル塩基対を参照)。
対になったDNA分子とRNA分子は室温では比較的安定していますが、2本のヌクレオチド鎖は、分子の長さ、ミスペアリングの程度(もしあれば)、およびGC含量によって決まる融解点を超えると分離します。GC含量が高いほど融解温度も高くなるため、 Thermus thermophilusなどの極限環境微生物のゲノムが特にGC含量が高いのは当然のことです。逆に、頻繁に分離する必要のあるゲノム領域(例えば、頻繁に転写される遺伝子のプロモーター領域)は、比較的GC含量が低くなっています(例えば、TATAボックスを参照)。PCR反応用のプライマーを設計する際には、GC含量と融解温度も考慮する必要があります 。
以下のDNA配列は、二本鎖DNAのペアパターンを示しています。慣例として、上鎖は5'末端から3'末端に向かって記述されるため、下鎖(相補鎖)は3'から5'の方向に記述されます。
ATCGATTGAGCTCTAGCGTAGCTAACTCGAGATCGCAUCGAUUGAGCUCUAGCGUAGCUAACUCGAGAUCGC標準的なワトソン・クリック型塩基対(A・T/UG・C)に加えて、特定の条件下では、異なる塩基配向、水素結合の数や形状を持つ塩基対形成が促進される場合もある。これらの塩基対形成は、局所的な主鎖形状の変化を伴う。
これらのうち最も一般的なのは、転写中に多くのコドンの3番目の塩基位置でtRNAとmRNAの間で起こるウォブル塩基対形成[ 13 ] 、および一部のtRNA合成酵素によるtRNAのチャージ中に起こるウォブル塩基対形成[ 14 ]である。これらは、一部のRNA配列の二次構造でも観察されている[ 15 ] 。
さらに、プリン塩基の異なる「面」が対合に使用される場合、フーグスティーン塩基対(通常、A*U/T および G*C と表記される)が発生することがあります。これは、標準的なワトソン・クリック対合と動的平衡にある一部の DNA 配列(例: CA および TA ジヌクレオチド)で発生します。 [ 12 ]また、一部のタンパク質-DNA 複合体でも観察されています。[ 16 ] tRNAには、プリンとピリミジンの両方が異なる「面」を使用する一種の逆フーグスティーン塩基対も存在します。[ 17 ] [ 18 ]
これらの代替塩基対形成に加えて、RNAの二次構造および三次構造では、幅広い塩基間水素結合が観察される。[ 19 ]これらの結合は、RNAの正確で複雑な形状、および相互作用パートナーとの結合にしばしば必要となる。[ 19 ]
ミスマッチ塩基対は、 DNA複製エラーや相同組換えの中間体として生成されることがあります。ミスマッチ修復プロセスは通常、正常なDNA塩基対の長い配列内の少数の塩基ミスマッチを認識して正しく修復する必要があります。DNA複製中に形成されたミスマッチを修復するために、テンプレート鎖と新しく形成された鎖を区別し、新しく挿入された誤ったヌクレオチドのみが除去されるように(突然変異の発生を避けるために)いくつかの特徴的な修復プロセスが進化しました。[ 20 ] DNA複製中のミスマッチ修復に用いられるタンパク質と、このプロセスの欠陥の臨床的意義については、記事「DNAミスマッチ修復」で説明されています。組換え中のミスマッチ修復プロセスについては、記事「遺伝子変換」で説明されています。
ヌクレオチドの化学的類似体は、本来のヌクレオチドの代わりに非標準的な塩基対を形成し、DNA複製やDNA転写におけるエラー(主に点突然変異)を引き起こす可能性がある。これは、それらの等立体化学によるものである。一般的な変異原性塩基類似体の1つは5-ブロモウラシルであり、これはチミンに似ているが、エノール型のグアニンと塩基対を形成できる。[ 21 ]
DNAインターカレーターと呼ばれる他の化学物質は、一本鎖上の隣接する塩基間の隙間に入り込み、塩基になりすますことでフレームシフト変異を誘発し、DNA複製機構がインターカレーション部位でヌクレオチドをスキップしたり、追加のヌクレオチドを挿入したりします。ほとんどのインターカレーターは大きな多環芳香族化合物であり、発がん性物質として知られているか、あるいは疑われています。例としては、臭化エチジウムやアクリジンなどがあります。[ 22 ]

D/R NA分子の長さを表す際によく用いられる略語は以下のとおりです。
一本鎖DNA/RNAの場合、ヌクレオチドは対になっていないため、ヌクレオチド単位(nt、またはknt、Mnt、Gntと略記)が用いられます。コンピュータの記憶容量単位と塩基を区別するために、塩基対にはkbp、Mbp、Gbpなどの単位が用いられることがあります。
センチモルガンは染色体上の距離を示すためにもよく使われますが、それに対応する塩基対の数は染色体交叉のパターンによって大きく異なります。ヒトゲノムでは、センチモルガンは約100万塩基対です。[ 24 ] [ 25 ]
非天然塩基対 (UBP) は、実験室で作成され、自然界には存在しない、設計されたDNAのサブユニット (またはヌクレオ塩基)です。自然界に存在する 2 つの塩基対、A•T (アデニン-チミン) と G•C (グアニン-シトシン) に加えて、新たに作成されたヌクレオ塩基を使用して 3 番目の塩基対を形成する DNA 配列が記述されています。Steven A. Benner、Philippe Marliere、Floyd E. Romesberg、Ichiro Hiraoらが率いるチームを含むいくつかの研究グループが、DNA の 3 番目の塩基対を探しています。[ 26 ]代替水素結合、疎水性相互作用、金属配位に基づくいくつかの新しい塩基対が報告されています。[ 27 ] [ 28 ] [ 29 ] [ 30 ]
1989年、スティーブン・ベナー(当時チューリッヒのスイス連邦工科大学に勤務)とそのチームは、試験管内でシトシンとグアニンの改変型をDNA分子に導入した。[ 31 ] RNAとタンパク質をコードするヌクレオチドは、試験管内で正常に複製された。それ以来、ベナーのチームは、原料を必要とせずに、ゼロから外来塩基を合成できる細胞を設計しようと試みている。[ 32 ]
In 2002, Ichiro Hirao's group in Japan developed an unnatural base pair between 2-amino-8-(2-thienyl)purine (s) and pyridine-2-one (y) that functions in transcription and translation, for the site-specific incorporation of non-standard amino acids into proteins.[33] In 2006, they created 7-(2-thienyl)imidazo[4,5-b]pyridine (Ds) and pyrrole-2-carbaldehyde (Pa) as a third base pair for replication and transcription.[34] Afterward, Ds and 4-[3-(6-aminohexanamido)-1-propynyl]-2-nitropyrrole (Px) was discovered as a high fidelity pair in PCR amplification.[35][28] In 2013, they applied the Ds-Px pair to DNA aptamer generation by in vitro selection (SELEX) and demonstrated the genetic alphabet expansion significantly augment DNA aptamer affinities to target proteins.[36]
In 2012, a group of American scientists led by Floyd Romesberg, a chemical biologist at the Scripps Research Institute in San Diego, California, published that his team designed an unnatural base pair (UBP).[29] The two new artificial nucleotides or Unnatural Base Pair (UBP) were named d5SICS and dNaM. More technically, these artificial nucleotides bearing hydrophobic nucleobases, feature two fused aromatic rings that form a (d5SICS–dNaM) complex or base pair in DNA.[32][37] His team designed a variety of in vitro or "test tube" templates containing the unnatural base pair and they confirmed that it was efficiently replicated with high fidelity in virtually all sequence contexts using the modern standard in vitro techniques, namely PCR amplification of DNA and PCR-based applications.[29] Their results show that for PCR and PCR-based applications, the d5SICS–dNaM unnatural base pair is functionally equivalent to a natural base pair, and when combined with the other two natural base pairs used by all organisms, A–T and G–C, they provide a fully functional and expanded six-letter "genetic alphabet".[37]
In 2014 the same team from the Scripps Research Institute reported that they synthesized a stretch of circular DNA known as a plasmid containing natural T-A and C-G base pairs along with the best-performing UBP Romesberg's laboratory had designed and inserted it into cells of the common bacterium E. coli that successfully replicated the unnatural base pairs through multiple generations.[26] The transfection did not hamper the growth of the E. coli cells and showed no sign of losing its unnatural base pairs to its natural DNA repair mechanisms. This is the first known example of a living organism passing along an expanded genetic code to subsequent generations.[37][38] Romesberg said he and his colleagues created 300 variants to refine the design of nucleotides that would be stable enough and would be replicated as easily as the natural ones when the cells divide. This was in part achieved by the addition of a supportive algal gene that expresses a nucleotide triphosphate transporter which efficiently imports the triphosphates of both d5SICSTP and dNaMTP into E. coli bacteria.[37] Then, the natural bacterial replication pathways use them to accurately replicate a plasmid containing d5SICS–dNaM. Other researchers were surprised that the bacteria replicated these human-made DNA subunits.[39]
The successful incorporation of a third base pair is a significant breakthrough toward the goal of greatly expanding the number of amino acids which can be encoded by DNA, from the existing 20 amino acids to a theoretically possible 172, thereby expanding the potential for living organisms to produce novel proteins.[26] The artificial strings of DNA do not encode for anything yet, but scientists speculate they could be designed to manufacture new proteins which could have industrial or pharmaceutical uses.[40] Experts said the synthetic DNA incorporating the unnatural base pair raises the possibility of life forms based on a different DNA code.[39][40]
The following sources have information on the free energy (thermodynamic measures of strength) of base pairs:
しかし、2 つの核酸塩基間の最小エネルギー水素結合状態が何であるかを知るだけでは十分ではありません。核酸分子の安定性は塩基スタッキングにも由来し、その強度は元のバージョンに対する修飾塩基によって変化する可能性があります。2 つの塩基の最適な水素結合状態は、核酸の骨格の不自然な量の曲がりを必要とする場合もあります。これらすべてが、核酸の二次構造の文脈における塩基対の「実効」強度に寄与するため、このような構造を予測するには、自由エネルギー (37 °C で) とエンタルピー (異なる温度に再スケーリングするため) の観点から塩基対を記述する「最近接」モデルが必要です。[ 42 ]
「最近傍」モデルのリストは、核酸構造予測§ 熱力学モデルで確認できます。
アクリジンが存在する場合の各突然変異事象は、単一の塩基対の追加または除去をもたらす。
×
10⁵
塩基
対の距離に相当する。
{{cite book}}: CS1 maint: パブリッシャーの場所 (リンク) (特に第6章と第9章を参照)