Showing posts with label RNA Biology. Show all posts
Showing posts with label RNA Biology. Show all posts

Saturday, August 08, 2026

RNA’s Next Act: From Biological Messenger to Programmable Nanomachine

 

Researchers around the world are learning to make RNA sense, compute, assemble, edit and organize living cells. The convergence of RNA biology and nanotechnology could reshape medicine, agriculture and synthetic biology.
https://thernablog.blogspot.com/

Researchers around the world are learning to make RNA sense, compute, assemble, edit and organize living cells. The convergence of RNA biology and nanotechnology could reshape medicine, agriculture and synthetic biology.

For decades, RNA occupied an awkward middle ground in biology. DNA stored genetic information; proteins performed most of the cell’s chemistry; RNA carried instructions between them.

That hierarchy has steadily collapsed.

RNA is now understood as an extraordinarily versatile molecule. It can catalyse reactions, recognize metabolites, regulate genes, form intricate three-dimensional structures and reorganize itself in response to its surroundings. Some RNAs act as switches. Others serve as scaffolds, molecular guides or components of cellular machines.

These properties are drawing RNA biology into an unexpected partnership with nanotechnology.

A recent Nature feature described this transformation particularly well: RNA's ability to fold, switch and reorganize is increasingly being exploited to build nanoscale biological technologies. Researchers are no longer interested only in discovering what an RNA molecule naturally does. They are beginning to ask what RNA can be engineered to do [1].

The distinction is important. It marks a transition from RNA biology as predominantly a science of discovery towards RNA biology as an engineering discipline.

And that transition is occurring globally.

Across laboratories in the United States, Europe, China, South Korea, Australia and elsewhere, researchers are developing RNA circuits, nanostructures, editing platforms, synthetic condensates, delivery vehicles and agricultural technologies. Individually, these advances belong to different specialties. Collectively, they suggest that RNA could become one of the principal programmable materials of twenty-first-century biology.

A molecule that can carry information and become machinery

RNA possesses an unusual combination of properties.

Its sequence stores information, much like DNA. But unlike the familiar textbook depiction of messenger RNA as a simple linear strand, RNA readily folds back upon itself. Complementary regions form stems, loops, bulges, junctions and elaborate tertiary structures.

Those structures matter because shape determines function.

An RNA molecule can expose or conceal a regulatory sequence. It can recognize another RNA, recruit a protein, bind a small molecule or switch conformation after encountering a particular chemical signal.

For nanotechnologists, this combination of information, structure and dynamics is particularly attractive.

DNA nanotechnology established that nucleic acids can be programmed to self-assemble into intricate structures. RNA potentially goes further because it is naturally produced within cells and participates directly in cellular regulation.

An RNA nanostructure therefore need not remain a passive molecular sculpture.

It could become machinery.

Turning RNA into a cellular switch

One of the clearest demonstrations of programmable RNA comes from riboswitches.

Natural riboswitches alter gene expression after binding specific metabolites. Their structures change in response to a chemical cue, affecting whether downstream genetic information is expressed.

Synthetic biologists are now rewriting this principle.

At the University of Konstanz in Germany, Jörg Hartig and colleagues developed engineered riboswitches based on bacterial xanthine aptamers that respond to oxypurinol, the active metabolite of the clinically used drug allopurinol. The work demonstrated strong chemically controlled regulation of gene expression in mammalian cells [2].

The long-term attraction is control.

Gene therapies generally aim to introduce or restore biological functions, but regulating therapeutic output after treatment can be difficult. An RNA switch introduces another regulatory layer: a therapeutic construct could, in principle, be activated or modulated pharmacologically.

That would make gene therapy less like installing a permanently active programme and more like installing a tunable biological system.

Other researchers are moving beyond individual switches.

At Pohang University of Science and Technology in South Korea, Jongmin Kim and colleagues have developed programmable RNA-based systems capable of processing multiple molecular inputs to regulate endogenous gene expression [3].

The analogy with electronic logic gates is useful, although biology is considerably messier than silicon.

A conventional engineered gene might respond to one trigger. A more sophisticated RNA circuit could require several conditions to be satisfied before generating an output.

A therapeutic cell might eventually detect multiple disease-associated signals and activate a treatment only when the appropriate combination is present.

In that sense, RNA begins to resemble a molecular decision-making system.

Building structures that cells manufacture themselves

An even more striking development is RNA origami.

The approach builds on a fundamental idea from nucleic-acid nanotechnology: predictable base-pairing interactions can be used to make nucleic-acid strands fold into designed geometries.

RNA introduces an additional possibility. Because cells naturally transcribe RNA, the information needed to construct a nanostructure can itself be genetically encoded.

Rather than manufacture a nanoscale object outside a cell and then attempt to deliver the completed structure, researchers could potentially provide the instructions and allow the cell to build it.

Work led by Fei Zhang and colleagues demonstrated this principle by designing RNA molecules that co-transcriptionally self-assemble within human-cell nuclei into rings, zigzag scaffolds, lattices and mesh-like architectures [4].

The achievement matters not simply because complex shapes can be produced inside cells.

Those shapes could ultimately become functional.

An RNA scaffold might recruit selected proteins.

Another could organize enzymes.

A structure positioned near chromatin might alter regulatory interactions.

Still another could provide the architecture for an intracellular biosensor.

The cell would no longer merely express an RNA sequence. It would manufacture a designed nanoscale object.

That possibility begins to erase the boundary between synthetic biology and nanofabrication.

The rise of artificial RNA organelles

Cells are spatially organized systems. Many biochemical reactions succeed because the correct molecules are concentrated in the correct location.

Yet not every cellular compartment has a membrane.

Biomolecular condensates can form through networks of interactions among proteins and nucleic acids, producing dense, dynamic compartments that remain physically distinct from their surroundings.

Researchers are now attempting to recreate this principle synthetically.

Giacomo Fabrini, Lorenzo Di Michele, Elisa Franco, Paul Rothemund and collaborators demonstrated the co-transcriptional production of programmable RNA condensates and synthetic organelle-like structures [5].

In 2026, Shiyi Li, Yuna Kim and colleagues extended this strategy to programmable artificial RNA condensates within mammalian cells [6].

Related work has now demonstrated nano-engineered RNA organelle-like assemblies in bacteria as well [7].

The implications extend beyond constructing unusual intracellular shapes.

Imagine creating an artificial compartment that concentrates several enzymes participating in one biochemical pathway. Instead of those enzymes diffusing independently throughout the cytoplasm, an RNA architecture could bring them together.

Another compartment might recruit particular RNA-binding proteins. A third could sequester molecules whose accumulation disrupts normal cellular function.

The ambition is therefore larger than regulating individual genes.

Synthetic RNA might eventually allow researchers to redesign parts of the physical geography of the cell.

When the nanostructure becomes the medicine

RNA nanotechnology is also blurring a traditional distinction in drug delivery: the boundary between therapeutic cargo and carrier.

Conventionally, a nanoparticle transports an active molecule.

But RNA structures can themselves be biologically active.

Researchers led by Hao Yan and Yung Chang at Arizona State University demonstrated that RNA-origami nanostructures can activate antitumour immunity. In mouse models, the engineered RNA structures functioned as potent immunostimulatory materials and produced antitumour responses through mechanisms involving innate and adaptive immunity [8].

This illustrates an important principle.

For an engineered RNA therapeutic, biological behaviour might depend not only on nucleotide sequence but also on shape, size, stability, molecular interactions and subcellular destination.

Future RNA medicines could therefore be designed at several levels simultaneously:

sequence, structure, delivery, immune recognition and biological function.

That is a substantially richer engineering problem than simply asking what protein an mRNA encodes.

Delivery remains the decisive bottleneck

But even the most sophisticated RNA architecture is useless if it never reaches the appropriate biological destination.

Delivery remains one of the central constraints on RNA biotechnology.

Lipid nanoparticles demonstrated dramatically that RNA can be protected and delivered effectively. The broader challenge, however, is much harder than merely encapsulating RNA.

Different applications require RNA to reach different organs, tissues and cell populations. Once inside a cell, some RNA molecules need to remain in the cytoplasm, whereas others may need access to particular intracellular environments.

Nanoparticle engineers therefore confront a cascade of barriers: extracellular stability, biodistribution, tissue penetration, cellular uptake, endosomal escape, intracellular release and eventual degradation.

A 2026 Nature Materials review describes this shift towards precision mRNA delivery, in which administration route, particle chemistry, targeting strategies and responsive release are engineered according to the biological destination rather than assuming that one formulation can serve every application [9].

This may prove to be one of nanotechnology's largest contributions to RNA biology.

The future is unlikely to belong to one universally optimal nanoparticle.

It is more likely to involve families of delivery architectures optimized for particular tissues, cell types and therapeutic tasks.

An exquisitely engineered RNA system delivered to the wrong cells is still a failed therapy.

From reading RNA to rewriting it

At the same time, RNA itself is becoming increasingly editable.

RNA editing is particularly attractive because it provides a way of changing biological information without permanently changing genomic DNA.

One major strategy exploits adenosine deaminases acting on RNA, or ADARs, which naturally convert adenosine to inosine in double-stranded RNA contexts.

Researchers can redirect endogenous ADAR activity using engineered guide RNAs.

Yuanfan Sun and colleagues recently showed that guide RNAs designed to mimic structural features of highly edited endogenous ADAR substrates can improve RNA base editing, illustrating how RNA structure itself can be engineered to recruit cellular editing machinery more effectively [10].

Researchers are also expanding the chemistry available for programmable RNA editing.

Yuan Zhuang, Qingguo Zhu, Chengqi Yi and colleagues developed AIM, a single-strand deaminase-assisted platform capable of A-to-I, C-to-U or simultaneous A+C editing within user-defined RNA regions [11].

Intriguingly, the engineering is also moving in the opposite molecular direction.

Hyeon Woo Im, Sangsu Bae and colleagues recently repurposed engineered ADAR domains for highly precise A-to-G DNA base editing within DNA–RNA hybrids. The work is not an RNA-editing technology itself, but it demonstrates how enzymes originating from RNA biology can be redesigned for entirely new information-processing roles [12].

The broader trend is unmistakable.

RNA is no longer simply being read as the output of gene expression.

It is becoming an editable information layer between genotype and phenotype.

Artificial intelligence enters RNA biology

The enormous RNA design space makes computation increasingly important.

Predicting RNA behaviour is difficult. RNA molecules are flexible, individual sequences can populate multiple conformations, and their structures can be influenced by ions, proteins, other nucleic acids and the local cellular environment.

Machine learning is beginning to address parts of this problem.

In 2026, Philip Fradkin, Bo Wang and colleagues reported Orthrus, a foundation model trained to learn evolutionary and functional representations of mature RNA molecules. The system outperformed several genomic foundation models on RNA-property prediction tasks and could distinguish functional differences among transcript isoforms [13].

At the structural level, Sumit Tarafder and Debswapna Bhattacharya introduced RNAbpFlow, a generative approach that incorporates base-pair information into three-dimensional RNA structure generation [14].

The eventual significance of such systems may extend beyond prediction.

The more transformative objective is inverse design.

Instead of asking:

What does this RNA sequence do?

a researcher might ask:

What RNA sequence should I build to obtain this function?

Design an RNA that recognizes this metabolite.

Design a scaffold that recruits these proteins.

Design an untranslated region that produces a particular expression profile.

Design a nanostructure that assembles only under defined intracellular conditions.

AI would generate candidate molecules; experiments would determine which ones actually work.

Machine learning would not eliminate experimental RNA biology.

It could profoundly change where experimentation begins.

Agriculture becomes another RNA-engineering frontier

Many of the same principles are appearing outside medicine.

Agriculture is becoming one of the most promising—and technically difficult—arenas for RNA nanotechnology.

Double-stranded RNA can initiate RNA interference and selectively suppress genes in plants, viruses, fungi and insect pests. That sequence specificity creates opportunities for crop-protection strategies fundamentally different from conventional broad-spectrum pesticides.

But the agricultural environment is unforgiving.

RNA applied to a leaf might encounter ultraviolet radiation, rain, nucleases, waxy barriers and cell walls before reaching the cellular compartment where gene silencing must occur.

Researchers have consequently investigated layered double hydroxides, carbon-based materials, mesoporous silica, chitosan formulations, lipid carriers and other nanomaterials as potential RNA-delivery platforms.

Yet a 2026 analysis in Nature Plants highlights a crucial distinction: stabilizing dsRNA on a leaf or increasing its accumulation in the apoplast does not demonstrate effective delivery to the plant cytoplasm. For antiviral applications in particular, reliable symplastic delivery remains incompletely understood [15].

This is precisely where RNA biology and nanotechnology must converge.

Sequence optimization cannot, by itself, solve a transport problem.

An effective carrier has to be designed around the biology of the organism, tissue and intracellular destination.

And another challenge is already appearing downstream.

RNA biopesticides must eventually move through regulatory systems developed largely for older classes of crop-protection chemicals. Researchers led by Sandya Gunasekara and Neena Mitter have argued that international regulatory harmonization will be important if dsRNA-based pesticides are to move efficiently towards widespread deployment [16].

The agricultural RNA revolution therefore depends simultaneously on molecular biology, nanomaterials, ecology, formulation engineering, field performance and regulation.

The convergence matters more than any single technology

RNA switches, RNA origami, artificial condensates, editing systems, machine-learning models and nanocarriers can appear to belong to separate scientific stories.

They probably do not.

Their convergence is the more important development.

Imagine a future therapeutic system.

A targeted nanoparticle first delivers an engineered RNA construct to a specific population of cells.

Inside those cells, the RNA folds into a predetermined architecture.

The structure recognizes a disease-associated molecular signature.

An RNA circuit evaluates the signal.

Only when the appropriate conditions are satisfied does the system recruit an editing enzyme or initiate production of a therapeutic protein.

Later, the RNA degrades and the programme disappears.

No research group has yet built this complete system.

But laboratories around the world are developing many of its components.

That is why the present period in RNA biology is so consequential.

The technologies are no longer advancing only along isolated tracks.

They are beginning to intersect.

A genuinely global scientific effort

The geography of this emerging field reflects its intellectual diversity.

Researchers in Germany are designing chemically controllable RNA switches.

South Korean scientists are building molecular logic systems and developing new uses for RNA-associated editing enzymes.

US laboratories are advancing RNA origami, intracellular nanostructures and immunologically active RNA materials.

British, European and American teams are collaborating on synthetic RNA condensates and artificial organelles.

Chinese researchers are expanding the chemistry and controllability of RNA editing.

Australian researchers and international collaborators are tackling the regulatory and agricultural dimensions of RNA biotechnology.

Computational researchers across multiple institutions are attempting to solve RNA sequence–structure–function relationships with increasingly sophisticated machine-learning models.

These efforts cross disciplinary boundaries as readily as geographical ones.

A nanotechnologist increasingly needs to understand RNA folding.

An RNA biologist may need materials science.

A synthetic biologist may need machine learning.

A plant biologist may need colloid chemistry and nanocarrier engineering.

A clinician may ultimately need all of them.

The field emerging from this convergence lies somewhere between molecular biology, materials science, computation and engineering.

RNA biology is changing its central question

The most important transformation may ultimately be conceptual.

Classical molecular biology usually begins with discovery.

Researchers identify a molecule and ask:

What does it do?

That question will remain fundamental.

But RNA engineering introduces another:

What could we make it do?

The distinction is profound.

One question seeks to understand biological systems as they exist.

The other seeks to construct biological behaviour.

Over the coming decade, RNA biology will increasingly involve both.

Scientists will continue discovering new RNAs, modifications, structures and regulatory pathways. Alongside them, however, an expanding engineering discipline will attempt to build RNA molecules that sense, compute, assemble, organize, edit and deliver biological information.

Some will become therapeutics.

Others may become intracellular sensors, synthetic organelles or programmable regulatory circuits.

Some could protect crops.

Others might become components of biological manufacturing systems that are difficult to envision today.

RNA was once described primarily as the molecule carrying DNA's instructions to the protein-making machinery of the cell.

That definition increasingly seems inadequate.

The deeper transformation now underway is that RNA is becoming something much more ambitious:

a programmable material from which researchers can begin to engineer biology itself.


References

1. Eisenstein, M. The dark horse of biology: how RNA is becoming a nanotool maker's dream. Nature 655, S2–S5 (2026). doi:10.1038/d41586-026-02180-6.

2. Hedwig, V. et al. Engineering oxypurinol-responsive riboswitches based on bacterial xanthine aptamers for gene expression control in mammalian cell culture. Nucleic Acids Research 53, gkae1189 (2025). doi:10.1093/nar/gkae1189.

3. Kang, H., Park, D. & Kim, J. Logical regulation of endogenous gene expression using programmable, multi-input processing CRISPR guide RNAs. Nucleic Acids Research 52, 8595–8608 (2024). doi:10.1093/nar/gkae549.

4. Chang, X. et al. Designer RNA nanostructures co-transcribed and self-assembled inside human cell nuclei. Nature Communications 17, 1055 (2026). doi:10.1038/s41467-025-67817-y.

5. Fabrini, G. et al. Co-transcriptional production of programmable RNA condensates and synthetic organelles. Nature Nanotechnology 19, 1665–1673 (2024). doi:10.1038/s41565-024-01726-x.

6. Li, S. et al. Programmable artificial RNA condensates in mammalian cells. Nature Nanotechnology 21, 821–830 (2026). doi:10.1038/s41565-026-02164-7.

7. Ng, B. et al. Expression of nano-engineered RNA organelles in bacteria. Nature Communications 17, 2752 (2026). doi:10.1038/s41467-026-69336-w.

8. Qi, X. et al. RNA origami nanostructures for potent and safe anticancer immunotherapy. ACS Nano 14, 4727–4740 (2020). doi:10.1021/acsnano.0c00602.

9. Deng, H., Li, L., Zhao, C. et al. Nanotechnology-mediated precision delivery of mRNA. Nature Materials (2026). doi:10.1038/s41563-026-02623-5.

10. Sun, Y., Cao, Y., Song, Y. et al. Improved RNA base editing with guide RNAs mimicking highly edited endogenous ADAR substrates. Nature Biotechnology 44, 464–476 (2026). doi:10.1038/s41587-025-02628-6.

11. Zhuang, Y., Zhu, Q., Wu, H. et al. Single-strand deaminase-assisted editing for functional RNA manipulation. Nature Biotechnology (2026). doi:10.1038/s41587-025-02956-7.

12. Im, H. W., Jeong, B., Lee, Y. et al. Engineered ADARs enable precision A-to-G base editing of DNA. Nature Biotechnology (2026). doi:10.1038/s41587-026-03223-z.

13. Fradkin, P., Shi, R. I., Dalal, T. et al. Orthrus: toward evolutionary and functional RNA foundation models. Nature Methods 23, 935–945 (2026). doi:10.1038/s41592-026-03064-3.

14. Tarafder, S. & Bhattacharya, D. RNAbpFlow: base pair-augmented SE(3) flow matching for conditional RNA 3D structure generation. Nature Methods 23, 1349–1358 (2026). doi:10.1038/s41592-026-03128-4.

15. Sede, A. R., Moorlach, B., Galli, M. et al. The unfulfilled potential of nanocarriers for RNA delivery in antiviral crop protection. Nature Plants 12, 1166–1178 (2026). doi:10.1038/s41477-026-02323-7.

16. Gunasekara, S., Fidelman, P., Fletcher, S. et al. The future of dsRNA-based biopesticides will require global regulatory cohesion. Nature Plants 11, 664–667 (2025). doi:10.1038/s41477-025-01953-7.

Wednesday, July 29, 2026

The Giant RNA Polymerase of CCHFV: A New Structural Window into Viral RNA Synthesis and Antiviral Design

 

https://thernablog.blogspot.com/
https://thernablog.blogspot.com/

Viral RNA polymerases are among the most important molecular machines in virology. They copy viral RNA, make viral transcripts, and control whether an RNA virus can successfully replicate inside a host cell. Because human cells do not use the same kind of processive RNA-dependent RNA polymerase for genome replication, these enzymes have long been attractive targets for antiviral drug discovery.

A recent study by Jia and colleagues, titled “RNA synthesis and substrate analog inhibition in the CCHFV polymerase,” provides a major structural and biochemical advance in this area. The work focuses on the L protein of Crimean-Congo hemorrhagic fever virus, or CCHFV, a tick-borne virus belonging to the Nairoviridae family. 

What makes this enzyme remarkable is its size. Nairoviridae L proteins are about 4,000 residues long, making them among the largest known viral polymerases. Until now, this enormous size came with a major mystery: why does this virus need such a large polymerase to perform a job that other RNA viruses accomplish with much smaller polymerase systems? Jia and colleagues address this by reporting structures of the full-length CCHFV L protein, including a 3.0 Å polymerase elongation complex.

The Nature study is important for two reasons. First, it gives us a clearer picture of how this giant viral enzyme organizes RNA synthesis. Second, it identifies nucleotide analogs with sofosbuvir-like ribose modifications that can specifically inhibit the CCHFV polymerase by immediate chain termination.

Why CCHFV polymerase matters

CCHFV is not an ordinary virus from a public-health perspective. It is a tick-borne biosafety level-4 pathogen, meaning it requires the highest level of laboratory containment. The virus causes Crimean-Congo hemorrhagic fever, a severe disease of major concern in endemic regions. Understanding how its polymerase works is therefore not only a structural biology question; it is also a foundation for antiviral discovery.

Like other segmented negative-sense RNA viruses, CCHFV depends on an RNA-dependent RNA polymerase, or RdRP, to copy and transcribe its RNA genome. The L protein contains multiple functional regions, including an endonuclease, the central RdRP module, and a cap-binding domain. These regions cooperate during viral transcription and replication. In segmented negative-sense RNA viruses, transcription often depends on “cap-snatching,” where the viral polymerase captures capped fragments from host RNAs and uses them as primers for viral mRNA synthesis.

The puzzle is that Nairoviridae L proteins are much larger than many related viral polymerase systems. Previous structures of segmented negative-sense RNA virus polymerases generally involved systems of about 2,000–2,500 residues, while Nairoviridae L proteins can reach 3,800–4,900 residues. The authors note that apart from an N-terminal OTU domain, the reason for this unusually large size had remained unclear.

A full-length view of a giant enzyme

To solve this problem, the researchers purified full-length CCHFV L protein and used cryo-electron microscopy to capture different structural states. They obtained apo and promoter-bound states, but the major breakthrough was the 3.0 Å elongation complex, which covered a much larger portion of the enzyme. This structure allowed the authors to define a more complete architecture of CCHFV L and to see how different regions cooperate during RNA synthesis.

One of the most interesting findings is that CCHFV L is not simply a larger version of other viral polymerases. It contains large additions and insertions in all three major functional regions. These additions reshape how the enzyme interacts with RNA. Two Nairoviridae-specific elements are especially important:

FID, or the fingers insertion domain, extends the downstream template RNA-binding path.

UPD, or the upstream product-binding domain, extends the upstream product RNA-binding path.

Together, these domains help explain why the CCHFV polymerase is so large. The extra mass is not random decoration. It appears to form additional RNA-binding paths and interaction networks that may help the enzyme handle long RNA products with sufficient processivity.

FID and UPD: two additions that change the RNA path

The study shows that FID lies near the downstream side of the RdRP active site and may help coordinate template RNA binding together with other polymerase regions. Structural analysis revealed positively charged residues in the relevant groove, consistent with a role in nucleic acid interaction. The authors propose that FID contributes not only to promoter binding but also to general downstream template RNA binding.

On the other side of the active site, UPD helps form an extended path for the upstream RNA product. The study identifies a tunnel-like route involving UPD, CBD, and mid-link regions, with positively charged residues positioned along the putative product RNA exit path. This suggests that the polymerase has evolved extra structural features to guide RNA as it emerges from the active site.

This is where the structural work becomes biologically meaningful. The researchers tested mutations in these interaction networks using a CCHFV minigenome assay. All 14 tested mutations reduced minigenome replication to varying degrees, and mutations affecting FID:RNA and UPD:RNA interactions had particularly strong effects, dropping replication below 20% of the wild-type level.

In simple terms, the extra domains are not just visible in the structure; they matter for viral RNA replication.

How the enzyme moves from initiation to elongation

The authors also propose a model for CCHFV RNA replication. In this model, the polymerase first recognizes the viral promoter and positions the 3′ end of the template RNA at the active site. As RNA synthesis progresses, the enzyme transitions into elongation. When the RNA duplex reaches roughly 10 base pairs, structural elements such as the lid and priming element move to accommodate the growing RNA duplex, and the lid helps separate template and product strands.

This model is useful because viral polymerases are not static machines. They must grip the promoter, initiate RNA synthesis, elongate the RNA chain, separate RNA strands, and eventually complete an entire replication cycle. The CCHFV L structure suggests that FID and UPD may help support processive elongation, possibly allowing the enzyme to synthesize the large L transcript of Bunyaviricetes.

The antiviral angle: sofosbuvir-like nucleotide analogs

The second major part of the study concerns nucleotide analog inhibitors. Nucleotide analogs work by mimicking natural nucleotide substrates. If a viral polymerase incorporates the analog into a growing RNA chain, the analog may disrupt further RNA synthesis.

The best-known example in this category is sofosbuvir, a nucleotide analog used to treat hepatitis C virus infection. Sofosbuvir’s active triphosphate form contains characteristic 2′-α-fluoro-2′-β-C-methyl ribose modifications that cause immediate chain termination in the hepatitis C virus polymerase.

Jia and colleagues asked whether similar chemistry could work against CCHFV RdRP. They found that nucleotide analogs carrying ribose-2′ modifications identical to sofosbuvir could be incorporated by CCHFV RdRP and then stop RNA synthesis immediately. Importantly, the same analogs were not incorporated by Lassa virus and Rift Valley fever virus polymerases in their assays, suggesting specificity for CCHFV among the tested systems.

The authors tested several analogs and found that all four base types with this ribose modification showed incorporation activity and chain-terminating behavior in the CCHFV system. Competition assays further supported the potential of these compounds, although different analogs varied in how strongly they competed with the corresponding natural nucleotides.

This does not mean that sofosbuvir itself is now a proven treatment for CCHFV infection. The study works at the enzyme and structural-biochemistry level. Drug development would still require prodrug optimization, cell culture testing, animal studies, pharmacokinetic evaluation, safety testing, and eventually clinical trials. But the work identifies a promising chemical logic: ribose 2′-α-fluoro-2′-β-C-methyl modification may be a useful starting point for anti-CCHFV nucleotide analog development.

Why this study matters for RNA biology

For RNA biologists, this work is exciting because it connects structure, mechanism, and inhibition in one system. The study does not merely show a beautiful cryo-EM structure. It links structural features to RNA-binding paths, tests their functional relevance through minigenome assays, and then uses active-site insight to explore antiviral inhibition.

It also reminds us that viral RNA polymerases are diverse. The familiar “right-hand” RdRP core is conserved, but viruses build many different accessory domains around that core. These additions can determine how the polymerase recognizes RNA, how it transitions between replication stages, how it separates strands, and how vulnerable it is to nucleotide analogs.

The CCHFV L protein is therefore more than a giant enzyme. It is a molecular example of how RNA viruses expand a conserved catalytic machine into a specialized replication platform.

Conclusions 

The new CCHFV polymerase structures help answer a long-standing question: why are Nairoviridae L proteins so large? The answer appears to lie in expanded RNA-binding architecture. Domains such as FID and UPD extend the paths of template and product RNA, helping organize the enzyme during replication. At the same time, the discovery that sofosbuvir-like nucleotide analogs can terminate CCHFV RNA synthesis provides a valuable starting point for antiviral research.

For The RNA Blog, this study is a reminder of why RNA biology remains one of the most dynamic areas of modern science. A single viral enzyme can teach us about evolution, molecular architecture, disease biology, and drug discovery. In the case of CCHFV, seeing the polymerase in action may be the first step toward learning how to stop it.



Wednesday, June 17, 2026

The Hidden Chaperones That Build RNA Silencing

The Hidden Chaperones That Build RNA Silencing

For years, RNA interference has been described like a clean molecular trick: give a cell a small RNA, let Argonaute hold it, and watch the matching message disappear.

But biology is rarely that simple.

The Hidden Chaperones That Build RNA Silencing For years, RNA interference has been described like a clean molecular trick: give a cell a small RNA, let Argonaute hold it, and watch the matching message disappear.  But biology is rarely that simpl
Thernablog.blogspot.com 


A new Nature study, “Structural basis for chaperone-guided assembly of RNA-induced silencing complex, shows that RISC assembly is not merely RNA loading. It is a carefully staged folding event. Argonaute does not simply grab a small RNA duplex. It must first be opened, held, stabilized, loaded, folded, and released.

At the center of this story is Argonaute, the protein engine of RNA silencing. In mature RISC, Argonaute carries one guide RNA strand and uses it to recognize target mRNAs. But before that final state, the small RNA arrives as a bulky duplex. The mature Argonaute structure is too compact to easily accept such a duplex. So the cell uses molecular chaperones.

Lee and colleagues identify an AGO–HSP90–p23 complex, which they call the AGO maturation complex, or AMC. This complex captures Argonaute in an RNA-free, pre-loading state. In this state, HSP90 and p23 hold AGO2 in a dramatically open conformation. The N domain is pulled away from the PAZ–MID–PIWI module, creating a widened, positively charged cleft that can receive a small RNA duplex.

This is the key visual message of the paper: Argonaute must be opened before it can become RISC.

The study also changes how we think about RNA itself. The RNA duplex is not only cargo. It acts almost like a folding cofactor. A duplex with a proper 5′ phosphate promotes productive AGO folding, while single-stranded RNA does not. The 5′ phosphate is especially important because it engages the MID domain, helping define which strand will become the guide. Duplex length also matters, with 22–23 nucleotide duplexes supporting efficient folding.

This has direct implications for siRNA therapeutics. Many approved siRNA drugs depend on chemical modifications such as 2′-fluoro and 2′-O-methyl substitutions. The paper shows that some modification patterns are compatible with AGO folding, while others can impair it. In particular, changes at guide-strand positions 2, 6, and 14 can influence how well the RNA supports Argonaute maturation.

The broader lesson is powerful: siRNA potency is not determined only by sequence, stability, or target accessibility. It may also depend on whether the RNA can help Argonaute fold correctly during RISC assembly.

This study gives the field a structural snapshot of a previously elusive intermediate. It shows HSP90 and p23 acting not as passive helpers, but as architectural guides. They hold Argonaute open, prevent premature collapse, and create a landing zone for duplex RNA. Once RNA binds, Argonaute can fold into a functional pre-RISC, eject the passenger strand, and become the mature silencing machine.

For RNA biology, this is a beautiful mechanistic advance.

For RNA therapeutics, it is more than beautiful. It is practical.

The AMC may become a platform for testing which siRNA designs, terminal chemistries, duplex lengths, and chemical modifications best support RISC assembly. That could move siRNA design from empirical screening toward more rational, structure-guided engineering.

RNA silencing begins with a guide strand. But this paper reminds us that before a guide can guide, the protein must be built correctly.

And behind that process stands a hidden workshop of chaperones.

Saturday, June 06, 2026

How Cells Process mRNA: Molecular Steps, Data Tools, and Disease Links

 

TheRNABlog

mRNA Processing as a System: From Nascent Transcript to Regulatory Network

A systems-biology view of how capping, splicing, 3-prime end formation, export, localization, translation, and decay work together to shape gene expression.

mRNA processing is the set of co- and post-transcriptional steps that convert nascent transcripts into mature mRNAs (capping, splicing, 3' cleavage/polyadenylation) and govern mRNA export, localization, translation and decay. A systems biology perspective treats these steps as an integrated network regulated by myriad RNA-binding proteins (RBPs) and feedback loops. High-throughput assays (e.g. RNA-seq, CLIP-seq, NET-seq, long-read and single-cell RNA-seq, ribosome profiling) have illuminated the genome-wide architecture and dynamics of this network. Quantitative models (deterministic ODEs, stochastic simulations, network models) capture aspects like splice-site selection and noise in gene expression. These approaches reveal how regulatory circuits and RNA modifications (e.g. m^6A) interconnect processing steps. Disruptions of mRNA processing underlie developmental programs and diseases (cancer, neurodegeneration, viral infection) by altering isoforms or global mRNA flux. We review the scope of mRNA processing, its molecular mechanisms, regulatory networks, and modeling/data frameworks. Key databases (Ensembl, ENCODE, GEO) and tools (alignment, CLIP analysis, network inference) are surveyed. Comparative and evolutionary trends in splicing diversity are considered (e.g. >60% of plant genes are alternatively spliced). Finally, we highlight open questions (e.g. integrating spatial/temporal data, modeling multi-step coupling) and future directions (e.g. single-cell isoform mapping, machine learning for RBP networks).

Scope and Definitions

The mRNA processing pathway comprises all steps from transcription to translation that shape an mRNA's sequence, localization, and lifespan. These include 5' capping, pre-mRNA splicing (removing introns), 3' end cleavage and polyadenylation, nuclear export, subcellular localization, translation, and mRNA decay. We focus on eukaryotic mRNAs (no specific organism assumed), noting that details vary (e.g. yeast has few introns, plants often use intron retention). In "systems" terms, we view processing not as isolated reactions but as a network of modules linked by shared factors and feedback. For example, the "exon junction complex" deposited by splicing influences both export and surveillance (nonsense-mediated decay). RNA-binding proteins (RBPs) often act at multiple steps, creating interlocking regulatory circuits. The processing network is thus hierarchical: transcription factor cues and chromatin impact splicing, splicing factors regulate export, and exported mRNAs may in turn regulate transcription factors, etc. This post-transcriptional regulatory network complements transcriptional networks and is crucial for cellular homeostasis and response.

Several authoritative resources define these processes. 5' capping is done by RNA triphosphatase/guanylyltransferase (RTC) during transcription initiation, enabling subsequent splicing and translation. The spliceosome (major and minor) removes introns and ligates exons; ~75% of human genes produce >=2 isoforms. Cleavage/polyadenylation at a poly(A) signal finishes the transcript and commits it to export. Quality-control pathways (e.g. the nuclear exosome) degrade aberrant RNAs (e.g. unspliced or with premature stops). We assume a generic eukaryotic cell by default; when examples specify, we note the organism (e.g. human ENCODE data or plant studies). Throughout, we integrate insights from genome-wide studies and database resources (Ensembl for annotations, ENCODE for RBP binding, GEO for data sets) to paint a comprehensive picture of the mRNA processing system.

Key Molecular Processes

5' Capping

Immediately after transcription initiation, the nascent RNA's 5' end is modified: a 7-methylguanosine cap is added by the capping enzyme complex. This cap protects the RNA and recruits factors for splicing and export. Systems studies show that co-transcriptional capping is tightly coupled to RNA Pol II's C-terminal domain (CTD) phosphorylation state. The cap-binding complex (CBC) remains bound through splicing and export, linking 5' capping to downstream steps. Defective capping leads to rapid decay.

Splicing

Pre-mRNA splicing is a hierarchical regulatory network mediated by ~200 proteins (snRNPs, SR/hnRNP proteins) that recognize splice sites and auxiliary elements. Spliceosomal assembly often occurs co-transcriptionally (influenced by Pol II speed and chromatin) and is regulated by combinatorial RBP binding. Global surveys (microarrays and RNA-seq) reveal pervasiveness: "~75% of human genes encode two or more splice isoforms". Alternative splicing (AS) creates transcript diversity by including/excluding exons, and is highly tissue-specific and signal-responsive. For example, neuronal RBPs like Nova and Rbfox mediate brain-specific splicing patterns. The "splicing code" - the set of cis-regulatory motifs and RBPs - has been studied via motif analyses and perturbations. A useful systems framework includes: (1) cataloging isoforms; (2) mapping splicing regulatory elements; (3) linking trans-acting RBPs to target networks; (4) integrating splicing with transcription and mRNP export; (5) relating splicing changes to signaling and disease. Recent large-scale studies follow these directions. For instance, systematic knockdown of >300 splicing regulators in human cells revealed specialized splicing networks and "extensive regulatory potential" of core spliceosome components - in other words, even core snRNPs have gene-specific regulatory roles. Thus, splicing is not simply constitutive; it is embedded in feedback loops (e.g. splicing factors auto-regulate their own pre-mRNAs), and networks of SR proteins/hnRNPs act akin to gene regulatory networks (see Table 3).

3' Cleavage and Polyadenylation

Termination of transcription is coupled to endonucleolytic cleavage and poly(A) tail addition. Core factors (CPSF, CstF, PAP) recognize the AAUAAA motif and downstream elements. Polyadenylation defines the mRNA's 3' end and influences stability and translation. Alternative polyadenylation (APA) is widespread: many genes have multiple cleavage sites, yielding mRNAs with different 3' UTRs or coding sequences. APA can be developmentally regulated and is influenced by the same RBPs that govern splicing. For example, some SR proteins and Nova also affect poly(A) site choice. Systems analyses show APA can alter networks (e.g. by changing miRNA binding sites in 3'UTRs). Viral factors can disrupt 3' processing: influenza NS1 binds CPSF30 and HSV-1 ICP27 blocks CPSF assembly, causing genome-wide readthrough transcription and host shut-off. (Viruses selectively spare their own mRNA processing.)

Nuclear Export

Processed mRNAs are packaged into mRNPs and exported through the nuclear pore. The NXF1/TAP pathway (often via the TREX complex and exon junction complex) is the primary route; CRM1/Exportin also handles some messages. Export is selective: only properly capped, spliced, polyadenylated RNAs bound by export adaptors can exit. For instance, the exon-junction complex (EJC) deposited on spliced mRNAs facilitates recruitment of export factors. Regulatory feedback exists: efficient export can affect Pol II recycling and gene looping, and conversely, transcription rates influence export kinetics. High-throughput fractionation studies (nuclear vs cytoplasmic RNA-seq) quantify export rates genome-wide; transcripts with suboptimal processing are enriched in the nucleus.

Localization

Once in the cytoplasm, many mRNAs are actively localized via interactions with transport granules and motor proteins. mRNA localization is crucial in development (e.g. embryonic axes, neuronal synapses). RBPs that bind 3'UTR "zipcodes" mediate transport; example: the beta-actin mRNA zip code binds ZBP1 to target cell protrusions. Systems-level data (e.g. spatial transcriptomics) show clustering of localized mRNAs encoding functionally related proteins. Localization and local translation form a regulatory loop: localized mRNA recruits translation machinery in situ, and translationally repressed granules may store RNAs until signals release them.

Translation Coupling

Translation often begins in the cytoplasm after export. It is coupled to earlier processing steps via RNP components. For instance, poly(A) tail length and binding of PABP enhance translation initiation; conversely, poor splicing can trigger nonsense-mediated decay (NMD) once translation terminates. The exon junction complex (EJC) left on mRNAs after splicing licenses proper translation but flags premature stops for NMD. Recent studies also suggest feedback from translation to RNA fate: stalled ribosomes can trigger mRNA decay (no-go decay) and influence nuclear events. Ribosome profiling (Ribo-seq) provides snapshots of translation genome-wide, allowing direct comparison of transcript and protein production (see High-Throughput Data below). In summary, the lifecycle of an mRNA is cyclical - its translation feeds back to decay and indirectly to re-initiation of transcription through gene looping (in yeast and some metazoans).

Regulatory Networks and Feedback

mRNA processing is governed by networks of RBPs and feedback loops that integrate cellular signals. RBPs are often multi-functional: large eCLIP maps show that many RBPs participate in more than one post-transcriptional process; for example, the Nova protein controls both alternative splicing and APA. The ENCODE eCLIP project mapped thousands of RBP-RNA binding sites, enabling the reconstruction of a genome-wide post-transcriptional regulatory network. They found RBPs connect diverse processes - splicing, polyadenylation, stability, localization and translation - into a unified system.

Feedback is built in at multiple levels. Auto-regulation: Many splicing factors regulate their own transcript splicing to maintain homeostasis (e.g. SR proteins and hnRNPs often splice-out poison exons in their own genes). Cross-talk: Splicing can influence transcription: Pol II pausing is affected by nearby splice signals, and conversely, transcription factors can recruit splicing factors. RNA surveillance loops: Faulty mRNAs are degraded, but NMD factors (UPF1/2) can also regulate the expression of splicing regulators. Signaling integration: Kinase signaling (e.g. SR protein phosphorylation by SRPK or CLKs) dynamically alters RBP binding, thus globally reshaping splicing networks in response to external cues.

Systems analyses often use network models to capture these interactions. For example, transcriptome-wide splicing networks have been inferred by perturbing RBPs or splicing factors and observing co-splicing changes. A systematic knockdown study performed systematic knockdowns of 305 spliceosome components, revealing specialized sub-networks for different core proteins. Similarly, RBP-RNA networks can be modeled as graphs where edges represent regulation of mRNA stability or translation; computational frameworks (e.g. Bayesian networks, correlation networks) have been applied to CLIP and RNA-seq data to predict novel RBP targets. In summary, mRNA processing is subject to rich regulatory architecture: cellular context and signaling modulate the components (RBPs, splice sites, polyA signals), which in turn feed back on mRNA fate. Table 3 lists key RBPs and complexes and their roles.

Quantitative Models of mRNA Dynamics

Mathematical modeling provides insights into mRNA processing kinetics and noise. Two broad approaches are deterministic vs stochastic models:

Deterministic (ODE) models assume continuous concentrations and mass-action kinetics. They are useful for average-case dynamics (e.g. average splicing rate, mRNA half-life). For instance, one can model transcription and splicing as sequential first-order reactions. These models scale well to genome-scale networks but neglect noise.

Stochastic models (Gillespie algorithms) incorporate discrete molecular events and noise, important when key factors are in low copy (e.g. a gene transcribed in bursts). Such models can capture cell-to-cell variability in mRNA levels and alternative isoforms. They often predict distributions of mRNA counts and can incorporate probabilistic splicing errors.

Kinetic models specifically characterize step-specific rates. For example, computational kinetic modeling of individual splice sites (with measured splicing half-lives) has revealed that splicing of long introns can take minutes, influencing co-transcriptional coupling. Models have also been used for polyadenylation site choice, where competition between sites is modeled as a rate process controlled by motif strength and RBP availability.

Network models abstract interactions qualitatively (Boolean or graph models). For example, RNA-protein interaction networks predict the effect of perturbing an RBP on downstream mRNA targets. Machine-learning models (deep learning) now attempt to predict splicing from sequence (SpliceAI) or to integrate multi-omic data (transcriptome + proteome).

Each modeling approach has trade-offs (Table 1). Deterministic models are computationally efficient but ignore noise; stochastic models are realistic but can be intractable for genome-scale. Kinetic models require many rate constants (often unknown). Logical or network models simplify complex networks but sacrifice dynamic precision. In practice, hybrids are used: e.g. deterministic ODEs for abundant components, stochastic for rare regulators, or coarse-grained network inference supplemented by detailed kinetics for key modules.

Table 1. Modeling approaches used for mRNA processing

Model type

Assumptions

Scale/Application

Strengths

Limitations

ODE (Deterministic)

Continuous concentrations, mass-action

Whole-cell averaged mRNA dynamics

Simple, analyzable; good for large-scale modeling of transcript abundance

Neglects molecular noise; requires parameter values

Stochastic (Gillespie)

Discrete events, random timing

Single-cell/molecule level

Captures cell-to-cell variability and low-copy effects

Computationally intensive for large networks

Kinetic (Compartmental)

Multi-step reaction rates

Single-gene or pathway kinetics

Can incorporate measured rates, good for detailed kinetics (e.g. splicing time)

Many parameters; often limited to one or few genes

Network/Boolean

Binary states or probabilities; qualitative

Regulatory network structure

Identifies key regulators and topology; integrates multi-omic data

No temporal dynamics; loses quantitative detail

Machine Learning

Data-driven; learns patterns

Isoform prediction, RBP binding

Captures complex, nonlinear patterns; uses big data

Requires large training sets; interpretability issues

High-Throughput Data Types and Analysis

Advances in sequencing and imaging have generated diverse datasets to probe mRNA processing globally (Table 2). Key technologies include:

Bulk RNA-seq (short reads): Measures transcript abundance and alternative splicing genome-wide. Typical output: tens to hundreds of millions of reads (e.g. Illumina). Resolution: exon or junction-level quantification. Analysis tools include aligners (STAR, HISAT), quantifiers (Salmon/Kallisto), and splicing tools (rMATS, LeafCutter). RNA-seq reveals gene expression, isoform ratios, and allelic or condition-specific splicing.

CLIP-seq (e.g. HITS-CLIP, iCLIP, eCLIP): Maps RBP-RNA interactions in vivo. Crosslinked RNA-protein complexes are immunoprecipitated and sequenced. Typical output: tens of millions of reads per RBP; resolution down to ~30nt footprints. Analysis identifies binding sites and motifs. ENCODE's enhanced CLIP (eCLIP) has catalogued binding for hundreds of RBPs.

NET-seq / GRO-seq: Captures nascent transcripts associated with active Pol II, mapping transcription and co-transcriptional splicing at nucleotide resolution. NET-seq (Native Elongating Transcript sequencing) provides single-nucleotide profiles of elongating Pol II, useful for studying splicing kinetics and polymerase pausing.

Long-read RNA-seq (PacBio, Oxford Nanopore): Reads >1 kb, often full-length transcripts. Allows direct observation of complete isoforms, concatenated splicing and poly(A) choices, and even base modifications (e.g. m^6A) in single molecules. Nanopore direct RNA sequencing has been used nanopore direct RNA sequencing to map full-length Arabidopsis mRNAs, revealing combinatorial diversity of TSS, splicing, poly(A) site, and tail length. Though lower throughput than short reads, long reads resolve complex isoforms and link events.

Single-cell RNA-seq (scRNA-seq): Profiles gene expression in thousands of cells, often with limited isoform resolution. Recent methods aim to capture isoforms: Smart-seq (full-length) vs 10x Genomics (3' end). Emerging single-cell isoform sequencing (scISO-seq) uses long reads on single-cell cDNA. These methods reveal cell-type-specific splicing programs and stochastic isoform variation.

Ribosome Profiling (Ribo-seq): Sequencing of ribosome-protected fragments provides codon-resolution maps of translation. It quantifies translation efficiency of each mRNA and can detect translated non-canonical ORFs. Comparison of Ribo-seq and RNA-seq yields direct coupling between transcript levels and protein synthesis.

Each data type has trade-offs (Table 2). For example, short-read RNA-seq is high-throughput and quantitative but fragments transcripts; long-read sequencing resolves isoforms but with lower depth and higher error rate. CLIP requires high quality antibodies and complex analysis.



Table 2. High-throughput data types for mRNA processing

Technology

Resolution

Throughput

Typical Outputs

Bulk RNA-seq (Illumina)

~30–150 bp reads; maps exons/junctions

High (10^7–10^8 reads/sample)

Transcript/gene expression; exon/junction counts; isoform abundance

Single-cell RNA-seq

Gene-level (3′-bias or full-length)

10^3–10^5 cells per run

Gene expression per cell; limited isoform info; cell clusters and states

Long-read RNA-seq (ONT/PacBio)

Full-length transcripts (kb)

Moderate (10^5–10^6 reads)

Complete isoform sequences; splicing patterns; poly(A) tails; base modifications

CLIP-seq (HITS/iCLIP/eCLIP)

~20–50 nt protein footprints

~10^7 reads per RBP

RBP binding sites (genome coordinates); binding motifs; RNA network maps

NET-seq/GRO-seq

Nucleotide resolution (nascent RNA)

Moderate

Pol II occupancy; co-transcriptional splicing events; pause sites

Ribosome Profiling

Codon-resolution (~30 nt footprints)

~10^7 reads/sample

Ribosome density on mRNAs; translated ORFs; translation efficiency

Ribo-Zero/PolyA-Seq

Genome/transcript end maps

High (10^7 reads)

Polyadenylation site locations (PolyA-Seq); non-polyadenylated transcripts (Ribo-Zero RNA-seq)

 

In data analysis, computational pipelines integrate these assays. For example, ENCODE/GEO repositories house thousands of RNA-seq and CLIP experiments. Bioinformatics tools (e.g. HTSeq, DESeq2 for RNA-seq; CLIPper, PureCLIP for CLIP) are used to quantify and statistically test processing differences. Machine learning and network inference tools (e.g. MEME, RBPmap, SpliceAI) aid motif discovery and splicing prediction. We recommend Ensembl/GENCODE for transcript annotation, and GEO/ArrayExpress to access relevant datasets.

Computational Tools and Databases

A multitude of software tools and databases support systems-level mRNA processing research. Key examples include:

Transcriptome annotation: Ensembl, GENCODE, and RefSeq curate gene models including splicing isoforms and poly(A) sites. These provide essential reference transcripts for mapping reads.

Sequence alignment: STAR and HISAT2 are splice-aware RNA-seq aligners; Salmon and Kallisto perform rapid transcript quantification by pseudo-alignment. For long reads, minimap2 aligns full-length cDNAs.

Splicing analysis: Tools like rMATS, SUPPA2, and LeafCutter identify differential splicing from RNA-seq data. The database VAST-DB compiles alternative splicing in vertebrates and tissues. RBPmap and ATtRACT provide RBP binding motif annotations.

CLIP analysis: PureCLIP, Paralyzer, and CLIPper call binding sites from CLIP-seq data. Databases like POSTAR and doRiNA aggregate CLIP results across RBPs and species.

3'-end processing: TAIL-seq analysis pipelines measure poly(A) tail lengths; APAlyzer and DaPars detect alternative polyadenylation from sequencing data. PolyA_DB and APADB catalogs APA sites.

Single-cell tools: STARsolo, CellRanger, and kallisto|bustools process scRNA-seq. For single-cell splicing, SpliZ and Velocyto estimate isoform variability.

Databases: The Gene Expression Omnibus (GEO) and EMBL-EBI ArrayExpress archive raw RNA-seq and CLIP-seq datasets. ENCODE and modENCODE portals provide richly annotated RBP binding and expression data. Domain-specific DBs include RBPDB (RNA-binding protein database) and doRiNA (database of RBP targets).

For network analysis, frameworks like WGCNA (for co-expression) and Graphia (for gene networks) can integrate multi-omic layers. Tools such as Cytoscape visualize RBP-RNA networks. Emerging platforms (e.g. EnrichRBP) automate integrative analysis of RBP function. Collectively, these computational resources enable reconstruction and interrogation of mRNA processing systems from diverse data.

Cross-Species and Evolutionary Perspectives

mRNA processing exhibits both conserved machinery and species-specific innovations. All eukaryotes perform capping, splicing, polyadenylation and export, but genome architectures differ markedly. Simple eukaryotes (yeasts) have few introns and limited alternative splicing, whereas multicellular eukaryotes show extensive AS. For instance, over 60% of Arabidopsis intron-containing genes are alternatively spliced, reflecting complex gene regulation in plants. Mammals and insects also have high AS rates; the Drosophila Dscam gene famously can produce thousands of isoforms. In contrast, yeast introns are rare and mostly constitutive.

Comparative genomics reveals that the core processing factors (snRNP proteins, CPSF, export factors) are broadly conserved, implying an early origin. However, the regulatory layers have expanded in complex organisms. Many RBPs present in vertebrates have no yeast homologs. Cross-species CLIP studies show some splicing regulators have conserved targets (e.g. SR proteins bind purine-rich motifs in animals and plants), but the bulk of AS patterns diverge with species. Evolutionary analyses indicate that many tissue-specific splice events are rapidly evolving, while core housekeeping splicing is conserved.

Polyadenylation signals (AAUAAA) are nearly universal in metazoans, though plants use A-rich variants. The coupling between splicing and 3' end processing is ancient: even plants show coordination. mRNA localization signals and RBPs (like zipcode-binding proteins) vary by lineage - for example, vertebrate neurons rely on different zip codes than yeast, which has simpler transport needs.

These differences have functional consequences. Alternative splicing and APA have been proposed to contribute to species diversity without increasing gene number. In development, organisms exploit these mechanisms differently: e.g. vertebrate embryogenesis involves extensive AS changes, while in Arabidopsis stress responses trigger specific splice variants. Systems studies often compare transcriptomes across species to identify lineage-specific regulatory networks. Future work in comparative epitranscriptomics (e.g. mapping m^6A across species) will further illuminate evolutionary trajectories of mRNA processing.

Roles in Development and Disease

Proper mRNA processing is essential for normal development and physiology. During development, regulated AS and APA create protein isoforms tailored to cell types. Examples include neuron-specific isoforms of neurotransmitter receptors and developmental stage shifts in 3'UTR length (longer UTRs in early embryogenesis, shorter in differentiating cells). RBPs like CELF, PTBP, and Hu proteins show developmental regulation, ensuring stage-specific splicing patterns.

Cancer: Many cancers exhibit mis-splicing and APA changes. Mutations in splicing factor genes are common in myeloid leukemias (e.g. SF3B1, U2AF1) and seen in solid tumors (TCGA analyses). Aberrant splicing can activate oncogenes or inactivate tumor suppressors. For instance, intron retention or exon skipping in apoptosis regulators can promote survival. APA shifts in cancer often truncate 3'UTRs, escaping miRNA repression and increasing oncogene translation. Large surveys (e.g. Kahles et al. 2018) show pan-cancer splicing signatures and RBP expression changes linked to tumor type. Targeting splicing (splice-switching oligonucleotides or SF3B inhibitors) is an emerging therapeutic strategy.

Neurodegeneration: Neurons heavily depend on mRNA processing. Mutations in RBPs (TDP-43, FUS, hnRNPA1) cause ALS/FTD; these proteins normally regulate neuronal splicing and RNA transport. Tau exon 10 mis-splicing underlies frontotemporal dementia. Widespread splicing dysregulation is observed in Alzheimer's and Parkinson's brains. mRNA localization is also critical in neurons - defects in localizing synaptic mRNAs can impair connectivity and learning.

Developmental and other disorders: Defects in core processing factors cause congenital diseases. For example, mutations in the U4atac snRNA (minor spliceosome) cause microcephalic osteodysplastic primordial dwarfism. Poly(A) signal mutations (e.g. FOXP3 AAUAAA→AUAAAG) lead to immunodeficiency. In viral infection, host mRNA processing is actively disrupted: as discussed above, viral proteins block cleavage/polyadenylation or even accelerate host mRNA decay to evade immunity. Some viruses rely on alternative splicing (e.g. HIV's multiple proteins from one transcript) or use unique poly(A) strategies (adenovirus uses very short poly(A) tails).

Single-gene disorders: Many monogenic diseases involve splicing errors (e.g. cystic fibrosis DeltaF508 creates an aberrant splice site; spinal muscular atrophy is due to SMN2 exon 7 skipping). Clinically, antisense therapies that redirect splicing (e.g. Spinraza for SMA) demonstrate the power of targeting this system.

Experimental and Modeling Gaps, Open Questions

Despite advances, significant gaps remain in our systems-level understanding. Integration across scales is incomplete: we lack unified models linking transcription dynamics to cytoplasmic translation outcomes. For example, how exactly does transcriptional bursting propagate to splicing noise and then to protein levels? Spatial context is underexplored: live-cell imaging (e.g. MS2 tagging of mRNA) shows granule assembly and transport, but genome-wide integration of spatial data (MERFISH or seqFISH of isoforms) is in its infancy. Single-cell complexity: while scRNA-seq profiles expression, single-cell isoform sequencing (long-read or linked reads) is just emerging. How heterogeneous is splicing within a "cell type"? Existing single-cell datasets often miss isoform-level detail, creating an analysis gap.

On the regulatory side, functional relevance of RBP binding sites is not fully known. CLIP maps hundreds of thousands of sites, but most lack characterized function. We need perturbation screens (e.g. saturating mutagenesis of UTRs) to link binding to outcome. Feedback mechanisms (e.g. how poly(A) tail length influences nuclear fate) need more quantitative data. Additionally, post-transcriptional modifications (m^6A, m^5C) are known to affect processing and stability, but the global networks of "writers, readers, erasers" in context of processing are still being mapped.

Modeling-wise, parameterization is a bottleneck. Many kinetic models assume constant rates, but in vivo rates vary by context. Direct kinetic measurements (e.g. metabolic labeling and nascent RNA-seq) provide some data, but integrating these into genome-scale models is challenging. Complex feedback loops pose theoretical challenges: for example, coupling of transcription termination with splicing through Pol II requires multi-scale simulation (chromatin, polymerase, RNP assembly) that current models cannot fully capture.

Finally, data biases and noise are issues. Short-read RNA-seq can misassign isoforms, and CLIP has false positives. Standardizing experimental protocols (e.g. benchmarks in CLIP-seq) and integrating replicates is ongoing. In summary, we need better data integration frameworks, more direct measurements of processing kinetics, and novel assays (e.g. simultaneous long-read sequencing of DNA, RNA, and proteins in single cells).

Future Directions and Recommendations

Looking ahead, multimodal single-cell technologies promise to revolutionize the field. Techniques combining long-read sequencing with single-cell resolution, or linking epigenetic state to transcript isoforms, will reveal cell-type-specific RNA processing landscapes. For example, single-cell nanopore RNA-seq is emerging. Integrating spatial transcriptomics (e.g. FISSEQ, MERFISH) with isoform resolution will map processing in tissue context, crucial for development studies.

Machine learning and data integration will grow in importance. Deep learning models (like SpliceAI) are already predicting splicing from sequence; expanding these to multi-step processing predictions (incorporating motifs, RBP expression, modifications) is a goal. Network inference algorithms that combine CLIP, expression, and phenotype data (e.g. CRISPR screens of RBPs) can build more accurate regulatory maps.

Experimentally, CRISPR-based screens targeting RBP binding sites or splice sites at scale will clarify functional networks. RNA-structure methods such as DMS-MaPseq and Nano-DMS-MaP and enhanced CLIP variants will improve our view of RNA secondary structure in vivo, informing processing mechanisms.

Finally, therapeutic targeting of the mRNA processing machinery is a growing frontier. Engineered RBPs and small molecules that modulate splicing, including SF3B-targeting compounds such as H3B-8800, have reached clinical testing. Understanding mRNA processing networks at systems level will better predict off-target effects of such interventions.

Table 3. Key RNA-processing regulators

Factor/Complex

Role in mRNA Processing

Capping enzymes (RNGTT, RNMT)

Add and methylate 5′ cap; recruit cap-binding proteins.

Spliceosome snRNPs (U1, U2, U4/U6, U5 complexes)

Core machinery for intron removal. Recognizes splice sites.

SR proteins (SRSF1-12)

SR-rich splicing factors; promote exon recognition and alternative splicing.

hnRNP proteins (hnRNP A/B, C, D, etc.)

Splicing repressors, often compete with SR proteins to regulate splice choice.

Polyadenylation factors (CPSF subunits, CstF, CFIm)

Recognize poly(A) signals; cleave pre-mRNA and recruit poly(A) polymerase.

Poly(A) polymerase (PAP)

Catalyzes poly(A) tail addition.

Poly(A) binding proteins (PABPN1, PABPC)

Bind poly(A) tails; regulate translation and tail length.

Nuclear export factors (NXF1/TAP, REF/Aly)

Mediate mRNP export through nuclear pore. Coupled to splicing via the TREX complex.

RNA decay enzymes (DCP2/DCP1 decapping, XRN1 exonuclease, exosome complex)

Remove cap or degrade from ends; perform quality control and mRNA turnover.

Regulatory RBPs (ELAVL/Hu proteins, FMRP, TIA1)

Bind specific sequences (e.g. AU-rich or G-quartets) to modulate stability, localization or translation.

Nonsense-mediated decay (NMD) factors (UPF1, SMG1)

Trigger decay of aberrant transcripts with premature stop codons; links to splicing (EJC-dependent).



Suggested figure: a lifecycle flowchart showing co-transcriptional capping, splicing, and polyadenylation in the nucleus; export through the nuclear pore; and cytoplasmic localization, translation, and decay, with RBPs and m6A marks acting across multiple stages.

Selected References

Core mechanisms and reviews

Rules of engagement: co-transcriptional recruitment of pre-mRNA processing factors. Current Opinion in Cell Biology, 2005. https://pubmed.ncbi.nlm.nih.gov/15901493/

Global analysis of mRNA splicing. RNA, 2008. https://pubmed.ncbi.nlm.nih.gov/18083834/

Transcriptional termination in mammals: Stopping the RNA polymerase II juggernaut. Science, 2016. https://doi.org/10.1126/science.aad9926

Modulation of mRNA 3-prime-End Processing and Transcription Termination in Virus-Infected Cells. Frontiers in Immunology, 2022. https://www.frontiersin.org/articles/10.3389/fimmu.2022.828665/full

Complexity of the Alternative Splicing Landscape in Plants. The Plant Cell, 2013. https://academic.oup.com/plcell/article/25/10/3657/6099545

Nanopore direct RNA sequencing maps the complexity of Arabidopsis mRNA processing and m6A modification. eLife, 2020. https://elifesciences.org/articles/49658

RBP networks and high-throughput assays

Principles of RNA processing from analysis of enhanced CLIP maps for 150 RNA binding proteins. Nature, 2020. https://pubmed.ncbi.nlm.nih.gov/32252787/

CLIP and complementary methods. Nature Reviews Methods Primers, 2021. https://doi.org/10.1038/s43586-021-00018-1

Transcriptome-wide splicing network reveals specialized regulatory functions of the core spliceosome. Science, 2024. https://pubmed.ncbi.nlm.nih.gov/39480945/

eCLIP Data Standards. ENCODE Project, accessed 2026. https://www.encodeproject.org/eclip/

Nano-DMS-MaP allows isoform-specific RNA structure determination. Nature Methods, 2023. https://www.nature.com/articles/s41592-023-01862-7

Modeling, tools, and databases

Stochastic gene expression and its consequences. Cell, 2008. https://pmc.ncbi.nlm.nih.gov/articles/PMC3118044/

Predicting Splicing from Primary Sequence with Deep Learning. Cell, 2019. https://doi.org/10.1016/j.cell.2018.12.015

EnrichRBP: an automated and interpretable computational platform for predicting and analysing RNA-binding protein events. Bioinformatics, 2025. https://academic.oup.com/bioinformatics/article/41/1/btaf018/7953276

GENCODE: The GENCODE Project. GENCODE, accessed 2026. https://www.gencodegenes.org/pages/gencode.html

Ensembl annotation. Ensembl, accessed 2026. https://grch37.ensembl.org/info/genome/genebuild/index.html

Gene Expression Omnibus. NCBI, accessed 2026. https://www.ncbi.nlm.nih.gov/geo/

Disease and therapeutic context

Comprehensive Analysis of Alternative Splicing Across Tumors from 8,705 Patients. Cancer Cell, 2018. https://pmc.ncbi.nlm.nih.gov/articles/PMC9844097/

Phase I First-in-Human Dose Escalation Study of the oral SF3B1 modulator H3B-8800 in myeloid neoplasms. Leukemia, 2021. https://www.nature.com/articles/s41375-021-01328-9

FDA approves first drug for spinal muscular atrophy. U.S. FDA, 2016. https://www.fda.gov/news-events/press-announcements/fda-approves-first-drug-spinal-muscular-atrophy

Nusinersen, an antisense oligonucleotide drug for spinal muscular atrophy. Nature Neuroscience, 2017. https://www.nature.com/articles/nn.4508

Tuesday, May 12, 2026

Why AlphaFold transformed protein biology, while RNA structure prediction remains one of biology’s most stubborn frontiers

 

The winner of the CASP14 protein-structure-prediction challenge was announced: AlphaFold, developed by Google DeepMind. The result was not merely better than previous tools. It was dramatically better. AlphaFold showed that artificial intelligence could predict many protein structures with near-experimental accuracy, solving a problem researchers had been chasing for decades.
TheRNABlog

RNA Function Follows Form — But RNA Refuses to Sit Still

The winner of the CASP14 protein-structure-prediction challenge was announced: AlphaFold, developed by Google DeepMind. The result was not merely better than previous tools. It was dramatically better. AlphaFold showed that artificial intelligence could predict many protein structures with near-experimental accuracy, solving a problem researchers had been chasing for decades.

Protein biology had its revolution.

RNA biology is still waiting for its equivalent moment.

That is the central tension behind Diana Kwon’s Nature Technology Feature, “RNA function follows form – why is it so hard to predict?” The article captures a frustrating truth: RNA is biologically essential, structurally fascinating, and increasingly important for medicine — but predicting its shape remains far harder than many people expected. 

RNA Is Not Just a Messenger

For decades, RNA was introduced in textbooks as a middleman: DNA stores genetic information, RNA carries the message, and proteins do the real work.

That explanation is now painfully incomplete.

RNA can regulate genes, catalyze reactions, guide protein complexes, control splicing, sense metabolites, organize cellular machinery, and influence disease. Ribozymes, riboswitches, long noncoding RNAs, microRNAs, guide RNAs, viral RNAs, circular RNAs, and therapeutic RNAs all remind us that RNA is not passive.

RNA does things.

But RNA does those things because it folds.

Its biological function depends on stems, loops, bulges, pseudoknots, junctions, long-range contacts, base stacking, ion coordination, and interactions with proteins or small molecules. In RNA biology, structure is not decoration. Structure is often the mechanism.

That is why the phrase “function follows form” matters.

An RNA molecule’s sequence tells us what it could become. Its structure tells us what it is actually capable of doing.

Why Proteins Were Easier for AI

AlphaFold’s success in protein structure prediction was built on several advantages.

Proteins have been studied structurally for a long time. Thousands of high-quality protein structures were already available in the Protein Data Bank. Protein sequences also carry rich evolutionary information: if two residues change together across evolution, they may physically interact in the folded protein. AlphaFold and related tools learned from these patterns at massive scale.

RNA does not offer the same easy path.

There are far fewer experimentally solved RNA structures than protein structures. Many RNAs are small, flexible, chemically sensitive, and structurally heterogeneous. RNA often depends on magnesium ions, cellular proteins, modifications, ligand binding, and environmental conditions to fold correctly. A sequence may not point to one stable structure. It may point to a shifting ensemble.

That makes RNA a much harder target for machine learning.

A protein often behaves like a molecule trying to reach a stable folded state.

RNA often behaves like a molecule negotiating among several states.

RNA’s Flexibility Is the Problem — and the Biology

The biggest mistake is to think RNA structure prediction is simply protein structure prediction with different building blocks.

RNA has its own grammar.

It has only four standard bases, but those bases can form canonical and noncanonical interactions. Its phosphate backbone is highly charged. Its folding can depend strongly on ions. It can form alternative secondary structures. It can switch conformations when binding a metabolite or protein. It can expose or hide regulatory regions depending on context.

This is not just a computational nuisance.

RNA flexibility is often exactly how RNA works.

A riboswitch must change shape to regulate gene expression. A viral RNA may remodel itself during infection. A guide RNA must fold into a form compatible with its protein partner. An mRNA may contain structural elements that affect translation, degradation, or immune recognition.

So the goal is not always to predict one “correct” RNA structure.

The goal may be to predict a population of possible structures, then understand which one matters under a specific biological condition.

That is much harder.

AlphaFold 3 Helps, But It Does Not End the RNA Problem

AlphaFold 3 expanded the modeling landscape by predicting biomolecular complexes involving proteins, DNA, RNA, small molecules, ions, and modified residues. That is a major advance because biology rarely happens molecule by molecule in isolation. Cells are crowded with interacting systems. 

But RNA structure prediction is still not solved.

AlphaFold 3 can model RNA-containing complexes, but RNA-only folding and RNA conformational ensembles remain difficult. Many RNA structures depend on experimental constraints, secondary structure priors, or specialized RNA-focused modeling approaches. The RNA field is therefore not simply waiting for one universal model to solve everything.

It is building a different toolkit.

The New RNA AI Toolkit

Several AI-based methods are now pushing RNA structure prediction forward.

Tools such as RhoFold+ use RNA language models and deep learning to predict RNA 3D structures from sequence. RhoFold+ was trained with large-scale RNA sequence information and designed to address the data scarcity that limits RNA modeling. 

Other methods, including trRosettaRNA and trRosettaRNA2, use RNA-specific structural logic, secondary structure information, and deep learning to improve 3D prediction. Recent work on trRosettaRNA2 emphasizes the value of secondary-structure-aware modeling and conformer prediction — a crucial feature because RNA often exists in multiple structural states rather than one fixed architecture. 

These methods suggest an important principle:

RNA prediction will probably not be solved by sequence alone.

It will require secondary structure priors, evolutionary signals, experimental probing data, cryo-EM maps, chemical constraints, molecular simulations, and AI models working together.

Experiments Still Matter

The rise of AI does not make RNA experiments obsolete.

It makes them more important.

Chemical probing methods such as SHAPE and DMS-based approaches can reveal which nucleotides are flexible, paired, exposed, or protected. Cryo-electron microscopy can capture larger RNA-containing complexes. NMR can reveal dynamics and local structure. X-ray crystallography can still provide atomic detail when crystals are available.

Each method sees a different part of the RNA story.

AI can propose models quickly. Experiments can test whether those models are real.

This is where RNA structure biology is heading: not toward purely computational prediction, but toward integrative structure determination.

A model becomes more trustworthy when it agrees with probing data, mutational analysis, cryo-EM density, biochemical function, and evolutionary conservation.

For RNA, evidence must converge.

Why This Matters for Medicine

RNA structure prediction is not just a technical problem for structural biologists. It is becoming central to biotechnology and medicine.

mRNA vaccines, siRNA drugs, antisense oligonucleotides, CRISPR guide RNAs, aptamers, ribozymes, circular RNAs, and self-amplifying RNAs all depend on folding behavior. A therapeutic RNA may fail because it folds incorrectly, exposes the wrong region, activates unwanted immune sensors, degrades too quickly, or binds inefficiently.

RNA-targeted small-molecule drugs are another major frontier. For a drug to bind RNA selectively, the RNA must present a recognizable structural pocket or motif. Without structural knowledge, RNA drug discovery becomes guesswork.

Better RNA structure prediction could help researchers design more stable RNAs, improve guide RNA performance, identify druggable RNA motifs, engineer riboswitches, and understand viral RNA elements.

In short, RNA structure prediction is not only about seeing molecules.

It is about designing them.

The Real Lesson from AlphaFold

AlphaFold changed protein biology because it made high-quality structural models widely accessible. It did not eliminate experiments, but it changed where experiments begin.

RNA needs a similar shift.

But the RNA version of that revolution may look different. It may not be a single model that predicts one final structure from sequence. It may be a network of tools that predicts secondary structures, tertiary folds, alternative conformers, RNA–protein complexes, RNA–ligand interactions, and experimentally testable structural hypotheses.

RNA does not sit still.

So RNA structure prediction must become comfortable with motion.

The next breakthrough will not simply tell us, “Here is the RNA structure.”

It will tell us:

Here are the structures this RNA can adopt. Here is when they appear. Here is what they bind. Here is how they regulate biology. Here is how we can redesign them.

That is the future RNA biology is moving toward.

Protein structure prediction had its AlphaFold moment.

RNA’s moment may be harder, slower, and messier.

But it may also be more interesting — because RNA is not merely a molecule with a shape.

It is a molecule with possibilities.


References / Sources

  1. Kwon, D. “RNA function follows form – why is it so hard to predict?” Nature, 2025. (Nature)

  2. Jumper, J. et al. “Highly accurate protein structure prediction with AlphaFold.” Nature, 2021. (Nature)

  3. Abramson, J. et al. “Accurate structure prediction of biomolecular interactions with AlphaFold 3.” Nature, 2024. (Nature)

  4. Shen, T. et al. “Accurate RNA 3D structure prediction using a language model-based deep learning approach.” Nature Methods, 2024. (Nature)

  5. Wang, W. et al. “The trRosettaRNA server for RNA structure prediction.” Nature Protocols, 2026. (Yang Lab)

  6. CASP16 RNA structure prediction assessment and Yang-Server/trRosettaRNA2 reporting. (Wiley Online Library)