Codondex: nucleotide

Showing posts with label nucleotide. Show all posts

Saturday, January 17, 2026

Genome Balance: Repeats, Immunity, and Cancer

Cancer is usually described as a disease of mutations. Genes break, pathways fail, and cells escape control. That framing has been powerful, but it misses a deeper layer that may reveal how it begins.

The human genome is not primarily a coding genome. It is a repeat genome. More than half of our DNA consists of repetitive elements, with Alu retroelements alone numbering over a million copies. These sequences are a defining feature of primate genomes and they create a unique biological problem that human cells must continuously manage. Recent work suggests that cancer may emerge, in part, when this management system loses balance.

Alu elements are short retrotransposons that readily form double‑stranded RNA stem‑loop structures when transcribed, particularly in antisense orientation within introns and untranslated regions. To the innate immune system, these structures resemble viral RNA. This means that normal gene expression in human cells constantly risks triggering antiviral immune responses against self‑derived RNA.

A striking recent study shows that human cells rely on active suppression to avoid this outcome. In Ku suppresses RNA‑mediated innate immune responses in human cells to accommodate primate‑specific Alu expansion, the authors demonstrate that the DNA repair protein Ku (Ku70/Ku80) plays an essential second role: binding Alu‑derived dsRNA stem‑loops and preventing activation of innate immune sensors such as MDA5, RIG‑I, PKR, and OAS/RNase L.

When Ku is depleted interferon and NF‑κB signaling are strongly activated, translation is suppressed, and cells undergo growth arrest or death. Notably, Ku levels scale tightly with Alu expansion across primates, and Ku is essential in human cells but not in mice. The implication is clear:

Human cell viability depends on continuous suppression of Alu‑derived innate immune activation.

Alu expression is not harmless noise, it is actively tolerated! Ku functions as a finite buffer that allows primate cells to tolerate structurally immunogenic RNA produced by repeat‑rich genomes. When structured RNA load increases simultaneously from endogenous repeat transcription and exogenous viral RNA infection, Ku becomes functionally saturated and redistributed, weakening nuclear retention and cytoplasmic buffering. This pressurizes the cell’s capacity to contain dsRNA stress, promoting escape of repeat‑derived RNA, activation of innate sensors, and eventual selection for immune‑tolerant states.

A second line of evidence connects this tolerance to cancer evolution. A 2025 bioRxiv preprint, p53 loss promotes chronic viral mimicry and immune tolerance, shows that loss of p53 permits transcription of immunogenic repetitive elements, generating signals that resemble viral infection. Rather than leading to effective immune clearance, this state becomes chronic. Tumor cells adapt by dampening innate immune responses and tolerating persistent repeat‑derived nucleic acids.

In this view, “viral mimicry” is not a one‑time immune alarm. It is a conditioning process repeat RNAs accumulate, immune pathways are activated, and progressively suppressed or rewired to allow survival. Cancer cells do not simply evade immunity, they learn to live with endogenous viral‑like signals.

These immune findings align with earlier evidence that repeat control begins at the level of genome structure itself. A 2022 Nature Communications study demonstrated that retroelements embedded within the first intron of TP53 act as cis‑repressive genomic architecture. Removing this intron increases TP53 expression, indicating that long‑embedded repeats contribute directly to regulating a core tumor suppressor gene.

Importantly, this repression is architectural rather than motif‑driven. The repeats do not act through a single conserved sequence, but through repeat‑dense structure.

Together, these findings suggest a layered system of control:

Structural repression of repeats within introns.
Immune suppression of repeat‑derived dsRNA.
p53‑dependent governance of both genome stability and immune signaling.

One long‑standing challenge in repeat biology is inconsistency. Different tumors show different repeat fragments. Even different regions of the same tumor can look unrelated at the sequence level.

From a traditional biomarker perspective, this appears discouraging. From a structural perspective, it is expected. Codondex analyses of repeat‑dense introns, including TP53 intron 1, show that cancer does not preserve specific Alu sequences. Instead, it perturbs repeat topology:

dominance and skew within intronic scaffolds,
stem‑loop‑prone architectures,
context‑specific fragmentation patterns.

The sequences vary. The instability regime does not. This is characteristic of a state change, not a discrete genetic event. Repeat‑dense introns behave like stress recorders. They integrate replication stress, chromatin relaxation, repair pathway bias, and immune tolerance history.

Unlike coding mutations, these signals are heterogeneous, region‑specific, and reflective of ongoing cellular state.

They are difficult to interpret with gene‑centric tools, but powerful when viewed architecturally.

Most cancer diagnostics ask:

What mutation is present? A repeat‑aware framework asks:

Has this tissue entered a stable state of repeat derepression coupled with immune tolerance?

That state may precede aggressive behavior, accompany treatment resistance, or mark transitions in disease evolution. Future prognostic approaches may therefore combine repeat‑topology instability metrics, repeat RNA burden, and evidence of immune decoupling from dsRNA load. Not to identify a single driver, but to detect loss of containment.

Alu repeats do not cause cancer on their own, but human cells must continuously restrain them, structurally and immunologically. Cancer appears, at least in part, when that restraint erodes and tolerance replaces control. Introns, long treated as background, may be one of the clearest places to see this shift, not because they encode instructions, but because they actively record genomic history and project it into a measure of present state.

Tuesday, June 1, 2021

Short Sequences of Proximally Disordered DNA

Oxford Nanopore Device Reducing Sequencing Cost

Relationships exist between short sequences of proximal DNA (SSPD) of a gene that when transcribed into RNA present stronger or weaker binding attractions to RNA binding proteins (RBP'S) that settle, edit, splice and resolve messenger RNA (mRNA). Responsive to epigenetic stimuli on Histones and DNA, mRNA are constantly transcribed in different quantity, at different times such that different mRNA strands are transported from the nucleus to cytoplasm where they are translated into and produce any of more than 30,000 different proteins.

Single nucleotide polymorphisms and DNA mutations can alter SSPD combinations in different diseased cells thus altering sequence proximity, ordering that affects transcribed RNA's attraction and optimal binding of RBP's. This may result in modified splicing of RNA, assembly of mRNA and slight or major variations in some or all translated protein derived from that gene.

The specific effects of these DNA variations, on the multitude of proteins produced are generally unknown. However, it remains important to understand their effects in disease, diagnosis and therapy. Typically these have historically been researched by large scale analysis of RBP on RNA as opposed to the more fundamental, yet underrepresented massive array of diseased variant DNA to mRNA transitions.

Most pharmaceutical research is directed to a molecular interference targeting an aberrant protein to cure widely represented or highly impactful disease conditions of society. Economic assessments generally influence government decisions to support research based on loss of GDP contribution by a specific disease in a patient cohort. However, in the modern multi-omics era top down research into protein-RNA activity is descending deeper into the cell to include RNA-mRNA and mRNA-DNA customizable therapies that will eventually resolve individually assessed diseases at a price that addresses much larger array of patient needs.

SNP's and other mutations can vary considerably in cells. These variations can cause instability during division and lead to translated differences that can ultimately drive cancerous cell growth to escape patient immunity. Like a 'whack-a-mole' game, pattern variation and mechanistic persistence eventually beat the player. Without effective immune clearance these cells can replicate into tumors and contribute to microenvironments that support their existence.

Link to video on tumor microenvironment https://youtu.be/Z9H2utcnBic

We thought to analyze DNA and mRNA transcripts from cells in tumors and their microenvironments to see if we could expose the SSPD disordered combinations that may have promoted sub-optimal RBP attractions and led to sustained immune escape. Given the complexity of DNA to mRNA transcription, for any given gene many distortions in gene data sets have to be filtered. To do that we focused on p53, the most mutated gene in cancer. We designed a method to compare sequences arrays of DNA and mRNA Ensembl transcripts, from the consensus of healthy patients to multiple cell samples extracted from different sections of a patients tumor and tumor microenvironment.

We previously identified and measured different levels of Natural Killer (NK) cell cytotoxicity, produced from cocultures with the extracted samples of each of the multiple sites of a biopsy. We will measure the different p53 transcript SSPD combinations associated with each sample and determine whether disordered SSPD's corelate with NK cytotoxicity from each coculture. We expect to identify whether biopsied tumor cells, ranked by SSPD's predict the cytotoxicity resulting from NK cell cocultures. We will narrow our research to identify the varied expressions of receptor combinations associated with degrees of cytotoxicity. We will test immune efficacy to lyse and destroy tumor cells. Finally we will test for adaptive immune response.

Our vision is for per-patient, predictable cell co-culture pairings, for innate immune cell education based on ranking DNA-mRNA combinations to lead to multiple effective therapies. The falling cost of sequencing and sophistication of GMP laboratories presently servicing oncologists may support a successful use of this analytical approach to laboratory assisted disease management.

Thursday, May 13, 2021

Non-Coding DNA Key Sequences

DNA Structural Inherency

Wind two strands of elastic, eventually it will knot, ultimately it will double up on itself. Separate the strands. From the point of unwinding, forces will be directed to different regions and the separation will approximately return to the wound state of the band. Do the same with each of 10 different bands or strings of any type, they will all behave in much the same way. For a given section of DNA being transcribed, the effect of separation will be much the same. For a given gene, there will be sequences that can tolerate force to greater or lesser degrees. For different transcripts, of a gene variation at those sequences may be crucial to the integrity of transcription machinery that separates DNA strands to initiate replication to RNA and for the outcome.

Cellular biology is enormously complex in all regards. The physics of molecular interaction, fluid dynamics, and chemistry combine in a system where cause and effect is near impossible to predict. At the most elementary level we hypothesize some non-coding DNA (ncDNA) possess structural inherencies that can be deployed to direct gene proteins and cell function for diagnosis or therapy.

Coding DNA and its regulatory, non-coding gene compliment is transcribed and spliced from a transcribed gene. Transcription to RNA, edited mRNA, spliced non-coding RNA and ultimately mRNA translation to protein can produce wide ranging, variable outcomes that may not be re-captured experimentally.

A single nucleotide polymorphism (SNP) or SNP combinations within a gene may affect the finely tuned balance that results. Under different environmental conditions this could be material to the protein produced. Additionally other mutations of the gene could add complexity to the environment and/or the resulting protein translation.

At this level of cellular biology, genetic DNA stores instruction for protein assemblies to produce new protein required for the fully functional cell. However, DNA's stored mutations can lead to different functional or non-functional versions of protein depending on many different factors. Relationships between ncDNA, including mutations and the transcripts' edited, protein coding mRNA may represent unexplored inherencies that can regulate the gene's mRNA or translated protein.

We built an algorithm to elaborately compare ncDNA sequences of multiple protein coding transcripts of the same gene. For each transcript it steps through every variable length ncDNA sequence (kmer) (specifically intron1), computes a signature for each and indexes it to the constant of the transcripts' mRNA signature. For each step these signatures order the kmers for each of the transcript's. The order is represented in a vector of all the transcripts being compared.

At millions of successive steps (depending on total intron 1 length's) transcripts mostly retain their vector ordering except, as expected at a kmer length change. Mostly transcript order in the vector does not change, occasionally a few positions change, vary rarely do all positions change. Position changes that cause another, like a domino effect are filtered out. For the rarest positions changes at a step, we look to the root causes in the kmer (sequence). We call this a Key Sequence because it is identified by the significance of changes to transcript positions in the vector compared to the vector at the next step.

Therefore, Key Sequences cause the most position changes between transcripts being compared by the algorithm. This relative measure is step dependent and Key Sequences are discovered by comparing transcript positions in the vector at the next step location. Logically, this infers a genes structural inherency discovered through ncDNA Key Sequence relationships to mRNA, to other transcripts, error in gene alignments, sequenced reads or the algorithm.

In assay testing we were able to predict and synthesize non-coding RNA Key Sequences that significantly reduced proliferation of HeLa cells. In our pre-clinical work, based on comparisons to transcripts of the TP53 we will be predicting the efficacy of cell and tissue selections that educate and activate Natural Killer cells.

If Key Sequences are inherent they could open a new frontier for diagnosis and therapy.

Codondex