Showing posts with label k-mer. Show all posts
Showing posts with label k-mer. Show all posts

Saturday, August 8, 2026

Do p53 Repeat Fields Synchronize Target Cells and Natural Killer Cells?


Natural Killer cells face a remarkable biological problem. They must continuously distinguish healthy cells from stressed, infected or malignant cells and decide which should be destroyed. Increasing evidence suggests that p53 participates directly in coordinating that decision. A 2024 review in Frontiers in Immunology brings together evidence that p53 affects not only the internal fate and external immune visibility of the target cell, but also the homeostasis, receptor signalling and functional state of the Natural Killer cell itself.

This creates an intriguing question for Codondex. If p53 coordinates several different components of the NK-target interaction, could the unusually dense repeat structures that Codondex detects within TP53 help identify genomic regions involved in regulating that coordination? We describe these structures as High-Density Nested Repeat Fields, or HDNRFs: regions in which repeated, overlapping and nested short sequences become unusually concentrated. HDNRFs are not simply conventional tandem repeats. They represent a broader sequence geography that becomes visible when DNA is interrogated simultaneously across different sequence lengths and positions.

The Codondex methodology approaches this problem without first needing to know the biological function of the sequence. It divides genetic sequence into overlapping components, preserves their relative positions and relationships, and identifies sequence combinations that recur or associate unusually strongly. The important distinction is that Codondex can identify unusual sequence architecture first and ask what that architecture means biologically afterwards. The progression is therefore from sequence, to pattern, to biological association and finally, if experimentally demonstrated, to mechanism.

We have already seen an intriguing example of this process within TP53. Earlier Codondex analysis converged on the intron 3 region containing the known PIN3 polymorphism, or Polymorphism Intron 3. PIN3 is a 16-base-pair duplication situated within an unusual region of TP53 associated with structural and regulatory features affecting p53 RNA processing. The importance of the PIN3 observation is not that Codondex discovered PIN3 itself. The polymorphism was already known. What was interesting was the route by which Codondex arrived there. The algorithm was not instructed to search for PIN3 or for a known p53 regulatory region. It converged on the location because of an unusual concentration of relationships within the sequence itself. Conventional biology then provided an independent reason why that location could matter.


PIN3 therefore provided an early example of the utility we are now trying to test more systematically with HDNRFs: whether unusual repeat geography can lead to biologically significant regulatory locations before we understand why those locations are important. More recent Codondex analysis has strengthened that question by revealing concentrations of repeated and overlapping sequences elsewhere within TP53 and other genes, suggesting that repeat density may represent a genomic characteristic worth measuring in its own right.

A second observation made the p53 question considerably more interesting. In our previous analysis of 48 tumour sections, Codondex-derived sequence rankings were compared with several experimental immune measurements, including Chromium-release Natural Killer cytotoxicity under unstimulated pNK, IL-2-stimulated pNK and sNK conditions, together with IFN-γ and other measurements. When the sequences were examined in different ways, including duplicate occurrence, repeated short sequences and tandem or overlapping repeat relationships. pNK repeatedly emerged as one of the more consistent candidate associations.

This was unexpected because the unstimulated pNK experimental signals themselves were frequently weak. Many tumour sections showed little or no measured pNK activity, while stimulated NK measurements often generated much stronger absolute signals. Yet pNK continued to appear when independent measures of Codondex sequence structure were compared with the experimental results. We previously discussed the possibility that this recurrence may therefore be more informative than the absolute strength of the pNK measurement suggests. A weak biological signal that repeatedly aligns with independently derived sequence characteristics may be pointing to an underlying state rather than simply reflecting the magnitude of immune activity.

The Frontiers in Immunology review provides a biological framework capable of explaining why this might occur. In the target cell, wild-type p53 can increase expression of NK-activating ligands including ULBP1 and ULBP2, which interact with NKG2D, and can influence PVR/CD155 and other components of activating versus inhibitory NK recognition. p53 also participates in the apoptotic machinery determining whether an NK attack successfully kills its target, including BAX-dependent mitochondrial pathways and death-receptor mechanisms.

But importantly, p53 is not operating only within the target. The same review describes p53-related regulation within NK cells themselves. p53 participates in NK-cell cycle and apoptotic control, influences receptor expression and regulates SAP, which couples SLAM-family signalling through Fyn and Vav-1 to NK activation. The authors explicitly separate the regulatory role of p53 in NK cells from its role in tumour cells, and conclude that these functions contribute to NK recognition and targeting of malignant cells.

This means that p53 potentially operates on both sides of the immune synapse. In the target, it can contribute to making a damaged cell visible, engageable and susceptible to killing. In the NK cell, p53-related pathways can influence whether the immune cell is functionally competent to recognize and respond to that target. Successful cytotoxicity may therefore depend upon the state of two interacting cellular programmes becoming appropriately aligned.

We have described this concept as p53-driven NK-target synchronicity. Until recently, however, the word “synchronicity” was principally a functional description of several p53-dependent variables moving together. New experimental work published in Molecular Systems Biology now gives that term a much more literal biological foundation.

Venkatachalapathy and colleagues demonstrated that p53 behaves as a dynamic cellular oscillator whose normally heterogeneous pulses can be experimentally phase-reset and synchronized across individual cells. Using time-lapse microscopy and controlled DNA-damage stimulation in MCF-7 cells, they showed that p53 oscillations could be brought into significantly greater temporal alignment by appropriately timed repeated stimuli. Two damage pulses spaced approximately 4.0 to 5.5 hours apart produced significantly increased synchronization compared with a single stimulus.

This is important because the synchronization was not merely cosmetic. Altering p53 pulse timing changed downstream gene-expression dynamics, and phase resetting affected cellular fate. The study showed that changes in p53 oscillatory frequency altered expression of p53 target genes and that greater synchronization reduced the probability that cells escaped cell-cycle arrest. In other words, p53 synchronization produced synchronized biological consequences.

This materially strengthens the hypothesis we have been developing. p53 should no longer be viewed only as a molecular switch whose biological significance is determined by how much p53 is present. Its timing, frequency and phase also contain biological information. The new work demonstrates experimentally that p53 dynamics can be reset, that cells can be brought into temporal alignment, and that changing those dynamics changes downstream molecular programmes and cell fate.

The implication for Codondex is potentially significant. A p53 HDNRF need not simply influence whether there is “more” or “less” p53. If repeat architecture has regulatory significance, it could conceivably affect p53 expression, RNA processing, response thresholds, pulse amplitude, oscillatory frequency or the persistence of downstream signals. PIN3 becomes particularly interesting in this context because Codondex independently converged on an intronic TP53 region already associated with regulation of TP53 processing. The question may therefore be broader than whether repeat fields alter p53 abundance. It may be whether they identify or influence the dynamic state of the p53 regulatory system.

This offers a new way of interpreting the repeated pNK result. Unstimulated pNK cytotoxicity may represent a relatively delicate basal relationship between the target cell and the Natural Killer cell. For killing to occur, the target must become sufficiently visible, activating signals must overcome inhibitory signals, the NK cell must be in a competent functional state, the two cells must successfully engage and the target must remain susceptible to the resulting cytotoxic attack. These conditions need not individually produce large experimental signals. Their biological importance may lie in whether they occur together.

The emerging hypothesis is therefore that the Codondex repeat signal is associated not simply with “NK activity,” but with the state in which the target and NK cell become appropriately coordinated. The target-side p53 programme can influence recognition ligands and apoptotic susceptibility, while the NK-side p53 programme can influence cellular homeostasis and activation pathways. The new Molecular Systems Biology work adds the crucial observation that p53 itself possesses a genuine temporal synchronization mechanism capable of coordinating downstream biological responses.

It is important not to claim more than the evidence demonstrates. The new study did not place an NK cell beside a tumour cell and show that p53 oscillations in the two cells phase-lock to one another. It synchronized p53 oscillations across MCF-7 cells using externally delivered DNA-damage pulses. It therefore proves that p53 is a synchronizable biological oscillator, but does not yet prove that NK and target-cell p53 oscillators synchronize during immune surveillance. That distinction defines the next experiment rather than weakening the hypothesis.

The research question can now be framed much more precisely. Tumour targets should be classified according to their relevant Codondex HDNRF state and then observed together with Natural Killer cells while p53 dynamics are measured in real time. Instead of measuring p53 only at a single time point, the experiment should measure p53 pulse timing, frequency, amplitude and phase in the target and, where technically possible, in the NK cell. These measurements should be followed simultaneously by target-cell ULBP1/2, PVR/CD155 and other recognition signals, NK activation and engagement, CD107a degranulation, IFN-γ production and ultimately target-cell killing.

The decisive observation would be whether these variables move together in time. Does a change in target-cell p53 state precede altered NK-recognition ligand expression? Does successful NK engagement occur preferentially during a particular target-cell p53 state? Does the NK cell itself undergo a corresponding p53-dependent state change? Do successful killing events show greater coordination between these trajectories than unsuccessful encounters? And most importantly for Codondex, does the strength or architecture of the relevant HDNRF predict any of these dynamic relationships?

The experimental design can then move from association to causality. Removing or inhibiting functional p53 should disturb the predicted relationship. Restoring p53 should restore at least part of it. Directly altering a candidate HDNRF using locus-specific sequence or epigenetic intervention could then determine whether the repeat field is merely a marker of the p53 state or participates in establishing it. If changing the HDNRF changes p53 dynamics and those changes propagate through NK recognition and cytotoxicity, the evidence would move substantially closer to a causal genomic mechanism.

PIN3, the recurring pNK association and the new p53 synchronization data now form three independent but potentially connected observations. PIN3 showed that Codondex could converge computationally on an unusual TP53 sequence region already recognized by conventional biology as regulatory. The tumour-section analysis repeatedly brought pNK to the surface despite weak absolute unstimulated NK signals. The new phase-resetting experiments demonstrate that p53 itself can operate as a synchronizable dynamic system whose temporal state changes downstream gene expression and cellular fate.

None of these observations alone proves that an HDNRF synchronizes a Natural Killer cell with its target. Together, however, they produce a considerably more specific and experimentally falsifiable hypothesis. A p53 repeat field may identify or influence a dynamic genomic state that helps coordinate p53-dependent programmes in the target and Natural Killer cell, increasing the probability that recognition, engagement, susceptibility and cytotoxic response occur together.

That hypothesis also provides a possible explanation for why pNK keeps appearing in the Codondex analysis. The algorithm may not be detecting the strength of an immune response. It may instead be detecting sequence architecture associated with the underlying state that permits natural NK surveillance to occur.

Codondex was designed to find relationships in genetic sequence before we necessarily know what those relationships mean. PIN3 was an early indication that concentrated sequence relationships could converge on known p53 regulatory biology. The recurring pNK evidence suggested a possible connection with Natural Killer surveillance. We now know experimentally that p53 itself is capable of phase resetting and synchronization, and that synchronized p53 dynamics can alter downstream biological outcomes.

The next question is therefore unusually clear: do p53 repeat fields participate in setting the dynamic conditions under which a target cell and a Natural Killer cell become synchronized for recognition and killing?

That is now an experiment worth doing.


Tuesday, June 2, 2026

The Hidden Topography of Gene Regulation


A gene is usually read as a linear instruction, a sequence running from promoter to exon, intron, splice junction, UTR and termination site, but Codondex suggests that a gene should also be read as chromosomal geography. Beneath the annotated map of exons and introns there is another terrain: a repeat-density topography formed by short DNA words that recur, overlap, nest inside longer words, cluster into local fields and rise into summits. These summits are not defined by conventional gene annotation. They are not necessarily exons, splice sites, enhancers or promoters. They are sequence-density formations inherent in the DNA itself. Codondex calls these nested formations High-Density Repeat Fields ("HDRF"s or HDRNF).

A HDRF is not simply a repeated sequence. It is discovered as a local field in which many short, non-trivial motifs recur through adjacent, overlapping and nested k-mer relationships. A repeated 8-mer may sit inside a repeated 9-mer, which sits inside a repeated 10-mer, which is carried through a population of longer 13–28-mers. The importance of the field is not that one short motif repeats many times in isolation. The importance is that the motif is embedded in a dense neighborhood of related repeating sequence words. Local DNA is therefore not merely repetitive; it is architecturally loaded. It carries a concentrated burden of repetitive sequence possibilities that can be read by chromatin, transcription factors, polymerase, splice machinery, RNA-binding proteins and, after transcription, by the nascent RNA environment or it inherently affect biological concentrations.

In this model, the genome is not flat text. It is landscape, and some parts of that landscape are loaded with encoded densities. For example, Introns are not empty space. They may contain ridges, basins and summits of repeat-density potential. The highest HDRF is the mountain in that landscape: the point where nested repeat architecture is most concentrated, where the gene’s internal sequence burden reaches its maximum, and where encoded DNA density may be most readily converted into biological concentration through chromatin exposure, transcription, RNA processing or synthetic mimicry. Codondex begins at that summit because the summit is where the gene has already concentrated its own sequence logic.

This is why HDRFs are best understood as chromosomal geography. A gene has valleys where nested repeat burden is low, ridges where motifs begin to cluster, plateaus where repeat families spread across local sequence, and peaks where the density of nested, overlapping, non-trivial motifs reaches a maximum. The highest peak in that landscape is the HDRF Summit: the local sequence region, Codondex represents computationally by a synthesis-length 28-mer, that carries the maximum nested-repeat burden within the gene or transcript region being analyzed.

The mountain analogy is useful because it does not overstate function. A mountain is real whether or not anyone climbs it. Likewise, an HDRF is real as sequence architecture whether or not the gene is actively transcribed at a given moment. The DNA contains the topography before transcription. Transcription does not create the field; transcription reads through it and may convert its encoded DNA density into RNA motif density. When chromatin opens, when polymerase traverses the region, when an intron is copied into pre-mRNA, when splice factors scan the nascent transcript, or when RNA-binding proteins engage the sequence, the latent geography may become regulatory opportunity.

This distinction is central. A high-frequency k-mer in a table does not automatically prove biological function. K-mer density is not itself biochemical concentration. But k-mer analysis can reveal a real feature of the genome geography: inherent sequence-density concentration. In DNA, this means an increased local density of potential interaction sites. In RNA, after transcription, the same encoded field may become a repeated motif substrate available for folding, binding, splicing, retention, decay or compartmental interaction. The biological question is therefore not whether every repeated word is functional. The stronger question is whether a gene’s highest-density nested repeat fields mark regions where regulatory potential is unusually concentrated.

This is especially important in first introns. First introns are often regulatory-rich, promoter-proximal and involved in early transcriptional architecture, chromatin accessibility, elongation and co-transcriptional processing. For example: In TP53 and MEN1, the intron 1 repeat landscapes suggest that transcript variants do not merely differ in length. They preserve different repeat-density fields. Even when transcript lengths are normalized, variant-specific clustering can remain because normalization rescales the sequence but does not erase its internal motif architecture. The gene’s repeat geography survives the scaling.

In introns of one TP53 transcript, for example, the short motif 'CCCAGCTA' emerges as a dominant repeat core. Its significance is not simply that this 8-mer appears frequently. The deeper signal is that CCCAGCTA is repeatedly nested inside adjacent and overlapping longer sequence contexts. It is surrounded by neighboring motifs that also recur. A 28-mer containing that core may therefore represent a compact summit of a broader HDRF: a local sequence unit carrying the densest accessible sample of the gene’s nested repeat architecture. The 28-mer is not chosen because 28 has mystical biological status; it is chosen because it is a practical synthetic length that can capture a local field of internal 8–12, 8–18 and 8–28 motif burden.

The computational task is therefore not merely to find the most frequent k-mer. That would overvalue trivial homopolymers and low-complexity tracts. The task is to compute the nested burden of each candidate window. For each 28-mer, Codondex sums the recurrence frequencies of all internal k-mers from length 8 to 28, with optional weighting for entropy, GC content, CpG content, palindromic potential, stem-capability, transcript conservation and non-triviality. Adjacent high-scoring 28-mers merge into a peak. The highest-scoring local maximum becomes the HDRF Summit.

This produces a different kind of gene map. Instead of asking only where the exons are, where the promoter is, or where the canonical splice junctions sit, Codondex asks: where is the gene’s highest encoded motif-density burden? Where are the repeat summits? Which short motifs form the summit core? Which adjacent motifs amplify the field? Which transcript variants carry the summit, and which exclude it? Does the summit sit in intron 1, in a UTR, near a splice boundary, inside a retained intron, in a GC-rich regulatory compartment, or in a low-complexity region that may influence chromatin rather than sequence-specific binding?

The biological implications are broad but must be stated precisely. HDRFs may contribute to regulation at the DNA level by increasing the effective local density of potential binding sites, altering DNA shape, influencing nucleosome preference, supporting chromatin-factor recruitment, contributing to methylation-associated architecture or affecting the probability of transcription-factor rebinding. They may contribute during transcription by shaping polymerase pausing, elongation or co-transcriptional splice recognition. They may contribute at the RNA level when the same density field is copied into pre-mRNA, creating repeated substrates for RNA-binding proteins, splice enhancers, splice silencers, intronic structure, R-loop tendency or RNA compartmental behavior.

The aggregate burden may also matter. A local HDRF is not isolated from the rest of the gene. A gene may contain multiple HDRF peaks, some sharing the same core motif family, some distributed across introns, some concentrated near the 5′ region, some sitting in transcript-specific compartments. The gene-level HDRF burden may shape the background geography within which the local summit operates. The summit is the highest mountain, but the surrounding range may affect its biological visibility. Context score estimates whether the summit is likely to be biologically exposed, transcribed, accessible or regulatory.

This framework also clarifies the possible role of synthetic DNA or RNA candidates. A synthetic 28-mer derived from an HDRF Summit does not reproduce the entire gene. It does not automatically carry the whole biological meaning of the chromosomal field. But it may act as a compact concentration mimic of the summit architecture. If introduced at sufficient copy number, in the correct chemical form and cellular compartment, it may present a dense version of a sequence field that the gene already carries internally. Its potential mechanism could be decoy-like, scaffold-like, guide-like, competitive, structural or binding-mediated. The hypothesis is not that any high-frequency motif will function. The hypothesis is that a summit-derived 28-mer is a rational candidate because it is selected from the strongest encoded motif-concentration point in the gene’s own geography.

HDRF geography therefore moves gene analysis away from the idea that regulation is only a list of known motifs at known annotations. It proposes that each gene carries an internal terrain of motif density. Some of that terrain may be silent, some structural, some regulatory, some transcript-specific, some disease-contextual. But the terrain exists. It can be measured. It can be ranked. It can be compared between transcript variants, genes, tissues and disease states.

In this model, the genome is not flat text. It is landscape. Introns are not empty space. They may contain ridges, basins and summits of encoded regulatory potential. The highest HDRF is the mountain in that landscape: the place where nested repeat architecture is most concentrated, where the gene’s internal sequence burden rises to its maximum, and where Codondex begins looking for the most compact representation of that hidden regulatory geography.


Thursday, March 26, 2026

When Processing, Not Presence, Determines Visibility


It is easy to assume that if a protein accumulates in a diseased cell, the immune system will eventually see it. In the case of p53, that assumption has always had an intuitive appeal. p53 is one of the central stress-response proteins in biology, frequently altered in cancer, often stabilized, and deeply woven into the molecular logic of cell fate. If any intracellular protein should become immunologically visible, it ought to be p53.

But the deeper one looks at antigen presentation, the less that simple view holds. What matters is not merely whether p53 is present. What matters is whether peptide fragments derived from p53 are generated in the right form, survive intracellular trimming, fit the preferences of a particular HLA groove, and remain stable enough on the cell surface to be interrogated by either a T cell or an NK-cell receptor system. The 2022 Codondex article, Expanding Treatment Horizons, was already moving in that direction by highlighting an underappreciated observation from the HLA-C ligandome literature: a TP53-derived peptide, TAKSVTCTY, was identified as a naturally presented ligand of HLA-C*02:02. That observation comes from Moreno Di Marco and colleagues’ immuno-peptidomics study, which also listed MAGEA3-derived peptides among ligands presented by the same allotype.

That point remains important, but it also needs sharpening. The HLA-C paper tells us that a TP53-derived peptide can be naturally presented by HLA-C02:02. It does not tell us that HLA-C02:02 is already a dominant or clinically validated p53 presentation route in the way that HLA-A02:01 has become. For that, the literature is far stronger on the HLA-A side. A substantial body of work has shown that **wild-type p53 peptides presented by HLA-A02:01**, especially the well-known p53(264–272) epitope LLGRNSFEV, can stimulate cytotoxic T-cell responses and can be recognized on tumor cells. This was shown in studies such as Chikamatsu et al. Hoffmann et al. Gnjatic et al. and later vaccine-oriented work including Svane et al and the broader review literature on p53-targeting vaccines. In other words, for HLA-A*02:01, p53 is not just a theoretical ligand source; it is already part of a fairly mature immunotherapeutic story.

The most useful contribution of the recent Nature paper, The DNA virome varies with human genes and environments, is that it sharpens the mechanistic frame through which both HLA-C02:02 and HLA-A02:01 should now be viewed. The paper is not a p53 paper. It does not center tumor antigens, and it does not establish anything directly about TP53 peptide presentation. What it does show, at population scale, is that viral DNA load is shaped not only by HLA variation but also by the antigen-processing machinery, especially ERAP1 and ERAP2. That matters because it shifts the center of gravity away from a simplistic “does the peptide bind?” model and toward a more realistic “does the peptide survive the whole processing pipeline?” model.

That shift is especially important for p53. The HLA-A02:01 literature had already hinted that presentation of the classic p53(264–272) epitope depends on more than sequence alone. The work by Kuckelkorn et al showed that generation of this epitope is influenced by the interferon-γ-inducible processing machinery and that a hotspot mutation at residue 273 can prevent proper generation of the epitope. This is a reminder that even for the most familiar p53/HLA-A02:01 peptide, presentation is a processing problem before it becomes a recognition problem. The Nature virome study widens that principle: inherited variation in antigen processing can have measurable biological consequences at human scale. Read together, these papers suggest that p53 visibility is governed not simply by the existence of a fitting sequence, but by whether intracellular processing delivers that sequence intact to the appropriate HLA molecule.

This is where the contrast between HLA-A02:01 and HLA-C02:02 becomes genuinely interesting. HLA-A02:01 has a long experimental trail behind it: peptides were mapped, CTLs were induced, tumors were shown to present certain epitopes, and vaccine studies were built on top of that scaffold. HLA-C02:02, by contrast, remains more conditional. The ligandome study establishes that TAKSVTCTY from TP53 can indeed appear on HLA-C02:02, and it also gives a broader view of the peptide preferences of that allotype. In that same work, HLA-C02:02 is described as favoring small aliphatic or hydrophilic residues at position 2, with additional motif features helping define its ligand space. That does not diminish the importance of the TP53 observation; it means the TP53 peptide should be treated as a real but selective presentation event rather than assumed to be broadly immunodominant.

The biology becomes even more layered because HLA-C is not simply a lower-profile version of HLA-A. HLA-C occupies a distinct place in immune regulation. Compared with HLA-A and HLA-B, HLA-C is generally expressed at lower surface levels and is more tightly integrated with KIR-mediated NK-cell regulation. That broader point is well summarized in the Nature Communications paper Structural and regulatory diversity shape HLA-C protein expression levels, which notes both the lower surface expression of HLA-C and its extensive functional relationship with KIRs. This makes HLA-C particularly interesting for p53 because a peptide displayed by HLA-C is not only a possible T-cell target; it is also part of a signaling surface read by NK cells.

That NK dimension turns out not to be merely background context. More recent work has shown that KIR recognition of HLA-C is often peptide-dependent. The point is made clearly in studies such as Sim et al. 2017 and Sim et al. 2023: the HLA-C molecule is not being read in a peptide-blind way. Inhibitory and activating KIRs can be strongly shaped by the identity of the peptide bound in the groove. That has profound implications for any discussion of TP53 peptides on HLA-C02:02. A TP53-derived peptide on HLA-C02:02 may not simply mark a cell for CD8 T-cell inspection; it may also alter the threshold for NK inhibition or activation. This is one of the most important places where the older Codondex article and the newer immunogenetic literature genuinely converge.

So the corrected reading is not that the 2026 Nature paper newly proves something specific about HLA-C*02:02 presenting p53. It does not. What it does is make the older HLA-C02:02 observation more meaningful by placing it inside a stronger mechanistic framework. The question is no longer only whether TAKSVTCTY can bind HLA-C02:02; the question is whether an individual’s processing machinery, inflammatory state, and HLA context allow that peptide to be generated, preserved, loaded, displayed, and then interpreted by either T cells or NK cells in a biologically consequential way. That is a more demanding question, but it is also a more interesting one.

This also helps explain why HLA-A*02:01 remains the more established p53 route. The A02:01 pathway has yielded peptides that are repeatedly recoverable in experimental systems, repeatedly recognized by CTLs, and repeatedly leveraged in translational work. The HLA-C02:02 pathway looks more contingent: real, but likely more dependent on peptide selection pressure, trimming, and the NK-facing consequences of peptide-loaded HLA-C. Seen this way, HLA-A02:01 is the clearer adaptive pathway, while HLA-C02:02 may be a narrower but potentially more intriguing bridge between tumor antigen presentation and innate immune tuning.

That may be the most useful lesson from putting these papers together. p53 is not simply “presented” or “not presented.” It passes through a filter. In HLA-A02:01, that filter has already produced a clinically legible signal. In HLA-C02:02, the signal is fainter, but perhaps more information-rich, because it may be read simultaneously by T cells and NK-cell receptor systems. If that is right, then the next real step is not more speculation about binding motifs alone. It is experimental work that directly compares TP53 peptide generation, ERAP dependence, surface abundance, and KIR/TCR consequences across HLA-A02:01 and HLA-C02:02 backgrounds. That is where the overlap becomes testable rather than merely suggestive.

Saturday, January 17, 2026

Genome Balance: Repeats, Immunity, and Cancer


Cancer is usually described as a disease of mutations. Genes break, pathways fail, and cells escape control. That framing has been powerful, but it misses a deeper layer that may reveal how it begins.

The human genome is not primarily a coding genome. It is a repeat genome. More than half of our DNA consists of repetitive elements, with Alu retroelements alone numbering over a million copies. These sequences are a defining feature of primate genomes and they create a unique biological problem that human cells must continuously manage. Recent work suggests that cancer may emerge, in part, when this management system loses balance.

Alu elements are short retrotransposons that readily form double‑stranded RNA stem‑loop structures when transcribed, particularly in antisense orientation within introns and untranslated regions. To the innate immune system, these structures resemble viral RNA. This means that normal gene expression in human cells constantly risks triggering antiviral immune responses against self‑derived RNA.

A striking recent study shows that human cells rely on active suppression to avoid this outcome. In Ku suppresses RNA‑mediated innate immune responses in human cells to accommodate primate‑specific Alu expansion, the authors demonstrate that the DNA repair protein Ku (Ku70/Ku80) plays an essential second role: binding Alu‑derived dsRNA stem‑loops and preventing activation of innate immune sensors such as MDA5, RIG‑I, PKR, and OAS/RNase L.

When Ku is depleted interferon and NF‑κB signaling are strongly activated, translation is suppressed, and cells undergo growth arrest or death. Notably, Ku levels scale tightly with Alu expansion across primates, and Ku is essential in human cells but not in mice. The implication is clear:

Human cell viability depends on continuous suppression of Alu‑derived innate immune activation.

Alu expression is not harmless noise, it is actively tolerated! Ku functions as a finite buffer that allows primate cells to tolerate structurally immunogenic RNA produced by repeat‑rich genomes. When structured RNA load increases simultaneously from endogenous repeat transcription and exogenous viral RNA infection, Ku becomes functionally saturated and redistributed, weakening nuclear retention and cytoplasmic buffering. This pressurizes the cell’s capacity to contain dsRNA stress, promoting escape of repeat‑derived RNA, activation of innate sensors, and eventual selection for immune‑tolerant states.

A second line of evidence connects this tolerance to cancer evolution. A 2025 bioRxiv preprintp53 loss promotes chronic viral mimicry and immune tolerance, shows that loss of p53 permits transcription of immunogenic repetitive elements, generating signals that resemble viral infection. Rather than leading to effective immune clearance, this state becomes chronic. Tumor cells adapt by dampening innate immune responses and tolerating persistent repeat‑derived nucleic acids.

In this view, “viral mimicry” is not a one‑time immune alarm. It is a conditioning process repeat RNAs accumulate, immune pathways are activated, and progressively suppressed or rewired to allow survival. Cancer cells do not simply evade immunity, they learn to live with endogenous viral‑like signals.

These immune findings align with earlier evidence that repeat control begins at the level of genome structure itself. A 2022 Nature Communications study demonstrated that retroelements embedded within the first intron of TP53 act as cis‑repressive genomic architecture. Removing this intron increases TP53 expression, indicating that long‑embedded repeats contribute directly to regulating a core tumor suppressor gene.

Importantly, this repression is architectural rather than motif‑driven. The repeats do not act through a single conserved sequence, but through repeat‑dense structure.

Together, these findings suggest a layered system of control:

  1. Structural repression of repeats within introns.

  2. Immune suppression of repeat‑derived dsRNA.

  3. p53‑dependent governance of both genome stability and immune signaling. 

One long‑standing challenge in repeat biology is inconsistency. Different tumors show different repeat fragments. Even different regions of the same tumor can look unrelated at the sequence level.

From a traditional biomarker perspective, this appears discouraging. From a structural perspective, it is expected. Codondex analyses of repeat‑dense introns, including TP53 intron 1, show that cancer does not preserve specific Alu sequences. Instead, it perturbs repeat topology:

  • dominance and skew within intronic scaffolds,

  • stem‑loop‑prone architectures,

  • context‑specific fragmentation patterns.

The sequences vary. The instability regime does not. This is characteristic of a state change, not a discrete genetic event. Repeat‑dense introns behave like stress recorders. They integrate replication stress, chromatin relaxation, repair pathway bias, and immune tolerance history.

Unlike coding mutations, these signals are heterogeneous, region‑specific, and reflective of ongoing cellular state.

They are difficult to interpret with gene‑centric tools, but powerful when viewed architecturally. 

Most cancer diagnostics ask:

What mutation is present? A repeat‑aware framework asks:

Has this tissue entered a stable state of repeat derepression coupled with immune tolerance?

That state may precede aggressive behavior, accompany treatment resistance, or mark transitions in disease evolution. Future prognostic approaches may therefore combine repeat‑topology instability metricsrepeat RNA burden, and evidence of immune decoupling from dsRNA load. Not to identify a single driver, but to detect loss of containment.

Alu repeats do not cause cancer on their own, but human cells must continuously restrain them, structurally and immunologically. Cancer appears, at least in part, when that restraint erodes and tolerance replaces control. Introns, long treated as background, may be one of the clearest places to see this shift, not because they encode instructions, but because they actively record genomic history and project it into a measure of present state.


Wednesday, August 13, 2025

Repeats as Signatures of Regulatory Potential


In the vast landscape of AI genomics, emerging analyses reveals non-coding DNA (ncDNA) as a treasure trove of regulatory information. At Codondex, our innovative k-mer-based approach uncovers how repetitive subsequences—short DNA fragments known as k-mers—serve as powerful signatures of regulatory potential. By viewing these repeats through a topological lens, we transform linear sequences into dynamic networks that highlight subtle distinctions in gene transcripts, offering new insights into gene regulation, isoform diversity, and disease mechanisms.

The Codondex Method: From Sequences to Topology

Codondex begins by "amplifying" ncDNA sequences associated with gene transcripts, generating all contiguous k-mers of length 8 or greater. For a gene like TP53, with its multiple isoforms (variants), we associate these k-mers with transcript-specific signatures derived from cDNA, mRNA or protein constants. The result? A rich dataset of subsequences, where repeats—identical k-mers appearing multiple times—emerge as key players.

Rather than treating DNA as a flat string, we interpret it topologically: k-mers as nodes in a graph, with repeats forming edges that indicate connections, clusters, and symmetries. Metrics like i-Score (normalizing contained k-mers by length) and inclusiveness (repeat frequency) rank these patterns, while cDNA or protein vectors capture fine distinctions. In our analyses of genes such as MEN1 and TP53, symmetries in repeat length and frequency stand out, unrelated to obvious features like reverse complements. These non-random patterns suggest repeats are not artifacts but deliberate signatures encoded for regulation.

Repeats as Regulatory Hotspots

How do these repeats signal regulatory potential? First, they often manifest as binding sites for proteins. Repetitive motifs can amplify affinity for transcription factors or splicing regulators. In TP53 introns, high-frequency k-mers align with p53-binding elements, potentially modulating tumor-suppressive isoforms. Variants with asymmetric repeats might weaken these interactions, leading to dysregulation in cancer.

Second, repeats influence secondary structures. Topologically, frequent repeats create "hubs" in the network, fostering DNA/RNA folds like hairpins that affect chromatin accessibility or mRNA stability. Our MEN1 intron1 study, analyzing 15 variants, revealed length-biased repeat clusters in scatter-graphs—despite length-agnostic algorithms—indicating structured motifs that differentiate stable from unstable transcripts. Disruptions from low-length repeats, as seen in TP53 vectors, act like regulatory "switches," fine-tuning expression in response to cellular stress.

Third, symmetries in repeats point to evolutionary conservation. Equal-length k-mers recurring with balanced frequencies form symmetric graphs, preserving robust modules across species. In MEN1, linked to endocrine tumors, these patterns suggest intron-driven adaptations for hormone regulation. Disruptions in variants could flag pathogenicity, enabling predictive modeling without coding-sequence reliance.

Real-World Implications and Validation

Our deep k-mer analysis, first detailed in a 2018 blog post, showcased MEN1 intron symmetries predicting protein outcomes, later validated through lab tests at Tel Aviv University. For TP53, stable vector positions disrupted by specific repeats correlated with isoform-specific roles, highlighting ncDNA's influence on cancer hallmarks.

This topological view empowers genomics: identifying regulatory elements for drug targeting, differentiating disease variants, and advancing precision medicine. At Codondex, we're excited to explore how these repeat signatures unlock ncDNA's secrets—join us in redefining genomic potential.


Wednesday, February 19, 2025

P53 - Stability and Life Or Disorder and Death!

There is something ancient about the struggle between order and disorder in biology. A cell does not merely live by dividing, signaling, and repairing itself. It lives by maintaining interpretability. Its genome must remain legible enough to be copied, restrained enough not to erupt into instability, and coherent enough that surrounding systems — especially the immune system — can still distinguish function from failure. In that sense, p53 is not simply a tumor suppressor in the narrow modern meaning of the phrase. It is closer to a molecular governor of biological intelligibility, one of the factors that helps determine whether stress remains containable or tips into forms of disorder.

The broader role has appeared repeatedly across Codondex discussions, from Expanding Treatment Horizons to Does SARS-CoV2 Strangle P53 to kill Natural Killer Immunity?, where p53 was already being read less as an isolated tumor suppressor and more as part of a wider immune and genomic control system.

That older and deeper role becomes clearer once p53 is viewed not only through apoptosis or cell-cycle arrest, but through its relationship with the repetitive genome. Work over the last decade has shown that p53 does not merely respond to genetic insult after the fact. It can directly repress human LINE1 retrotransposons by binding the 5′UTR and promoting local repressive chromatin, and more recent work has extended that picture by showing p53-dependent restraint of LINE1-associated RNA-DNA hybrid states as well. One of the great guardians of cellular integrity is therefore also engaged in policing one of the genome’s recurrent internal threats: mobile and semi-mobile repetitive sequence that, when released from restraint, can destabilize chromosomal order and provoke inflammatory consequences. p53 directly represses human LINE1 transposons and p53-mediated regulation of LINE1 retrotransposon-derived R-loops both push that picture into sharper focus.

But the relationship runs in both directions. Transposable elements are not only targets of p53; they have also helped shape the p53 regulatory landscape itself. A substantial body of work has shown that human retrotransposons contain p53 responsive elements, meaning that the repetitive genome has donated part of the sequence architecture through which p53 now reads and regulates stress. This is one of those places where the older division between “functional genome” and “junk” becomes difficult to maintain. Repetitive sequence has not only threatened order. It has also contributed to the grammar by which order is defended.

Once that is appreciated, the immune side of the problem begins to look less like a separate field and more like a continuation of the same one. If p53 helps determine whether genomic instability remains suppressed, then it also helps determine whether such instability becomes visible to immune surveillance. Emerging views now frame p53 as a major regulator of NK-cell tumor immunosurveillance, not because NK cells are somehow subordinate to p53, but because p53 influences so many of the target-cell properties that NK cells are built to read: stress ligands, metabolic distress, microenvironmental signals, and the broader state of cellular legitimacy.

One of the clearest examples is the p53-dependent induction of ligands visible to NK activation pathways. When wild-type p53 is induced in tumor cells, NK cells can be alerted through upregulation of the NKG2D ligands ULBP1 and ULBP2. That is an important bridge. It means p53 is not only preserving internal order; it can also help convert intracellular stress into a surface-readable signal that tells NK cells something has gone wrong. The target is no longer merely unstable. It becomes interpretable to innate immune surveillance.

A parallel bridge exists through antigen processing. p53 has also been shown to increase MHC class I expression by upregulating ERAP1, a trimming enzyme involved in preparing peptides for class I presentation. That does not collapse NK and T-cell biology into one another, but it does reinforce the larger point: p53 influences whether a distressed cell remains hidden, partially legible, or fully exposed to immune scrutiny. It affects not just whether the cell survives, but how clearly that cell can be judged by the systems around it.

The recombination thread can also be preserved, though it needs to be understood in the right register. Mature NK cells do not generate their recognition receptors through classical V(D)J recombination in the way T and B cells do. Their receptors are fundamentally germline encoded. Yet that is not the end of the recombination story. Work on NK ontogeny has shown that a history of RAG expression in progenitors and NK precursors marks functionally distinct NK subsets later in the periphery, with consequences for fitness, survival, and responsiveness. Recombination machinery therefore leaves a developmental trace on NK biology, even if it does not build the mature receptor repertoire in the adaptive sense.

That developmental nuance matters because it prevents the argument from becoming either too weak or too strong. Too weak, and the relationship between recombination-linked stress and NK function disappears. Too strong, and NK cells are mistakenly described as if they were just another rearranging lymphocyte lineage. The better reading is that recombination biology, DNA damage response history, and developmental programming can shape the later functional competence of NK cells without making their mature surveillance logic identical to that of T cells.

A similar caution, and opportunity, appears in the KIR story. One of the most interesting findings in NK regulation is that KIR expression can be governed by bidirectional promoter logic and antisense transcription. In particular, KIR antisense transcripts processed into a 28-base PIWI-like small RNA have been linked to transcriptional silencing, while related work on KIR antisense lncRNAs and probabilistic promoter switching suggests that inhibitory receptor expression is shaped by a layered interaction between transcription, antisense regulation, and epigenetic commitment. This does not establish a direct p53→piRNA→KIR3DL1 pathway. But it does show that the NK lineage is not insulated from the wider world of small RNA restraint and genome-governed silencing.

That is where the Codondex theme begins to re-emerge. If p53 sits in one part of the cell as a governor of repetitive-element restraint, and if NK inhibitory receptor choice is itself touched by small-RNA and antisense-mediated silencing logic, then the two systems may not be identical, but they may still rhyme. Both are concerned with the management of unstable potential. Both are concerned with whether latent disorder is allowed to become active. Both are concerned with which signals are permitted to surface and which are held in reserve. This is not yet a single proven pathway. It is a systems-level parallel supported by a growing amount of molecular detail.

There is another reason the p53–NK connection deserves attention. In some settings, p53 activation appears able to convert repetitive-element biology into something resembling a warning flare. Pharmacologic activation of p53 has been linked to antiviral-like and immune-stimulatory states, and the broader literature now places p53 within a network that can enhance NK recognition and tumor destruction through multiple channels rather than one single canonical mechanism. The significance of that shift should not be underestimated. It means p53 is no longer best understood only as the decider of cell fate from within. It is also a participant in the communication of cellular fate to the outside world.

So the central question remains a fruitful one. Is p53 merely a brake on instability, or is it also part of the language by which instability becomes visible to elimination? The literature increasingly favors the second possibility. p53 restrains transposable elements. p53 shapes stress-ligand display. p53 influences antigen processing and MHC-I expression. p53 intersects with developmental and regulatory processes that matter to NK-cell competence and target recognition. The picture that emerges is not of a single linear circuit, but of a pressure point where genome integrity, immune legibility, and cellular fate begin to converge.

In that light, p53 can still be read as this article’s central character without overstatement. When p53 function is preserved, a cell under strain is more likely to remain ordered, to arrest, to die cleanly, or to become visible enough for immune removal. When p53 function is lost, not only does instability grow, but the cell’s interpretability may degrade with it. Disorder then becomes doubly dangerous: more abundant internally, and more ambiguous externally. That may be one of the deeper meanings of p53 in cancer and perhaps in biology more generally. It is not only a guardian against mutation. It is one of the means by which life keeps disorder readable.


Wednesday, May 17, 2023

Immune Synchronization

Stem Cell

Navigating the regulatory regimes that govern drug safety can be challenging. But, rigorous standards are more relaxed in the lesser used track for autologous and/or minimally manipulated cell treatments. Toward meeting the challenges of this minimal regulation track, the wide-spectrum of NK cells, of the innate immune system, are compelling candidates to address complex cellular and tissue personalization's or conditions of disease. One effect of cell function on NK cell potency occurs via aryl hydrocarbon receptor (AhR) dietary ligands, potentially explaining numerous associations that have been observed in the past.

The AhR was first identified to bind the xenobiotic compound dioxin, environmental contaminants and toxins in addition to a variety of natural exogenous (e.g., dietary) or endogenous ligands and expression of AhR is also induced by cytokine stimulation. Activation with an endogenous tryptophan derivative, potentiates NK cell IFN-γ production and cytolytic activity which, in vivo, enhances NK cell control of tumors in an NK cell and AhR-dependent manner.

A combination of ex vivo and in vivo studies revealed that Acute Myeloid Leukemia (AML) skewed Innate Lymphoid Cell (ILC) Progenitor towards ILC1's and away from NK cells as a major mechanism of ILC1 generation. This process was driven by AML-mediated activation of AhR, a key transcription factor in ILC's, as inhibition of AhR led to decreased numbers of ILC1's and increased NK cells in the presence of AML.

Activation of AhR also induces chemoresistance and facilitates the growth, maintenance, and production of long-lived secondary mammospheres, from primary progenitor cells. AhR supports the proliferation, invasion, metastasis, and survival of the Cancer Stem Cells (CSC's) in choriocarcinoma, hepatocellular carcinoma, oral squamous carcinoma, and breast cancers leading to therapy failure and tumor recurrence.

Loss of AhR increases tumorigenesis in p53-deficient mice and activation of p53 in human and murine cells, by DNA-damaging agents, differentially regulates AhR levels. Activation of the AhR/CYP1A1 pathway induces epigenetic repression of many tumor suppressor and tumor activating genes, through modulation of their DNA methylation, histone acetylation/deacetylation, and the expression of several miRNAs. 

p53 is barely detectable under normal conditions, but levels begin to elevate and locations change particularly in cells undergoing DNA damage. The significant network effect of p53 availability and its mutational status in cancer makes it the worlds most widely studied gene. 

From 48 sequenced samples of two different tumors, Codondex identified 316 unique Key Sequences (KS) of the TP53 Consensus. 9 of these contained the core AhR 5′-GCGTG-3′ binding sequence, and some overlapped p53 quarter binding sites as illustrated below;

Key Sequence                                                                           

GGATAGGAGTTCCAGACCAGCGTGGCCA (intron1) AhR [1699,1726], p53 @ [1706,1710]

AAAAATTAGCTGGGCGTGGTGGGTGCCT (intron1) AhR [1760,1787], p53 [1783,1787]

AAAAAAAATTAGCCGGGCGTGGTGCTGG (intron6) AhR [12143,12170]

GAGGCTGAGGAAGGAGAATGGCGTGAAC (intron6) AhR [12195,12222]

We propose that DNA damage liberates transposable DNA elements that are normally repressed by p53 and other suppressor genes. The p53 repair/response also includes increased cooperation between p53 and AhR, which further influence transcription, mRNA splicing or post-translation events. Repeated damage, at multi-cellular scale, may proximally bias ILC's toward NK cells capable of specific non-self detection, through localized ligand, receptor relationships that trigger cytolysis and immune cascades. 

KS's are a retrospective view of transcripts ncDNA elements, ranked by cDNA that may reflect inherent bias that can be used to direct NK cell education. One way to accomplish minimal manipulation may be to leverage patient immunity by educating autologous NK cells with computationally selected tumor cells, identified by KS alignments to the index of past experiments that expanded and triggered a more desirable immune response. Customizable immune cascades, capable of managing disease or preventatively supporting a desired heterogeneity being the primary objective. 


Wednesday, November 17, 2021

Retroviral Defense And Mitochondrial Offense


Chromosomal DNA has played host to the long game of viral insertions that repeat and continue as a genetic and epigenetic symbiosis along its phosphate and pentose sugar backbone. But, the bacterial origin of mitochondria and its hosted DNA also promotes its offense. 

Research suggests that retrovirus insertions evolved from a type of transposon called a retrotransposon. The evolutionary time scales of inherited, endogenous retroviruses (ERV) and the appearance of the zinc finger gene that binds its unique sequences occur over same time scales of primate evolution. Additionaly the zinc-finger genes that inactivate transposable elements are commonly located on chromosome 19. The recurrence of independent ERV invasions can be countered by a reservoir of zinc-finger repressors that are continuously generated on copy number variant (CNV) formation hotspots.

One of the more intiguing aspects of prevalent CNV hotspots on chromosome 19 are their proximity to killer immunoglobulin receptor gene's (KIR's) and other critical gene's of the innate immune system.

Frequently occuring DNA breaks can cause genomic instability, which is a hallmark of cancer. These breaks are over represented at G4 DNA quadruplexes within, hominid-specific, SVA retrotransposons and generally occur in tumors with mutations in tumor suppressor genes, such as TP53. Cancer mutational burden is shaped by G4 DNA, replication stress and mitochondrial dysfunction, that in lung adenocarcinoma downlregulates SPATA18, a mitochondrial eating protein (MIEAP) that contributes to mitophagy. 

Genetic variations, in non-coding regions can control the activity of conserved protein-coding genes resulting in the establishment of species-specific transcriptional networks. A chromosome 19 zinc finger, ZNF558 evolved as a suppressor of LINE-1 transposons, but has since been co-opted to singly regulate SPATA18. These variations are evident from a panel of 409 human lymphoblastoid cell lines where the lengths of the ZNF558 variable number tandem repeats (VNTR) negatively correlated with its expression. 

Colon cancer cells with p53 deletion were used to analyze deregulated p53 target genes in HCT116 p53 null cells compared to HCT116-p53 +/+ cells. SPATA18 was the most upregulted gene in the differential expression providing further insight to p53 and mitophagy via SPATA18-MIEAP.

p53 response elements (p53RE) can be shaped by long terminal repeats from endogenous retroviruses, long interspersed nuclear repeats, and ALU repeats in humans and fuzzy tandem repeats in mice. Further, p53 pervasively binds to p53REs derived from retrotransposons or other mobile genetic elements and can suppress transcription of retroelements. The p53- mediated mechanisms conferring protection from retroelements is also conserved through evolution. Certainly, p53 has been shown to have other roles in DNA  context, such as playing an important role in replication restart and replication fork progression. The absence of these p53-dependent processes can lead to further genomic instability. 

The frequency of variable length, long or short nucleotide repeats and their locations within a gene may be key to the repression of DNA sequences that would otherwise cause genomic instability or protein expressions that would eat bacterial mitochondria or destroy its cell host. 

The complexity of variable length insertions is made evident when exhaustively analyzing a simple length 12 sequence for the potential frequency of each of its variable length repeats starting from a minumum variable length of 8.

Then, for TGTGGGCCCACA(12)

All possible internal variable length combinations from and including length 8:

TGTGGGCC(8)|GTGGGCCC(8)|TGTGGGCCC(9)|TGGGCCCA(8)|GTGGGCCCA(9)|TGTGGGCCCA(10|GGGCCCAC(8)|TGGGCCCAC(9)|GTGGGCCCAC(10)|TGTGGGCCCAC(11)|GGCCCACA(8)|GGGCCCACA(9)|TGGGCCCACA(10)|GTGGGCCCACA(11)|TGTGGGCCCACA(12)

For example, reviewing length (8) only:

TGTGGGCC (8) occurs 5 times

GTGGGCCC (8) occurs 8 times

TGGGCCCA (8) occurs 9 times

GGGCCCAC (8) occurs 8 times

GGCCCACA (8) occurs 5 times

Any repeat can be ranked based on its ocurrence within all possible combinations of a given sequence, known as the repeats' iScore rank. This illustrates a potential useful statistical ranking that, subject to biology may describe a repeats inherency to be more or less effective, in increments of the gene sequence. 

Repression of the most active sequences, especially in context of repeats may result in genetic variation.