
How to Kill a Shapeshifter – A New Weapon?
In 2023, I returned to the cancer biology field after over 10 years away and the first thing that struck me was how much the field had changed, namely how much data was being generated. My PhD work took an approach that was common for pre-clinical, academic cancer research at the time. The usual flow for molecular cancer research involved identifying an interesting molecular change (in my case, an upregulation of the Notch pathway signaling protein Jagged1), confirming this change in diverse human tumor samples and data sets, identifying a model system, knocking it down or perturbing this biomarker in some way, maybe making a mouse model of it, and finally looking at the impacts on cancer biology when you disrupt the molecule. In other words, a reductionistic approach to oncology, understanding the key molecular drivers in a certain cancer and focusing research around that one specific molecular change. Targeting treatment to an individual type of cancer based on its specific genetic changes or other molecular features is called precision oncology.
More than 10 years later, precision oncology is still the name of the game but now everything is focused on as much data as possible. The reductionist single gene, single target approach yielded to massive datasets derived from a whole rainbow of advanced “seq” and “omics” techniques designed to look at genetic, transcriptomic, epigenetic, proteomic, and other data types in millions of individual cancer cells, sometimes looking at multiple different molecular classes simultaneously (plus spatial and temporal combinations). In other words, oncologists are cranking out more data than ever before and refining their picture of what cancer is at the molecular level to higher and higher resolutions. AI, allegedly skilled at making sense of complexity and drawing connections between gargantuan amounts of data, could very well help decipher the mysteries of cancer at the molecular level. Demonstrable benefits led by AI have already occurred in the oncology field and some of the most anticipated benefits of AI applications to biology are predicted as well. But is the hype warranted? What has AI accomplished so far, what are the gaps, and what are the predictions for the future? What is AI-led research finding that previous decades of human-led research have not? Is usage of biological data in AI systems safe, ethical, and controllable?
Cancer is a protean and resilient disease. There is no single cancer but hundreds of diseases that all share similar behaviors. Genetic mutation, either inherited, caused by environmental factors, or random errors, is the underlying mechanism of all cancers. Looking for these specific genetic perturbations in a specific cancer, while standard in the field these days, was itself a huge advancement over the previous decades. Indeed, the first true targeted cancer therapy against a discovered molecular change, HER2 in breast cancer, occurred in 1998 (Herceptin) followed by the approval of Gleevec for chronic myelogenous leukemia (CML) in 2001, and more than a hundred more since then. Cancer care was getting personal and the impacts on patient survival were significant (in combination with many other advances in cancer care). For example, CML used to be a death sentence but now with a family of molecules following after Gleevec, CML is one of the most treatable cancers. Despite these many advancements, treatment of many cancers remains dismally intractable.
One reason is that cancers are not homogenous. Even within a single cancer in a single patient, there are many different lineages of cancer cells and different driver mutations. The targeted therapies against the primary genetic drivers could be effective in killing 99% of cancer cells. Unfortunately, the 1% of remaining cells could comprise dozens of other subtypes avoiding the silver bullet of the therapy due to some other mutations present in them not targeted by the drug. These dormant lineages waiting in the wings then divide, mutate, evolve, and expand. Cancer recurrence/relapse usually occurs because the cancer is a different beast than before with new mutations the old drug doesn’t work against. Even worse, there may be even more sub-types within that cancer waiting for their turn to grow. “Cure” in cancer means a relentless battle with a shapeshifting enemy that is excellent at evading that thing that you designed to kill it.
The Big Data Problem in Oncology
Deciphering cancer at the individual level requires a lot of understanding of all the meaningful molecular changes occurring in that cancer, not just the most prominent ones. The big-data, multi-modal approach I described above is tackling this challenge by looking at everything from the immune microenvironment, to diverse tissue types, to cancer DNA floating around in the blood (circulating tumor DNA, or ctDNA), and so much more.

Another thing that struck me upon my return to the field was that every presentation using big data approaches tended to end the same way: a Uniform Manifold Approximation and Projection (UMAP) plot that showed how all the thousands of individual cancer cells that were just sequenced clustered together. UMAP is a way of reducing the dimensionality of these huge data sets and compressing the biggest trends into visual patterns (Figure 1. for example). The clusters on a UMAP plot correspond to different cell populations with statistically related molecular patterns. The biological consequences of these different clusters? Unknown. Therapeutic implications? Also unknown. What does all this data mean? How do these clusters actually translate into an understanding of cancer biology in such a way that we can develop new therapies, or new ways of detecting or tracking cancer, or targeting existing therapies to the right people? Of course, insights are being drawn from all this data by people much smarter than me and tremendous work is happening on all these fronts (check out the AACR 2025 Cancer Progress Report for a summary of some accomplishments over the previous year). But my point is that if the amount of data being generated is only going to increase, then the human ability to comprehend it all needs some serious help. Maybe AI can do this better. AI models may be able to take the data, parse it, and figure out all the nuanced subtypes of cancer cells occurring in a single patient, and figure out which ones are the most dangerous or which therapies will work the best.
With the advent of AI, I believe the field is changing once again. Oncologists understand the power of big data, and many brilliant computational biologists identified many prominent signals that have led to hundreds of targeted therapies, but what all those UMAP plots translate into still eludes the field. UMAP is itself a machine learning approach and many other AI algorithms have been deployed in cancer biology for years. The most advanced AI models like large language models (LLMs), neural networks, and others (to be reviewed in a future post) seem to be exceptionally good at drawing patterns and making unexpected connections from the chaos of big data sets. But how exactly does AI apply to the cancer biology field? What is happening now with AI and cancer, what is it expected to do, and what are the limitations, both technical and practical?
AI Applications in the Cancer Care Continuum
AI in cancer can impact nearly every aspect of the cancer care continuum [Fig 2, insert image of cancer care continuum] from diagnosis to molecular characterization to treatment selection to development of new therapeutics. These fall roughly into three buckets which I’ll quickly summarize here for awareness. How exactly researchers are using AI in each of these areas requires a much deeper interrogation of individual papers. I’ll save that for future posts. The bread and butter of cancer biology is in these details, but I believe a brief high-level survey of where oncologists are already applying AI is still useful.

- Prevention/diagnosis
The most effective way to increase survival from cancer is to identify it as early as possible. A promising avenue for AI could include screening of healthy patients that lack any biomarkers normally used to detect cancer. An exciting hypothesis is that AI may be able to identify patterns in certain blood tests or other health information routinely collected that predict the early stages of cancer that would otherwise be missed. Related, AI tools have already been developed to assist in routine screenings for skin, breast, and colorectal cancer. ctDNA is a major advancement in recent years for early detection. For example, an ML tool called GRAIL Galleri looks at methylation status in ctDNA to detect multiple cancer types and has gained CLIA certification. Finally, the diagnosis/characterization and classification of a cancer is probably the most familiar application for AI in cancer. For example, AI models trained on skin cancer images can diagnose new cancers to the same level as dermatologists, or maybe even better.
- Optimizing current treatments and clinical trials
I mentioned the growing catalogue of targeted therapies and known driver mutations in hundreds of different cancers. How can one oncologist be expected to possibly keep up with the rapid changes and discoveries in the field and integrate the cumulative findings of researchers across the world? As discussed above, AI may be a powerful tool in this molecular characterization. An oncologist needs to know if the molecular signature predicts a slow growing tumor or an aggressive one and a possible prognosis. AI is being put to use on predicting patient outcomes and identifying the molecular signatures that are better/worse for a patient. For example, a model using genetic and transcriptomic data can predict favorable colorectal cancer outcomes. Similarly, AI may be able to discover new biomarkers that predict a positive response to a particular therapy. This is a hot area of research and many AI models fed on diverse data types have shown promise (for example, predicting immunotherapy response and therapy response in colorectal cancer and breast cancer) . Finally, matching the right patient to a clinical trial is complicated since a multitude of inclusion criteria must be met in order for the patient to be enrolled. AI involvement in matching patients with an appropriate trial could improve the efficiency of this process and some tools have already been developed. Assuming new AI tools gain FDA approval and demonstrate proven impacts on patient outcomes, diffusion and adoption of any of these new molecular tools will likely take time. Genomic profiling, one of the important things that can be done after a confirmed cancer diagnosis, is still not used in 100% of cases and it has been around for more than a decade.
- Discovery and future treatments
Much of the discussion above on big data and cancer and all the research going on feeds directly into this category. AI models integrating this data could help in discovering new cancer biology mechanisms and molecular vulnerabilities of cancer never before considered. New AI models built to design experiments and hypotheses may also accelerate the work of their human scientist counterparts. Predicting effective treatment combinations is another particularly exciting area. With the expansion of new cancer drugs, there can be synergistic effects where two drugs can produce much greater responses than each one individually. Given the huge numbers of combinations, running all possible combination trials is impossible, but viable combinations may be able to be predicted by AI. For example, an AI model has been developed that can predict drug combination responses in breast cancer. Finally, AI is already being deployed to design new and better drugs. Many big pharma companies are betting big on this as evidenced by the investments many are making. However, AI may be able to come up with a million excellent drugs, but that doesn’t mean the whole therapeutic pipeline, from pre-clinical to clinical to manufacturing to patient, accelerates too (yet).
Challenges, Risks, and Limitations
So far, AI has been employed in each area above and has even led to FDA approval for some AI-based cancer screening tools. Big pharma companies are investing heavily into AI and partnering with the big AI labs, who are themselves forming their own biology divisions. But within the context of all these advances for cancer biology and for AI in general, the golden future promised by the AI evangelists is hardly a guarantee. Besides what the technology can do itself, I can imagine quite a few potential limitations and hurdles. This list is not an attempt to be comprehensive but just a few things that jump out at me now. As I learn more, I imagine this list will grow and change.
- Data quality and data availability
There is no shortage of biological data. The amount of genomic data uploaded into public repositories is on the order of petabytes. Powerful LLMs like ChatGPT may have been trained on similar amounts of data; one estimate puts GPT-4 at about 1 petabyte. So why don’t we have ChatGPT-like models for biology then? Unlike text data, there is a lot of redundancy in human DNA. All those petabytes of genetic data collected over the past 20 years are likely 99.9% the same because after all, most humans are 99.9% identical to each other at the genetic level. Due to this redundancy, the information efficiency of genetic data is so much less than that of text data. Deciphering unique and meaningful gene variants requires feeding an AI model a lot of repetitive DNA sequences. And storing all those data and reading all those redundant genomes is expensive and may not necessarily yield anything too interesting. Also, text data, regardless of where it comes from, is all structured pretty much the same way. The same cannot be said of biological data. DNA sequencing data is not the same as histology images which is not the same as epigenetic data and so on. How is an AI model trained on these different data sets? How is it structured to draw parallels between very different types of data? This is my biggest knowledge gap but is the most critical parameter in the technical capabilities of AI applied to biology.
- Data annotation
The mantra of the AI age is garbage in, garbage out. Meaning, if you train an AI model with garbage data, then the model gives you garbage responses. The same is true of biological data and likely one of the biggest gaps holding back AI models in biology. Directly related to the points above, biological data without the experimental context or relevance of that data curated with a biological outcome is almost useless. You can feed an AI model the sequencing results from thousands of cancer cells from thousands of patients, but if you don’t tie that data to say, the survival rate or how that patient responded to a drug, how will the AI be able to decipher which of those pieces of data is the most meaningful? How do you annotate these huge amounts of biological data so that it is not just a bunch of big data sets but data linked to the outcome or interest you want the AI to interpret for you? Sequencing may have gotten 200,000 times cheaper since 2001, but figuring out what a gene does costs roughly what it did in 1998, when Herceptin was first approved. Some data sets are being produced now solely for the purpose of overcoming this limitation for the express purpose of training AI models. How are researchers curating these data to maximize AI output?
- Data governance
How do we access this data, standardize it, make it available to any researcher who wants to improve models while guaranteeing privacy and patient safety? The umbrella that encompasses the usage of data is what I’ll call data governance. How to structure, organize, and make biological data into standard formats and standard repositories is nothing new and has been a topic for years that has produced many useful public databases but again, this is still in its infancy. This is worth a deep dive in a future post but where data is stored and how it is accessed and by whom will be a pivotal discussion as AI models become more prevalent in biological research. Biology is a hodgepodge of standards for how the data is prepared, how data sets are annotated, the permissions for accessing data, and much more. Without standardization, there could be consequential batch effects that make data sets incompatible, never mind an AI model trained on that data. Of course, researchers must consider the privacy and ethical implications of an enormous amount of people’s personal health and biological data going into AI models. Given controversies surrounding training LLMs on the collective human output while a tiny number of people and companies profit from it, the current public outcry may pale in comparison to what happens when AI companies start using all the bio data in the world. As everyone knows, your genomic signature is unique to you, the most powerful identifier you have. HIPAA, the data privacy law, was itself only codified into law in the 1990s and almost certainly is inadequate to confront data fed into AI models.
- Model interpretability and clinical utility
Even some of the top AI researchers don’t fully understand why a model produces one answer over another. A field of AI research called interpretability allows researchers to proverbially pop the hood on the model and take a look. What does this practically look like for an AI model trained on cancer bio data? Is it safe or ethical to make a prediction on a treatment for a patient if you have no idea why the AI model came to that conclusion? How do you test the reliability and repeatability of these discoveries? I’m sure AI researchers have tools to address interpretability issues, but for me, it is one huge potential drawback of any AI-led approach and something all biomedical researchers need to think about when deciding to use these tools. And what if an AI model hallucinates an incorrect prediction, or even more terrifying, actively attempts to deceive its handlers (not beyond the realm of possibility given some startling new research coming out of the AI labs)? These false positives could not only waste scientific resources and time, but also put patients at risk if such aberrant behavior is not fully understood and appropriately controlled (if at all possible). Ultimately, regardless of model interpretability, most AI predictions on treatments and new molecular targets will need to be tested the old-fashioned way: with time-consuming and expensive randomized, controlled clinical trials.
- Regulatory Issues
As someone who currently works in the current good manufacturing practices (cGMP) space, when it comes to making a pharmaceutical product that is intended for humans, the controls and documentation required to guarantee safety are nothing short of staggering. Even something as conceptually simple as filling a finished drug into a vial requires systematic and controlled documentation for every step of the process. This includes analytical testing, equipment qualifications, sterility testing, cleaning validations, and much more. Indeed, every parameter of the process must be controlled, documented, and justified with evidence. And this level of rigor is just for the last stage of preparation of the final drug product, never mind all the prior steps it took to get there. This reality applies to any biomedical product. The path from discovery to patient is a long, complicated, and expensive one. AI may help streamline and improve some of these steps but right now, I am skeptical about it actually accelerating this path. Granted, this is an area that I am still learning about so maybe I’m wrong? But one thing is certain: the already overworked regulators at the FDA may have their hands full with a whole slew of new AI-based tools and how you actually prove that what an AI model comes up with is safe for humans. Likely, the same onerous approval process and same high-bar burden of proof will be required (i.e. clinical trials, rigorous documentation and evidence, etc.).
Conclusions
Cancer bio has not yet had its “GPT-3” moment (at least not to my knowledge), but that doesn’t mean it’s not soon on the horizon. I think the frontier AI labs are racing to build the most powerful AI models imaginable. For biology though, I think diffusion and experimentation with the technology as it exists is the bigger bottleneck, not the capability. So far, there are no therapies that are primarily designed and directed by AI that have achieved FDA approval (though there are some AI-designed drugs currently in late-stage clinical trials). The above limitations further argue that the radical changes we could see in cancer and other fields of biology may not be apparent for years. Not because AI is not powerful enough to help now, but because of the pure data, regulatory and testing hurdles involved. Another possibility is that AI models may hit a wall of capability based on the amount of useful data currently available. Is AI really the revolution it promises to be or just another tool? I’m optimistic about the future of biology and the advances AI can bring, but as with anything in science, you need to prove it first.